Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Apr 13;16:17120. doi: 10.1038/s41598-026-42748-w

Analysis of hyperparameter optimization effects on lightweight deep models for real-time image classification

Vineet Kumar Rakesh 1,2,✉,#, Soumya Mazumdar 3,#, Tapas Samanta 1,2,#, Hemendra Kumar Pandey 1,2, Amitabha Das 4, Sarbajit Pal 5
PMCID: PMC13230992  PMID: 41974758

Abstract

Lightweight convolutional and transformer-based networks are increasingly used for real-time image classification on resource-constrained hardware, yet their practical performance is highly sensitive to training hyperparameters. This work systematically quantifies how controlled hyperparameter choices affect both accuracy and deployability for seven modern lightweight backbones–ConvNeXt-Tiny, EfficientNetV2-S, MobileNetV3-L, MobileViT v2 (S/XS), RepVGG–A2, and TinyViT-21M trained from scratch on a class-balanced 90K/10K subset of ImageNet-1K under a standardized 300-epoch protocol. We isolate the effects of learning-rate magnitude and cosine scheduling, optimizer selection (SGD vs. AdamW where appropriate), and progressively stronger regularization via RandAugment, Mixup, CutMix, and label smoothing, complemented by constrained automated searches (Optuna and population-based training). Beyond training-time analysis, we add a deployment-focused evaluation: inference latency and throughput are benchmarked on an NVIDIA L40s GPU across batch sizes 1–512, and edge feasibility is examined via Edge CPU Platform under sustained workloads. Results show that hyperparameter tuning without architectural modification yields consistent accuracy gains (Inline graphic Top-1 over baseline) and reveals architecture-dependent stability regions. Several models deliver strong real-time operating points: MobileNetV3–L and RepVGG–A2 achieve very low latency with high throughput on GPU, while edge tests highlight the limited benefit of batching on low-power CPUs and the importance of latency-centric model choice. The code and logs may be seen at: https://github.com/VineetKumarRakesh/lcnn-opt.

Subject terms: Engineering, Mathematics and computing

Introduction

Real-time image categorization on edge and peripheral devices requires deep learning models that achieve strong predictive performance under strict constraints on compute, memory, and latency. This has motivated sustained interest in lightweight architectures–typically under roughly 30 million parameters–that aim to balance accuracy with deployability. Contemporary lightweight model families span efficient convolutional neural networks and hybrid CNN–transformer designs, each offering different trade-offs in representational capacity, optimization behavior, and inference efficiency. However, despite rapid architectural progress, practical performance in real-world deployments remains highly sensitive to training hyperparameters, and comparisons across models are often confounded by differences in training recipes rather than architecture alone. As a result, automated machine learning (AutoML)1 techniques have been increasingly explored to reduce manual tuning effort and standardize training pipelines across architectures.

This work systematically studies seven widely used lightweight backbones selected for architectural diversity and broad adoption: EfficientNetV2-S2, ConvNeXt-Tiny3, MobileNetV3-L4, MobileViT v25, TinyViT-21M6, and RepVGG–A27. To enable controlled iteration while preserving class diversity, all models are trained and evaluated on a class-balanced subset of ImageNet–1K consisting of 90,000 training images (with a held-out evaluation split) sampled uniformly across all 1,000 classes8. This experimental design supports rapid and reproducible comparisons under identical constraints, though absolute accuracy values are not intended to be directly compared with full ImageNet benchmarks reported in prior work6,9.

The central hypothesis of this study is that systematic hyperparameter optimization can yield consistent improvements in both accuracy and deployment-relevant behavior for lightweight models without architectural modification. Rather than treating automated hyperparameter optimization (HPO) as a black-box procedure, we isolate and quantify the sensitivity of key training decisions across architectures under a shared training budget. Specifically, we examine the impact of initial learning-rate scale and scheduling, optimizer choice (SGD with momentum versus AdamW where appropriate), and progressively stronger regularization via modern augmentation and smoothing strategies. In this context, Optuna10 provides a flexible, sampler-based framework for efficient hyperparameter search under limited budgets, while Population-Based Training (PBT)11 enables joint optimization of hyperparameters and weights through online evolutionary adaptation during training.

Across all studied architectures, the initial learning rate emerges as a dominant factor in determining final performance. Moderate increases–ranging from 0.001 toward 0.1 depending on the optimizer and model–consistently improve Top-1 accuracy, while overly large values destabilize optimization and degrade convergence, indicating a bounded region of stable training. Generalization further improves with a composite augmentation pipeline incorporating RandAugment, Mixup, CutMix12, and label smoothing, producing steady gains over minimal baseline recipes. These findings align with prior observations that training configuration choices can materially influence final performance, often rivaling architectural changes in their impact7,13,14.

To ensure efficient training throughput and consistent optimization behavior, experiments employ large-batch training (batch size 512) on an NVIDIA L40s-class GPU, enabling stable convergence under cosine learning-rate schedules and modern regularization. We observe architecture-dependent optimization dynamics: larger or more expressive backbones, such as TinyViT-21M6 and EfficientNetV2-S2, often reach strong accuracy earlier in training, particularly under cosine scheduling, while hybrid or transformer-leaning models, including ConvNeXt-Tiny3 and TinyViT-21M, frequently benefit from faster early-stage convergence when optimized with AdamW. With well-tuned schedules and regularization, however, both SGD and AdamW converge to comparable final accuracy across models.

Overall, the results demonstrate that careful hyperparameter selection yields consistent absolute accuracy improvements of approximately 1.5–2.5% across all evaluated lightweight backbones. These findings underscore that training configuration is a first-order determinant of real-time deployment viability alongside architectural design. By evaluating multiple lightweight models under identical experimental constraints and explicitly characterizing hyperparameter sensitivity, this work helps disentangle architectural effects from training-recipe effects and provides practical guidance for training lightweight models intended for edge deployment.

This work makes four main contributions: (1) a controlled, architecture-spanning study of seven widely used lightweight backbones (CNN and hybrid CNN–transformer) trained under an identical compute budget and experimental protocol; (2) an ablation-driven characterization of hyperparameter sensitivity—including initial learning rate, cosine scheduling, optimizer choice (SGD vs. AdamW), and modern regularization and augmentation components (RandAugment, Mixup, CutMix, and label smoothing)—to disentangle training-recipe effects from architectural effects; (3) evidence that systematic tuning without architectural modification yields consistent absolute Top-1 accuracy gains (approximately 1.5–2.5%) across all evaluated models on a class-balanced ImageNet–1K subset; and (4) deployment-oriented benchmarking that connects training choices to practical efficiency, reporting latency and throughput on an NVIDIA L40s across batch sizes and CPU-only feasibility results on a Raspberry Pi 4 under sustained inference workloads.

This paper is constructed as follows: Section 2 surveys related work; Sections 3 detail model design and setup; Sections 4 to 5 present analysis and results respectively; Section 6 presents the limitations and Section 7 closes with practical suggestions.

Related work

Lightweight image classification models

To balance accuracy and low latency, several efficient convolutional neural network (CNN) architectures have been introduced. Notable early models include SqueezeNet15, MobileNets4,16,17, ShuffleNet18, and EfficientNetV213. MobileNetV3, developed using Neural Architecture Search, integrates h-swish activation and squeeze-and-excitation modules to improve the accuracy-latency trade-off. The MobileNetV3-L4 variant (5.4M parameters) achieves 75.2% top-1 ImageNet–1K accuracy, outperforming MobileNetV2 by 3.2% with 20% reduced latency4.

EfficientNetV2 scales depth, width, and resolution jointly to optimize performance under resource constraints. EfficientNetV2-S2 (22M parameters) enhances this design using progressive image resizing and better augmentation, reaching 83.9% top-1 accuracy on ImageNet–1K2. RepVGG–A27, a VGG-style architecture re-parameterized post-training, achieves over 80% accuracy using modern augmentations. The RepVGG–A27 (25M parameters) matches or exceeds ResNet-50 in throughput and accuracy on ImageNet.

Transformers, though resource-intensive, have shown strong performance in image classification19. MobileViT9 combines lightweight CNNs and self-attention, achieving 78.4% top-1 accuracy with just 5.6M parameters. MobileViT v2 improves inference and accuracy using separable self-attention, attaining 75.6% accuracy with 3M parameters5.

TinyViT-21M6, distilled from Swin Transformers, achieves up to 84.8% top-1 accuracy with pretraining. Trained from scratch, it maintains strong performance (83.1%) and excels on tasks like COCO object detection (50.2 mAP). ConvNeXt-Tiny3, a modern CNN inspired by transformers, uses large kernels and LayerNorm. ConvNeXt-Tiny3 (29M parameters) achieves 82.1% top-1 accuracy at Inline graphic resolution, scalable to 87.8% for larger variants3, proving CNNs remain competitive with modern training techniques.

Hyperparameter optimization and training strategies

Model performance heavily depends on hyperparameters and training strategies. Augmentation methods such as AutoAugment20 and RandAugment21 improve accuracy by 1–2% on ImageNet. Mixup18 and CutMix12 enhance generalization and robustness. For instance, CutMix improved ResNet-50 accuracy from 76.3% to 78.6%. Label smoothing22, by softening target labels, typically adds 0.2–0.5% gains.

Many lightweight models integrate these strategies. ConvNeXt-T adopted the DeiT-style training with Mixup, CutMix, RandAugment, and label smoothing3. EfficientNetV2 used AutoAugment and label smoothing13. RepVGG–A27 benefited from extended training with Mixup and aggressive augmentation, surpassing 80% accuracy7. MobileNetV3 used simpler augmentations, suggesting room for further gains via modern techniques.

Learning rate scheduling is critical. Cosine annealing23, now common in ConvNeXt-T and transformer training, smoothly decays the learning rate and enhances convergence stability. Often paired with warm-up phases, it avoids the abrupt changes seen in step schedules. Optimizer choice also matters: while SGD with momentum works well for CNNs, AdamW24 is preferred for transformers due to faster convergence. Models like ConvNeXt-T use AdamW with gradient clipping, cosine schedules, and weight decay (typically 0.05–0.1).

Batch size plays a crucial role. Larger batches can speed up training but may require learning rate adjustments. Following the linear scaling rule25, increasing batch size requires proportional learning rate increases and gradual warm-up for stability. For example, moving from batch size 256 to 1024 can shorten training time while maintaining accuracy with appropriate tuning.

This study emphasizes that hyperparameter optimization is central to high–performance lightweight models. Rather than proposing new architectures, the paper systematically evaluates existing ones under optimized training regimes. Effective augmentation, regularization, learning rate scheduling, and optimizer settings collectively enhance model accuracy. Bacanin et al.26 used a firefly algorithm to optimize CNN hyperparameters for brain tumor MRI classification. Iqbal et al.27 demonstrated real–time coronary artery disease detection using streamlined CNNs with tuned hyperparameters. These studies reinforce that training optimization is vital for deploying efficient deep learning models in real–world, resource–constrained environments. While these models achieve strong accuracy, their performance under controlled hyperparameter variation remains underexplored motivating this work. An overview of the iterative hyperparameter-optimization workflow used in this study is shown in Fig. 1.

Fig. 1.

Fig. 1

Iterative Workflow of Hyperparameter Optimization (HPO).

Methodology

To reach high accuracy with low computational delay, numerous efficient convolutional neural network (CNN) designs have been developed. Early families of models designed for mobile and embedded vision tasks include SqueezeNet15, MobileNets4,16,17, ShuffleNet18, and EfficientNetV213.

Lightweight model selection

Based on their effectiveness, popularity, and design diversity, seven cutting-edge architectures were chosen to investigate how hyperparameter optimization affects lightweight deep learning models. A overview of each model, including parameter counts, reported ImageNet–1K Top-1 accuracy, and initial training setups, is given in Table 1. Selection was based on architectural diversity covering CNN, hybrid, and transformer families, with parameter counts under 30 M to ensure comparability in edge constraints. EfficientNetV2-S2 is a convolutional model with 22 million parameters, trained using neural architecture search and sophisticated methods. It achieves 83.9% Top-1 accuracy on ImageNet-1K2. ConvNeXt-Tiny3 is a contemporary ConvNet with 29 million parameters and 4.5 GFLOPs, improved using Transformer-style training techniques. It achieves 82.1% Top-1 accuracy on ImageNet-1K3. MobileViT v2 (XS)9 is a hybrid CNN-Transformer model with 2.9 million parameters and 75.6% Top-1 accuracy5. MobileViT v2 (S)5 is a scaled-up version of MobileViT v2, with 5–6 million parameters and enhanced accuracy of 78–79% on ImageNet5. MobileNetV3-L4 is a conventionally effective CNN with 5.4 million parameters and 219 MFLOPs, achieving 75.2% Top-1 accuracy on ImageNet4. TinyViT-21M6 is a Vision Transformer model with 21 million parameters, achieving 84.8% accuracy with distillation-based pretraining6. RepVGG–A27 is a VGG-like model with 25 million parameters, achieving 78.4% accuracy with baseline training and 80.4% with vigorous augmentation7.

Table 1.

The summary includes the number of parameters, ImageNet-1K Top-1 accuracy under original training (224Inline graphic224 unless noted), and notable training hyperparameters used in original works.

Model Params (M) Top-1 accuracy Original training highlights
ConvNeXt–T3 29 82.1% 300 epochs, AdamW, cosine LR, RandAug, Mixup, CutMix, LS 0.1
EfficientNetV2 -S2 22 83.9% 350 epochs, RMSProp, progressive resize 224Inline graphic480, RandAug, LS 0.1
MobileNetV3–L4 5.4 75.2% 300 epochs, RMSProp, cosine LR, AutoAug, SE modules
MobileViT v2 (S)5 5.6 78.5% 300 epochs, AdamW, cosine LR, heavy augmentation
MobileViT v2 (XS)9 2.9 75.6% 300 epochs, AdamW, ImageNet-1K pretrain, RandAug, LS 0.1
RepVGG–A27 25 78.4–80.4% 120–240 epochs, SGD, step/cosine LR, AutoAug, Mixup, LS 0.1
TinyViT–21M6 21 84.8% 210+90 epochs, AdamW, cosine LR, heavy Aug, distillation

System configuration

GPU benchmark platform

All experiments were conducted on an NVIDIA L40s GPU (48 GB, CUDA 12.6) using PyTorch 2.5.1 with automatic mixed precision (AMP). The software stack includes Python 3.10.18 managed via Anaconda, and experiments were implemented in PyCharm Community Edition v2025.1.2.

Edge device (Raspberry Pi 4)

To evaluate real edge-device behavior, inference benchmarks were additionally performed on a Raspberry Pi 4 Model B (2018) with a Broadcom BCM2711 quad-core Cortex-A72 (ARMv8) CPU at 1.8 GHz and 4 GB LPDDR4-3200 RAM. We report CPU-only inference at 224Inline graphic224 resolution under the same preprocessing pipeline, and we note that optimized runtimes (e.g., TFLite/ONNX Runtime) and quantization may change absolute latency/FPS.

Desktop CPU benchmark platform

Inference benchmarks on CPU were performed on a desktop system equipped with an Intel(R) Core(TM) i7-10700 CPU @ 2.90 GHz, 32 GB RAM, and an Intel(R) UHD Graphics 630 (iGPU). All CPU measurements are reported using PyTorch inference on the CPU (no discrete GPU acceleration), with input resolution Inline graphic.

Dataset

To reduce computational burden while preserving diversity across classes, we used a representative subset of ImageNet-1K comprising approximately 90,000 training and 10,000 validation images uniformly sampled across 1,000 categories. This design choice was motivated by the primary goal of the study: to compare hyperparameter sensitivity trends consistently across multiple lightweight architectures and optimization strategies under a fixed and feasible experimental budget. Running full ImageNet-1K training for 300 epochs across many configurations (learning rates, batch sizes, optimizers, and augmentation stacks) and across multiple models would be prohibitively expensive and would limit the breadth of controlled ablations and repeated trials. The subset therefore enables (i) extensive ablation coverage, (ii) repeatability across runs (e.g., multi-seed evaluation where needed), and (iii) fair, like-for-like comparisons across model families using the same data protocol. Importantly, we do not claim that the resulting absolute accuracies are directly comparable to full ImageNet-1K benchmarks. Rather, the subset is intended to preserve relative performance trends and sensitivity patterns under consistent conditions. To support representativeness, we retained class balance and evaluated distributional similarity to the original label distribution (via per-class sample entropy and KL-divergence computed over labels). Top-1 accuracy is reported as the primary metric, with Top-5 accuracy included for additional context. All models were trained from scratch on this 90,000-image subset for 300 epochs.

Training configuration and pre-processing

Unless otherwise specified, all models were trained using stochastic gradient descent (SGD) with momentum of 0.9, an initial learning rate (LR) of 0.1 (scaled appropriately), cosine-annealing learning rate schedule over 300 epochs, and weight decay of Inline graphic. For architectures originally trained with AdamW (e.g., ConvNeXt-Tiny3 and TinyViT), we followed the authors’ recommended AdamW settings: an initial learning rate of Inline graphic, Inline graphic, Inline graphic, and the model-specific weight decay used in the original training recipe. Training was performed using a global batch size of 512, enabling efficient utilization of GPU resources and stable convergence during model optimization. Input images were normalized using ImageNet–1K statistics and augmented with random resized cropping and horizontal flipping as the baseline. Additional augmentations (RandAugment, Mixup, CutMix) and regularization (Label Smoothing) were incrementally introduced to isolate their effects. All pipelines used timm implementations for consistency.

Implementation of ablation study

An extensive ablation study was performed by altering one hyperparameter at a time from a fixed baseline configuration to assess its individual impact on model accuracy. Key factors analyzed include:

  • Initial Learning Rate and Scheduler: Values from 0.001 to 0.100 were tested. Cosine annealing schedules provided smoother convergence and consistently better final accuracy than step decay. A short 5-epoch warm-up was applied for high learning rate settings.

  • Batch Size: All primary ablation and optimization experiments were conducted using a fixed global batch size of 512 (with learning-rate scaling applied as appropriate). Under this controlled setting, little difference in final accuracy was observed when scaling was performed correctly, while batch size 512 offered the best trade-off between training speed, numerical stability, and hardware utilization on the NVIDIA L40s GPU. A separate batch-size sweep was conducted solely for diagnostic analysis and is reported independently in Section 4.2.

  • Optimizer: SGD was effective for CNN-based models (e.g., MobileNetV3-L, RepVGG–A2), while AdamW showed superior convergence for transformer-based and hybrid models (e.g., ConvNeXt-T, MobileViT v2 (S/XS).

  • Data Augmentation: The impact of augmentations was studied cumulatively. RandAugment provided early accuracy gains; Mixup and CutMix improved generalization further; Label Smoothing offered consistent minor gains with no training cost.

  • Training Epochs: Although 300 epochs were used as the standard, models such as MobileNetV3-L and RepVGG–A27 continued to improve under stronger augmentations and were therefore trained for the full 300 epochs using the standardized schedule.

All experiments used consistent data pipelines and PyTorch-based logging, enabling reproducibility and direct comparison across settings.

Ablation study: hyperparameter effects

This paper investigates the impact of hyperparameters and training methodologies on model efficacy, concentrating on seven typical models: ConvNeXt-Tiny3, MobileViT v2 (XS)9, and MobileNetV3-L4. These models exemplify convolutional networks, transformer hybrids, and mobile-optimized convolutional neural networks. The research revealed that alternative models had comparable behaviors in response to changes in hyperparameters. The main quantitative findings for learning rate and augmentation experiments are encapsulated in Tables 2 and 5 for the selected models. This section examines the impact of critical hyperparameters and training strategies on model performance for lightweight real-time image classification. The analysis focuses on seven representative models–ConvNeXt-T, EfficientNetV2-S2, MobileNetV3-L, MobileViT v2 (S), MobileViT v2 (XS), RepVGG–A2, and TinyViT-21M–spanning convolutional backbones, transformer hybrids, and mobile-optimized architectures. Experimental outcomes are presented through quantitative results in Tables 2 and 5, and graphical insights in Figures 2 and 3.

Table 2.

Top-1 validation accuracy (%) across three initial learning-rate regimes (Inline graphic) under a fixed 300-epoch training budget with cosine annealing (no restarts). The table is used to compare end-of-training performance across LR regimes, while fixed-stage convergence behavior is analyzed in Table 3.

Model LR = 0.001 LR = 0.010 LR = 0.100
ConvNeXt-Tiny3 83.81 83.73 83.61
EfficientNetV2-S2 88.31 88.50 88.29
MobileNetV3-L4 86.94 87.00 86.92
MobileViT v2 (S)5 87.53 87.53 87.08
MobileViT v2 (XS)9 87.36 87.12 85.81
RepVGG–A27 88.41 88.45 88.22
TinyViT-21M6 90.82 90.94 90.75

Table 5.

Top-1 validation accuracy (%) of representative models trained for 300 epochs on the ImageNet–1K subset. Results show cumulative augmentation ablations under a fixed manual training configuration (Baseline Inline graphic RandAug Inline graphic Mixup Inline graphic CutMix Inline graphic Label Smoothing), alongside best-performing configurations obtained via automated hyperparameter optimization (Optuna and Population-Based Training [PBT]) under constrained search budgets. For Optuna and PBT, the reported value is the best validation accuracy and the corresponding trial number is shown in brackets, where T denotes the trial index. All Optuna and PBT results are reported using a total of 20 trials. Bold values indicate the highest accuracy achieved for each model across augmentation or optimization strategies.

Model Baseline Manual Augmentations Optuna PBT
+ RandAug + Mixup + CutMix + Label Smooth
ConvNeXt-Tiny3 83.85 86.24 86.90 88.50 88.00 90.40 (T12) 88.07 (T6)
EfficientNetV2-S2 88.50 91.34 92.72 92.63 92.56 84.34 (T9) 89.53 (T6)
MobileNetV3-Large4 86.99 89.15 90.97 90.45 90.20 72.76 (T13) 85.40 (T6)
MobileViT v2 (S)5 87.83 89.91 91.47 92.63 91.28 66.19 (T12) 79.52 (T6)
MobileViT v2 (XS)9 87.36 88.88 90.56 90.18 90.31 89.16 (T4) 84.47 (T6)
RepVGG–A27 88.45 89.61 91.54 91.48 91.43 83.56 (T11) 88.01 (T6)
TinyViT–21M6 90.94 92.11 93.30 93.35 93.84 91.99 (T19) 79.88 (T6)

Fig. 2.

Fig. 2

Accuracy vs Learning Rate for all evaluated models under consistent training settings.

Fig. 3.

Fig. 3

Validation accuracy (%) across epochs (1 to 300) for each model on ImageNet–1K subset.

Learning rate and scheduler

The learning rate (LR) is one of the most influential hyperparameters governing both convergence dynamics and final model accuracy28. In the original analysis, validation accuracy was summarized using the epoch of peak performance within a training window. While such peak-epoch reporting is useful for diagnosing optimization stability, it can conflate convergence speed with end-of-training performance and may introduce selection bias when validation curves are noisy. To address this concern, the revised analysis emphasizes fixed-stage accuracy trends across training (early, mid, and late stages) and the final-epoch behavior, thereby decoupling convergence dynamics from final performance.

Learning rates were evaluated across three logarithmically spaced regimes: Inline graphic. All models were trained for a fixed budget of 300 epochs using cosine annealing without restarts, ensuring comparable optimization trajectories across learning-rate settings. The resulting LR decay is summarized at representative epochs in Table 3, showing the transition from large early updates to small late-stage refinement.

Table 3.

Merged learning-rate schedule and fixed-stage validation dynamics in baseline runs. For each sampled epoch E, we report the instantaneous learning rate (LR), Top-1 validation accuracy (Val Acc, %), and the absolute validation accuracy change Inline graphic. For epoch 10, Inline graphic equals Val Acc@10. For later stages, Inline graphic.

Model Epoch 10 Epoch 60 Epoch 120 Epoch 200 Epoch 260 Epoch 300
LR Val Acc Inline graphic LR Val Acc Inline graphic LR Val Acc Inline graphic LR Val Acc Inline graphic LR Val Acc Inline graphic LR Val Acc Inline graphic
ConvNeXt-Tiny3 0.099778 4.56 4.56 0.090757 78.11 73.55 0.065951 81.59 3.48 0.025462 82.95 1.36 0.004548 83.69 0.74 1.27Inline graphic 83.85 0.16
EfficientNetV2-S2 0.099778 61.52 61.52 0.090757 85.25 23.73 0.065951 86.58 1.33 0.025462 87.79 1.21 0.004548 88.39 0.60 1.27Inline graphic 88.30 −0.09
MobileNetV3-Large4 0.099778 53.85 53.85 0.090757 82.37 28.52 0.065951 84.34 1.97 0.025462 86.20 1.86 0.004548 86.84 0.64 1.27Inline graphic 86.88 0.04
MobileViT v2 (S)5 0.099426 67.44 67.44 0.082995 84.04 16.60 0.057327 86.17 2.13 0.040436 87.36 1.19 0.032546 87.28 −0.08 0.024527 87.53 0.25
MobileViT v2 (XS)9 0.045392 60.05 60.05 0.094489 82.26 22.21 0.049101 85.27 3.01 0.033298 86.59 1.32 0.008881 87.08 0.49 0.000889 87.30 0.22
RepVGG–A27 0.099778 53.51 53.51 0.090757 83.91 30.40 0.065951 85.78 1.87 0.025462 87.50 1.72 0.004548 88.24 0.74 1.27Inline graphic 88.42 0.18
TinyViT-21M6 0.099778 76.36 76.36 0.090757 87.99 11.63 0.065951 89.50 1.51 0.025462 90.44 0.94 0.004548 90.86 0.42 1.27Inline graphic 90.80 −0.06

Rather than relying solely on best-epoch accuracy, learning-rate sensitivity is interpreted using validation accuracy progression at fixed training stages, as reflected in the convergence trajectories (Figure 3) and quantified in Table 3. Across all architectures, validation accuracy exhibits a rapid increase during the initial training phase, followed by gradual stabilization. Importantly, this early improvement is quantified using the measured gains up to epoch 60:

graphic file with name d33e892.gif

which ranges from 69.42 to 84.36 percentage points across the evaluated models (Table 3). After epoch 60, improvements become progressively smaller (typically Inline graphic percentage points per interval beyond epoch 120), consistent with a refinement regime induced by the decayed step size under cosine scheduling.

End-of-training behavior across LR regimes is summarized in Table 2. Moderate learning rates (Inline graphic) yield comparable final accuracy across architectures, whereas smaller learning rates slow early convergence without providing measurable gains under the same 300-epoch budget. Conversely, learning rates beyond the stable region reduce optimization stability and degrade performance, which is consistent with the sharp decline observed in the accuracy-versus-LR trend (Figure 2). Overall, these results indicate a bounded LR region that supports stable convergence under a fixed training budget.

The revision improves rigor by removing peak-epoch selection and reporting fixed-stage trends; this directly separates convergence dynamics from final performance. When computational budget permits, each Inline graphic configuration can be repeated across multiple random seeds and summarized with mean ± 95% confidence intervals to quantify variability in peak and final accuracy.

All of the models show a rapid increase in validation accuracy in the early training phase (up to epoch 60). Then, the learning rate steadily makes their performance more stable. This pattern indicates that feature learning starts off quickly and subsequently slows down, but it keeps continuing through architectures.

Batch size scaling

All primary ablation experiments were trained using a fixed batch size of 512 per GPU, leveraging the full memory capacity of the NVIDIA L40s GPU (48 GB). This large-batch regime provided stable gradient estimates, enabled proportionally scaled learning rates, and facilitated high-throughput training. The training stability and consistent convergence illustrated in Figures 2 and 3 are attributable in part to this controlled large-batch configuration. Larger models such as TinyViT-21M6 and EfficientNetV2-S2 converged in fewer epochs, reaching stable accuracy plateaus earlier in training. These observations align with established practices suggesting that batch size and learning rate must be jointly scaled to preserve training dynamics. Mixed-precision training further improved resource efficiency, allowing all models to be trained without memory bottlenecks.

To further analyze batch-size sensitivity, we conducted a separate diagnostic batch-size sweep in which otherwise fixed optimized configurations were evaluated across multiple batch sizes. In this auxiliary analysis, models consistently achieved high accuracy across batch sizes, with smaller batch sizes (e.g., 32) occasionally yielding marginally higher peak validation accuracy. This behavior is consistent with prior observations that increased gradient noise at smaller batch sizes can improve generalization. Importantly, these results do not redefine the primary training configuration, but rather contextualize the robustness of the models to batch-size variation.

Table 4 summarizes the best accuracies obtained during evaluation, where the maximum Top-1 and Top-5 scores for each model correspond to the most effective batch size configuration. It can be observed that while all architectures deliver competitive results, TinyViT-21M6 achieves the highest Top-1 (90.94%) and Top-5 (97.74%) accuracies, with ConvNeXt-Tiny3 slightly trailing behind.

Table 4.

Model accuracy results from a diagnostic batch-size sweep. All models were trained under the same optimized hyperparameter configuration, and the reported batch size corresponds to the setting that yielded the highest validation accuracy during this auxiliary evaluation. These results are reported to analyze batch-size sensitivity and are distinct from the fixed batch-size (512) configuration used in the primary ablation experiments.

Model Batch Size Top-1 (%) Top-5 (%)
ConvNeXt-Tiny3 32 83.85 95.09
EfficientNetV2-S2 32 88.50 97.15
MobileNetV3-Large4 32 86.99 96.93
MobileViT v2 (S)5 32 87.82 97.19
MobileViT v2 (XS)9 32 87.36 96.80
RepVGG–A27 32 88.45 97.16
TinyViT-21M6 32 90.94 97.74

These findings reinforce that batch size is not merely a computational parameter but also influences optimization dynamics. While smaller batch sizes may yield marginally higher peak validation accuracy in diagnostic sweeps, larger batch sizes provide superior stability, throughput efficiency, and reproducibility. Accordingly, a batch size of 512 was adopted as the fixed configuration for all primary experiments in this study. Together, these results highlight the importance of distinguishing between batch-size sensitivity analysis and baseline training methodology when evaluating real-time image classification models.

Larger models such as TinyViT-21M6 and EfficientNetV2-S2 converged in fewer epochs, with higher accuracy plateaus observed in early training phases. These observations align with established practices suggesting that batch size and learning rate must be jointly scaled to maintain training dynamics. Mixed-precision training further contributed to resource efficiency, allowing deeper models to be trained without memory bottlenecks. Although this study evaluated a single batch size of 512, this choice was informed by empirical guidelines for large-batch training and the capabilities of the NVIDIA L40s GPU (48 GB Memory), which ensured stable convergence, efficient resource utilization, and consistent throughput across all models. Future work may investigate scaling behaviors across varied sample sizes and hardware platforms, the results indicate its applicability for large-scale, real-time classification tasks.

Data augmentation and regularization

The effects of data augmentation and regularization strategies are summarized in Table 5. We evaluate a cumulative augmentation pipeline in which widely used techniques–RandAugment, Mixup, CutMix, and Label Smoothing–are added sequentially to a fixed baseline training setup. This approach enables a controlled examination of progressively stronger regularization, without assuming that the benefits of individual augmentations are strictly additive. As shown in Table 5, computational constraints limited the Optuna and PBT experiments to 20 trials each. While a larger search budget could yield higher absolute performance, the observed trends and the conclusions drawn from the analysis are expected to remain unchanged. Manual augmentation outperforms automated optimization because it incorporates strong prior knowledge, exhibits stable long-horizon training behavior, and applies architecture-aware regularization. In contrast, automated methods are restricted by limited trial budgets and relatively coarse exploration of the search space.

The chosen sequential augmentation order is for interpretability and does not guarantee optimality across all architectures. Overall, augmentation substantially improves Top-1 accuracy across all evaluated architectures relative to the baseline, confirming its importance for generalization in lightweight models. However, the magnitude and monotonicity of gains vary by both method and architecture, highlighting the presence of interaction effects between augmentation strategies and model capacity.

This behavior can be attributed to overlapping regularization mechanisms. Mixup and CutMix both enforce linear behavior in feature space through label interpolation, while CutMix additionally disrupts spatial structure. In smaller or strongly regularized models, such aggressive perturbations may induce mild underfitting, whereas higher-capacity architectures (e.g., RepVGG–A27 and TinyViT-21M) are better able to absorb these transformations and consistently benefit from composite augmentation pipelines. Label smoothing further illustrates diminishing returns: when applied on top of strong sample-level augmentations, its effect is often neutral or modest, as confidence regularization is already implicitly provided.

The table also reports results from automated hyperparameter optimization using Optuna and Population-Based Training (PBT). These configurations explore broader search spaces under constrained optimization budgets and are not restricted to the manual augmentation pipeline, explaining performance differences relative to hand-tuned composite augmentation. Accordingly, Optuna and PBT results should be interpreted as complementary rather than directly comparable augmentation ablations.

Figure 3 corroborates these findings by showing smoother convergence trajectories and higher final accuracy plateaus for models trained with richer augmentation strategies, despite occasional non-monotonic improvements at intermediate stages.

Taken together, these results demonstrate that while data augmentation is a powerful tool for improving accuracy without architectural changes, its effectiveness depends on model capacity and interaction effects between regularizers. For practitioners with limited tuning budgets, RandAugment and Mixup provide the most consistent gains across architectures, while CutMix and Label Smoothing are most beneficial for larger or more expressive models. Despite being trained on a 90K-image subset (approximately Inline graphic of ImageNet–1K), the relative trends observed here remain consistent and provide a reliable basis for rapid experimentation and hyperparameter selection in real-time and resource-constrained settings.

Results and discussion

This section presents a detailed analysis of the performance improvements achieved through systematic hyperparameter optimization across seven lightweight deep learning models. All models were trained for a fixed duration of 300 epochs using the same subset of ImageNet-1K (90,000 training images and 10,000 validation images), ensuring a fair and controlled comparison. The choice of 300 epochs was not arbitrary: five of the seven evaluated architectures originally reported training schedules of approximately 300 epochs in their respective papers, while the remaining models were trained for longer durations. To place all architectures on a common and comparable footing, we standardized the training length to 300 epochs across all experiments.

All experiments were conducted on a high-performance computing node equipped with an NVIDIA L40s GPU (48 GB), CUDA 12.6, and PyTorch 2.5.1, using automatic mixed precision (AMP). Optimizer and weight decay follow the baseline definitions, with architecture-consistent settings (e.g., SGD for CNN backbones; AdamW where used by original authors).Specifically, we adopted the following shared configuration: input size: 224, batch size: 512, workers: 8, pin_memory: true, RandAugment: true, Mixup: 0.2, CutMix: 1.0, label smoothing: 0.1, optimizer: SGD/AdamW(model-dependent), Learning rate: 0.1 for SGD runs and Inline graphic for AdamW runs, momentum: 0.9, scheduler: cosine annealing, epochs: 300, and minimum learning rate: Inline graphic. These values are consistent with those reported or implied in the original model publications and widely adopted training practices for ImageNet classification, as summarized in Table 1. For automated hyperparameter optimization using Optuna and Population-Based Training (PBT), the search space was restricted to augmentation-related parameters and limited to 20 trials per model due to computational constraints; while a larger search budget may further improve absolute performance, the observed trends and comparative conclusions remain consistent. The best trial values are reported in the table 6.

Table 6.

Best augmentation hyperparameters suggested by Optuna and PBT for each model. For Optuna, values are taken from the best trial (trial index shown as T). For PBT, only augmentation enablement is reported due to unavailable coefficient logs. Best validation accuracy (%) is taken from Table 5.

Model Optuna (best trial) PBT (best member)
T RandAug MixUp CutMix Label Smooth Best Val Acc T RandAug Label Smooth Best Val Acc
ConvNeXt-Tiny3 T26 True 0.110 0.472 0.057 90.40 T6 True False 88.07
EfficientNetV2-S2 T9 True 0.271 0.269 0.130 84.34 T6 True True 89.53
MobileNetV3-Large4 T20 False 0.198 0.518 0.097 72.76 T6 True True 85.40
MobileViT v2 (S)5 T15 True 0.387 0.174 0.181 66.19 T6 True True 79.52
MobileViT v2 (XS)9 T4 True 0.135 0.420 0.014 89.16 T6 True True 84.47
RepVGG–A27 T11 False 0.062 0.626 0.074 83.56 T6 True True 88.01
TinyViT–21M6 T19 True 0.132 0.396 0.171 91.99 T6 True True 79.88

The input resolution of Inline graphic was selected as it is the canonical ImageNet setting for lightweight architectures, enabling direct comparison while maintaining moderate computational cost. A batch size of 512 was chosen to exploit the full memory capacity of the L40s GPU and to provide stable gradient estimates during training. RandAugment was enabled to introduce stochastic yet low-overhead transformations without requiring policy search. Mixup and CutMix were applied to improve robustness through sample interpolation and spatial patch mixing, while label smoothing mitigated overconfident predictions and improved calibration. Weight decay followed the baseline definitions in Section 3.4 (SGD runs: Inline graphic; AdamW runs: Inline graphic–Inline graphic, as in the original training recipes). SGD with momentum 0.9 was selected due to its demonstrated stability and effectiveness for convolutional and hybrid architectures. A high initial learning rate of 0.1, decayed via cosine annealing toward a minimum of 1e–5, enabled rapid early convergence while preserving fine-grained optimization in later epochs; this schedule consistently outperformed step-based decays in both stability and final accuracy. In the table 6, Optuna reports explicit numeric values for all augmentation coefficients. PBT logs only record which augmentations were enabled; therefore MixUp and CutMix magnitudes are not shown. Best validation accuracy values are taken from Table 5.

Training all models for a fixed 300-epoch schedule ensured complete convergence while avoiding architecture-specific stopping criteria that could bias comparisons. By fixing these hyperparameters–each motivated by prior work and original author recommendations–we ensure that observed improvements in Top-1 and Top-5 accuracy arise from architectural efficiency and hyperparameter interaction effects, rather than tuning bias. Under this controlled setting, the selected configuration, particularly the combination of cosine scheduling, composite augmentation, and SGD optimization, yielded consistent gains of approximately 1.5–3.5% over baseline configurations.

Finally, to complement training-time analysis, we extend the evaluation with a newly added inference and deployment study, including GPU-based throughput measurements and edge-device benchmarking on a Raspberry Pi 4. These results connect hyperparameter optimization with real-world deployment constraints, strengthening the practical relevance of the analysis for resource-constrained and real-time applications.

Performance gains from hyperparameter optimization

Hyperparameter optimization significantly improved model accuracy across all evaluated architectures. Table 7 delineates the maximum attained Top-1 and Top-5 validation accuracies, minimal latency, and peak throughput (FPS) recorded for each model. All outcomes are obtained from logs of the training sessions and assessment iterations. In addition to accuracy and inference metrics, the table also reports the number of parameters (Params (M)) and floating-point operations (FLOPs (G)), providing a quantitative measure of model complexity and computational cost.

Table 7.

Optimized performance metrics on the ImageNet-1K 90K-image subset. Reported accuracy corresponds to the final validation epoch. Latency and FPS denote the minimum and maximum values, respectively, measured over 10,000 inference iterations across the evaluated batch sizes. The symbol B indicates the batch size at which peak inference performance was observed (maximum FPS/minimum latency). Arrows denote whether higher (Inline graphic) or lower (Inline graphic) values are preferable for each metric.

Model Top-1 (%) Top-5 (%) Latency (ms)Inline graphic FPSInline graphic Params (M)Inline graphic FLOPs (G)Inline graphic
ConvNeXt-Tiny3 83.85 95.09 0.51 Inline graphic 1964.99 Inline graphic 28.566 4.456
EfficientNetV2-S2 88.50 97.15 0.31 Inline graphic 3226.66 Inline graphic 21.305 2.85
MobileNetV3-Large4 86.99 96.93 0.10Inline graphic 10034.10Inline graphic 4.178 0.215
MobileViT v2 (S)5 87.82 97.19 0.40 Inline graphic 2516.01 Inline graphic 4.878 1.412
MobileViT v2 (XS)9 87.36 96.80 0.33 Inline graphic 3007.27 Inline graphic 1.359 0.362
RepVGG–A27 88.45 97.16 0.26 Inline graphic 3862.14 Inline graphic 28.206 5.685
TinyViT–21M6 90.94 97.74 0.59 Inline graphic 1687.04 Inline graphic 21.206 4.091

Table 7 reflects the final optimized configuration for each model after applying all effective hyperparameter tuning strategies. These values are not directly comparable to ablation-stage results in Table 2, which isolate individual factors.

TinyViT-21M6 achieved the highest Top-1 accuracy (90.94%) and Top-5 accuracy (97.74%), while also maintaining competitive GPU inference performance (Table 7). EfficientNetV2–S and RepVGG–A2 followed closely with Top-1 accuracies of 88.50% and 88.45%, respectively. MobileViT v2 (S)5 and MobileViT v2 (XS)9 earned Top-1 scores of 87.82% and 87.36%, respectively, providing a balanced tradeoff between performance, parameter count, and computational overhead.

Inference speed and latency trends

To evaluate deployment feasibility, we benchmarked each model for inference performance. FPS and latency were measured on the same L40s GPU using PyTorch (eager mode, AMP), with input size 224Inline graphic224 and batch sizes– 1, 16, 32, 64, 128, 256, 512. Latency is reported as time per image in milliseconds, and FPS is computed as effective throughput. FPS and latency metrics were averaged across 10,000 runs. Variability (standard deviation < 1.5%) was minimal due to the controlled hardware setup. Although variability across repeated runs was minimal (standard deviation < 1.5%) under the controlled GPU setup, these benchmarks do not capture thermal or power-related effects that arise during sustained inference on edge devices. Moreover, all GPU measurements were performed in PyTorch eager mode with automatic mixed precision (AMP). In practical deployments, optimized runtimes such as TensorRT, ONNX Runtime, or TFLite may significantly alter latency and throughput characteristics. Accordingly, the reported GPU benchmarks should be interpreted as framework-consistent reference measurements rather than deployment-optimal results. A detailed study of runtime-specific optimizations, thermal behavior, and power efficiency on edge hardware is left for future work.

Edge inference on Raspberry Pi 4

To complement the high-throughput GPU benchmarks and to better reflect edge deployment constraints, we additionally evaluated inference performance on a Raspberry Pi 4 Model B (4 GB RAM). This platform represents a commonly used low-power edge device with strict thermal and power budgets. All models were evaluated using CPU-only inference with input resolution Inline graphic, and latency was measured under sustained inference workloads to capture realistic edge behavior.

Unlike the GPU benchmarks, which emphasize peak throughput, the Raspberry Pi evaluation focuses on per-image latency and sustained execution stability. Due to limited memory bandwidth and the absence of dedicated acceleration, batch size was fixed to 1, reflecting typical real-time edge deployment scenarios. Measurements were averaged over repeated runs after an initial warm-up period to reduce transient effects.

In addition to single-image inference, we analyzed the impact of batch size on throughput and latency on the Raspberry Pi 4 to understand whether batching offers practical benefits on low-power edge CPUs. Figures 6 and 7 report the relationship between batch size, effective throughput (FPS), and per-image latency.

Fig. 6.

Fig. 6

Batch Size vs Throughput (FPS) on Raspberry Pi 4. Throughput increases modestly with batch size and saturates quickly, reflecting limited CPU parallelism on edge devices.

Fig. 7.

Fig. 7

Latency vs Batch Size on Raspberry Pi 4. Latency increases sharply at larger batch sizes due to queuing and memory pressure, limiting batching benefits for real-time edge inference.

Unlike the GPU results, increasing batch size on the Raspberry Pi yields only modest throughput gains and leads to rapidly increasing latency. This behavior reflects the limited parallelism and memory bandwidth available on general-purpose ARM CPUs, where batching primarily increases queuing delay rather than computational efficiency. For real-time edge applications, batch sizes greater than one offer limited benefit and may violate latency constraints.

These observations highlight a fundamental distinction between server-class GPUs and edge CPUs: while batching is effective for maximizing GPU throughput, edge deployments typically favor batch size 1 to maintain predictable latency, thermal stability, and power efficiency. Accordingly, batch-size scaling should be treated as a deployment-specific optimization rather than a universally transferable strategy.

As expected, lightweight CNN-based models such as MobileNetV3-Large4 and RepVGG–A27 exhibited the lowest latency on the Raspberry Pi, making them more suitable for continuous edge inference. Transformer-based and hybrid architectures incurred higher latency, reflecting their greater reliance on memory access and attention operations, which are less efficient on general-purpose CPUs. While absolute latency values differ significantly from GPU results, the relative ordering of models remains consistent with their computational complexity.

Thermal considerations are particularly important on embedded platforms. During sustained inference, we observed gradual frequency scaling on the Raspberry Pi CPU under prolonged workloads, indicating that thermal throttling can influence long-running performance. This highlights that edge deployment feasibility depends not only on raw model efficiency but also on thermal design and power management, factors that are outside the scope of GPU-centric benchmarks.

These results emphasize that while GPU benchmarks are useful for understanding upper-bound performance, CPU-based edge evaluations provide complementary insight into real-world deployability. We therefore present the Raspberry Pi results as a qualitative reference for edge suitability rather than a direct comparison with GPU throughput figures.

Desktop CPU inference benchmark (Intel i7-10700)

To complement the GPU and Raspberry Pi evaluations, we benchmarked CPU-only inference on a desktop Intel(R) Core(TM) i7-10700 CPU @ 2.90 GHz with PyTorch CPU execution at Inline graphic input resolution. For each model and batch size, we report per-image latency (ms/image) and effective throughput (FPS). Batch sizes follow powers of two to capture both latency-centric (batch size 1) and throughput-centric (larger batch) operating regimes.

Across all models, throughput improves from batch size 1 to a small-to-moderate batch regime (typically 4–16) due to amortization of framework overhead and better utilization of CPU threading/vectorization. However, unlike GPUs, CPU throughput saturates quickly as cache capacity and memory bandwidth become the dominant bottlenecks. Beyond the saturation point, increasing batch size yields diminishing returns in FPS and often increases total batch time substantially, which is undesirable for real-time latency constraints.

For MobileViT v2 (S/XS) and TinyViT-21M, the benchmark run for batch size 512 was initiated but did not produce a stable timing record in the log (i.e., no corresponding [RESULT] entry was recorded). In practice, very large batches on CPU can become infeasible for hybrid/transformer-leaning models because intermediate activation tensors (e.g., attention projections and feature maps across stages) scale with batch size and can trigger severe memory pressure. On a CPU, this manifests as large transient allocations, cache thrashing, and potential paging (swap), which can cause the run to stall or be terminated by the runtime/OS. Since CPU throughput already saturates at much smaller batch sizes, we cap these curves at batch size 256 where valid measurements were obtained and report only completed points in the plots. The desktop CPU throughput and latency trends across batch sizes are summarized in Figures 8 and 9, respectively.

Fig. 8.

Fig. 8

CPU FPS vs batch size. Throughput (frames per second) on the desktop CPU across batch sizes.

Fig. 9.

Fig. 9

CPU latency vs batch size. Per-image latency (ms/image) on the desktop CPU. MobileNetV3-Large4 provides the lowest latency at batch size 1 and remains stable across larger batches, while larger/hybrid models show higher latency and earlier saturation.

Model suitability for real-time use

Among all tested models, MobileNetV3–L and RepVGG–A2 exhibited the greatest balance of high accuracy and low latency as shown in figure 4 and 5. At batch size 1, MobileNetV3–L maintained latency under 3.2 ms, while scaling to over 9800 FPS at batch size 512. These qualities make such models excellent for embedded deployment in mobile, surveillance, or industrial robotics applications.

Fig. 4.

Fig. 4

Throughput (FPS): FPS across varying batch sizes [1, 16, 32, 64, 128, 256, 512] for each model on NVIDIA L40s. Values averaged over 10,000 iterations per batch.

Fig. 5.

Fig. 5

Inference Latency: Inference latency (ms) across batch sizes on an NVIDIA L40s GPU with 224Inline graphic224 input resolution. Latency measured in PyTorch eager mode using AMP.

Reproducibility and fair comparison

All experiments were conducted under controlled hardware and software environments, with the same input pipeline and augmentation strategies. The use of a class-balanced subset, though not equivalent to full ImageNet–1K, provides a reproducible and time-efficient benchmark for comparative hyperparameter evaluation. Full training logs and benchmarking scripts are included in the supplementary repository.

Augmentation interaction effects and diminishing returns

The augmentation ablation presented in Table 5 utilizes a cumulative design to assess the incremental impact of various regularization techniques. However, due to overlapping mechanisms among data augmentations like RandAugment, Mixup, CutMix, and Label Smoothing, their effects are not strictly additive, potentially resulting in diminishing returns or mild regression with multiple strong regularizers. For instance, while both Mixup and CutMix promote linearity through label interpolation, the latter may disrupt spatial coherence, leading to underfitting in models like MobileViT v2-XS. In contrast, larger models, such as TinyViT-21M6 and RepVGG–A2, benefit from CutMix, indicating that model capacity is crucial for handling spatial changes. Label Smoothing also shows limited benefits when strong sample-level augmentations are applied. These findings emphasize that augmentation strategies interact with model capacity and architecture, and highlight that improvements are not guaranteed with increased regularization. The sequence in Table 5 is representative of a standard augmentation pipeline and not an optimal combination for all architectures.

Limitation

A key limitation of the study is conducting experiments on a class-balanced subset of 90 K images from ImageNet–1K due to GPU constraints, affecting the direct comparability of accuracy values with full-scale benchmarks. The study’s primary focus is on analyzing relative hyperparameter sensitivity and interaction patterns–such as convergence stability and trade-offs with batch size–rather than on achieving dataset-specific performance records. While optimal hyperparameter values may change with dataset size, previous studies suggest that qualitative trends remain consistent across subsets. Thus, findings serve as guidance for efficient hyperparameter selection and robust regimes, with final adjustments needed for full dataset training.

Conclusion

This study isolates and quantifies hyperparameter sensitivity across multiple lightweight architectures, providing a reproducible benchmark for real-time deployment analysis. By training on a carefully curated, class-balanced 90,000–image subset of ImageNet–1K, we systematically assessed the influence of various hyperparameter strategies—including learning rate schedules, data augmentation methods, and optimizer selection—on both model accuracy and inference efficiency.

Our results demonstrate that hyperparameter tuning alone, without modifying the model architecture, can lead to substantial performance gains. Across the evaluated architectures–ConvNeXt–Tiny, EfficientNetV2 –S, MobileNetV3–Large, MobileViT v2 (S and XS), RepVGG–A2, and TinyViT-21M–optimized configurations resulted in Top-1 accuracy improvements ranging from 1.5% to 3.5% compared to baseline settings. These gains were achieved through the combined use of cosine learning rate scheduling, modern augmentations such as CutMix and RandAugment, and adaptive optimizers like AdamW. Notably, TinyViT-21M6 achieved the highest Top-1 accuracy of 90.94%, followed closely by EfficientNetV2-S2 and RepVGG–A2, with 88.50% and 88.45%, respectively.

In addition to classification accuracy, we benchmarked all models for real-time deployment feasibility using throughput (frames per second) and latency (milliseconds per image) metrics. These benchmarks were conducted on an NVIDIA L40s GPU using PyTorch’s eager mode with Automatic Mixed Precision (AMP). Batch sizes ranged from 1 to 512, reflecting both low-latency and high-throughput scenarios. Results showed that models such as MobileNetV3-L and RepVGG–A27 consistently delivered below 5 ms average latency and throughput exceeding 9,000 FPS at higher batch sizes, making them strong candidates for deployment on edge devices, mobile platforms, and embedded systems with strict latency requirements.

However, while our findings are promising, it is important to acknowledge the limitations of the study. The use of a 90K-image subset, although balanced and reproducible, does not fully replicate the complexity and diversity of the complete ImageNet-1K dataset. As such, absolute accuracy values should not be directly compared with those reported in original papers trained on the full dataset. Instead, our study focuses on relative improvements due to hyperparameter optimization and their implications for model efficiency and deployability.

To support reproducibility and future research, all training logs, evaluation scripts, YAML configuration files, and dataset sampling protocols have been included in the supplementary materials. These artifacts allow for the exact replication of experiments and facilitate transferability of our findings to other domains such as medical imaging, robotics, and low-power AI systems.

The research illustrates that even with confined datasets and conventional training hardware, considerable performance and efficiency benefits are feasible with systematic hyperparameter optimization. While our benchmarks were conducted on a high-end NVIDIA L40s GPU, future work should explore deployment feasibility on edge CPUs (e.g., Raspberry Pi 5, ARM Cortex-A series) or mobile NPUs (e.g., Qualcomm Hexagon, Apple Neural Engine). Such analysis would help assess the generalizability of optimized lightweight models across diverse hardware settings. Our subset-based assessment approach gives a useful benchmark for lightweight model selection and optimization. The insights discovered here are particularly relevant for real-time applications in limited contexts, and we hope they serve as a platform for subsequent explorations into efficient deep learning at the edge.

Supplementary Information

Acknowledgements

This work was supported by the Variable Energy Cyclotron Centre (VECC), Department of Atomic Energy (DAE), Government of India (GoI), and the Homi Bhabha National Institute (HBNI), Department of Atomic Energy (DAE), Government of India (GoI), for providing comprehensive facilities and technical support essential to this research. The Department of Atomic Energy (DAE), Government of India (GoI), is also appreciated for financing the open-access publishing of this study. The authors appreciate the VECC library for their essential support over the course of this research.

Author contributions

V.K.R. conceived the core idea and methodology, performed all training and benchmarking experiments, and generated the visualizations. S.M. led the evaluation process, created ablation plots and tables, and contributed to writing and editing the manuscript. T.S. implemented the benchmarking scripts for latency and throughput analysis, assisted with profiling, and prepared reproducibility resources. H.K.P. curated the ImageNet-1K subset, developed the data loader and augmentation pipeline, and handled training log management. A.D. contributed to model selection, reviewed the training configurations for hardware efficiency, and validated the final results. S.P reviewed and approved the final version of the manuscript. All authors reviewed and approved the manuscript.

Funding

The Department of Atomic Energy (DAE), Government of India (GoI), is also appreciated for financing the open-access publishing of this study.

Data availability

All source code, configuration files, and supplementary scripts used in this study are publicly available at https://github.com/VineetKumarRakesh/lcnn-opt. The ImageNet-1K dataset is available under its original license and cannot be redistributed by the authors. Due to institutional data-sharing restrictions, detailed training logs and key result files are not openly available but can be accessed upon reasonable request to the corresponding author at vineet@vecc.gov.in.

Declarations

Competing interests

The authors declare no competing interests. Several authors are affiliated with VECC and HBNI (under the Department of Atomic Energy, Government of India), which supported the open-access publication charges; this support did not influence the study design, results, interpretation, or conclusions.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally to this work: Vineet Kumar Rakesh, Soumya Mazumdar and Tapas Samanta.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-026-42748-w.

References

  • 1.Salehin, I. et al. AutoML: A systematic review on automated machine learning with neural architecture search. Journal of Information and Intelligence2, 52–81. 10.1016/j.jiixd.2023.10.002 (2023). [Google Scholar]
  • 2.Tan, M. & Le, Q. V. Efficientnetv2: Smaller models and faster training. arXiv preprint arXiv:2104.00298 (2021).
  • 3.Liu, Z. et al. Convnext: Revisiting resnets at scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4817–4827 (2022).
  • 4.Howard, A. et al. Searching for mobilenetv3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 1314–1324 (2019).
  • 5.Mehta, S., Nguyen, N., Rastegari, M., Shapiro, L. & Hajishirzi, H. Separable self-attention for mobile vision transformers (mobilevitv2). Transactions on Machine Learning Research (2023).
  • 6.Wu, K. et al. Tinyvit: Fast pretraining distillation for small vision transformers. In European Conference on Computer Vision (ECCV), 68–85 (2022).
  • 7.Ding, X. et al. Repvgg: Making vgg-style convnets great again. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13733–13742 (2021).
  • 8.Russakovsky, O. et al. Imagenet large scale visual recognition challenge. Int. J. Comput. Vis.115, 211–252 (2015). [Google Scholar]
  • 9.Mehta, S. & Rastegari, M. Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer. arXiv preprint arXiv:2110.02178 (2021).
  • 10.Akiba, T., Sano, S., Yanase, T., Ohta, T. & Koyama, M. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’19, 2623–2631 10.1145/3292500.3330701(Association for Computing Machinery, New York, NY, USA, 2019).
  • 11.Li, A. et al. A generalized framework for population based training. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’19, 1791–1799, 10.1145/3292500.3330649(Association for Computing Machinery, New York, NY, USA, 2019).
  • 12.Yun, S. et al. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 6023–6032 (2019).
  • 13.Tan, M. & Le, Q. V. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (ICML), 6105–6114 (2019).
  • 14.Loshchilov, I. & Hutter, F. Sgdr: Stochastic gradient descent with warm restarts. In Proceedings of the 5th International Conference on Learning Representations (ICLR) (2017).
  • 15.Iandola, F. N. et al. Squeezenet: Alexnet-level accuracy with 50 fewer parameters and 0.5mb model size. arXiv preprint arXiv:1602.07360 (2016).
  • 16.Howard, A. G. et al. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017).
  • 17.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A. & Chen, L.-C. Mobilenetv 2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4510–4520 (2018).
  • 18.Zhang, H., Cisse, M., Dauphin, Y. N. & Lopez-Paz, D. mixup: Beyond empirical risk minimization. In Proceedings of the 6th International Conference on Learning Representations (ICLR) (2018).
  • 19.Dosovitskiy, A. et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the 9th International Conference on Learning Representations (ICLR) (2021).
  • 20.Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V. & Le, Q. V. Autoaugment: Learning augmentation policies from data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 113–123 (2019).
  • 21.Cubuk, E. D., Zoph, B., Shlens, J. & Le, Q. V. Randaugment: Practical automated data augmentation with a reduced search space. In Advances in Neural Information Processing Systems (NeurIPS), 18613–18624 (2020).
  • 22.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J. & Wojna, Z. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2818–2826 (2016).
  • 23.Loshchilov, I. & Hutter, F. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983 (2016).
  • 24.Loshchilov, I. & Hutter, F. Decoupled weight decay regularization. In Proceedings of the 7th International Conference on Learning Representations (ICLR) (2019).
  • 25.Goyal, P. et al. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677 (2017).
  • 26.Bacanin, N., Bezdan, T., Venkatachalam, K. & Al-Turjman, F. Optimized convolutional neural network by firefly algorithm for magnetic resonance image classification of glioma brain tumor grade. J. Real-Time Image Process.18, 1085–1098 (2021). [Google Scholar]
  • 27.Iqbal, T., Khalid, A. & Ullah, I. Explaining decisions of a lightweight deep neural network for real-time coronary artery disease classification in magnetic resonance imaging. J. Real-Time Image Process.10.1007/s11554-023-01411-7 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Goodfellow, I., Bengio, Y. & Courville, A. Deep Learning (MIT Press, 2016). http://www.deeplearningbook.org.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

All source code, configuration files, and supplementary scripts used in this study are publicly available at https://github.com/VineetKumarRakesh/lcnn-opt. The ImageNet-1K dataset is available under its original license and cannot be redistributed by the authors. Due to institutional data-sharing restrictions, detailed training logs and key result files are not openly available but can be accessed upon reasonable request to the corresponding author at vineet@vecc.gov.in.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES