Skip to main content
Sensors (Basel, Switzerland) logoLink to Sensors (Basel, Switzerland)
. 2026 Feb 11;26(4):1157. doi: 10.3390/s26041157

DKTransformer: An Accurate and Efficient Model for Fine-Grained Food Image Classification

Hongjuan Wang 1, Chenxi Wang 1, Xinjun An 1,*
Editor: Sheryl Berlin Brahnam1
PMCID: PMC12944211  PMID: 41755097

Abstract

With the rapid development of dietary analysis and health computing, food image classification has attracted increasing attention. However, this task remains challenging due to the fine-grained nature of food categories. Different classes are visually similar, whereas samples within the same class exhibit large appearance variations. Existing methods often rely excessively on either global or local features, limiting their effectiveness in complex food scenes. To address these challenges, this paper proposes DKTransformer, a lightweight hybrid architecture that combines Vision Transformers (ViT) and convolutional neural networks (CNNs) for fine-grained food image classification. Specifically, DKTransformer introduces a Local Feature Extraction (LDE) module based on depthwise separable convolution to enhance local detail modeling. Furthermore, a Multi-Scale Dilated Attention (MSDA) module is designed to capture long-range dependencies with reduced computational cost while suppressing background interference. In addition, an Efficient Kolmogorov–Arnold Network (EfficientKAN) is employed to replace the conventional feedforward network, further reducing parameter redundancy. Experimental results on three public food image datasets—ETH Food-101, Vireo-Food-172, and ISIA Food-500—demonstrate the effectiveness of the proposed method. In particular, DKTransformer achieves a Top-1 accuracy of 92.71% on the ETH Food-101 dataset with 47 M parameters and 7.21 G FLOPs. Moreover, DKTransformer attains 90.70% Top-1 accuracy on Vireo-Food-172 and 66.89% on Food-500, indicating strong generalization across different food styles and dataset scales. These results suggest that DKTransformer achieves a favorable balance between accuracy and efficiency for fine-grained food image classification.

Keywords: food recognition, ViT, local feature, lightweight

1. Introduction

Given the fundamental role of food in human health, food computing has emerged as a prominent research area within the multimedia field in recent years [1]. As one of its core tasks, food image retrieval addresses the challenge of dynamic category expansion (e.g., incorporating new food types) by enabling efficient matching of query images against large-scale databases [2,3]. Unlike general-purpose image classification tasks (e.g., ImageNet), where categories exhibit large semantic and visual differences, food image analysis is a fine-grained classification problem. In this setting, different food categories are often visually similar, while samples within the same category can have significant appearance variations. As shown in Figure 1, under this setting, existing food recognition methods still face two major challenges: (1) irrelevant information interference [4], where food images often contain non-subject regions such as tableware and side dishes, which can distract the model from focusing on core ingredients; (2) ambiguity of fine-grained differences [5], where similar food products present diverse visual appearances due to variations in cooking styles, while dissimilar food categories may share highly similar ingredient compositions, making it difficult for traditional classification networks to capture subtle yet discriminative visual cues.

Figure 1.

Figure 1

Example food images. The red dashed boxes highlight visually similar food categories with subtle inter-class differences.

To address the above problems, existing research mainly relies on convolutional neural networks (CNNs) [6,7] to extract local features or Vision Transformers (ViT) [8] to model global dependencies. However, both types of methods have significant limitations: CNNs lack the ability to model long-distance semantic associations, while the multi-head self-attention (MSA) mechanism of ViT, although capable of capturing global context, suffers from high computational complexity and potential loss of local details, making it difficult to apply directly to lightweight scenarios. To this end, this paper proposes DKTransformer, a hybrid architecture that combines the advantages of ViT and CNN, to realize efficient fine-grained classification through three core improvements:

(1) Local Feature Extraction (LDE) module: enhances the capture of local texture and edge features by embedding depthwise separable convolution [9] within the token sequence;

(2) Multi-scale dilated attention module (MSDA) [10]: utilizes a sparse attention mechanism with multiple dilation rates to cover a larger receptive field while reducing computation and suppressing background clutter interference;

(3) Efficient Kolmogorov–Arnold Network (EfficientKAN) [11]: employs parameter sharing with sparse activation regularization to reduce the number of parameters while guaranteeing the model’s expressiveness.

The proposed DKTransformer is evaluated on several large-scale public food datasets. Detailed experiments and model results analyses demonstrate that the proposed network is reliable and efficient for lightweight food image classification tasks.

2. Related Works

2.1. CNN-Based Fine-Grained Food Image Classification Study

CNNs mimic the human visual system through a hierarchical feature extraction mechanism (from underlying edges/textures to higher-level semantic features) [12] and demonstrate strong potential in food image recognition tasks. Earlier studies mainly implemented feature extraction through pre-trained model migration, e.g., Ming et al. [13] directly utilized ResNet to extract global visual features, and McAllister et al. [14] fused deep features from ResNet-152 and GoogleNet for food classification. While such methods leverage the generalization ability of pre-trained models, global feature-dominated representations struggle to capture fine-grained differences between ingredients (e.g., the texture difference between a steak and a pork chop). To further improve results, researchers have proposed fine-tuning strategies, where network parameters are tuned on the target dataset (e.g., ShuffleNetv2 [15], MobileNetV2 [16], MobileNetV3 [17]) to enhance the model’s sensitivity to local discriminative features. However, the inherent local receptive field limitation of CNNs leads to insufficient long-range semantic association modeling, restricting classification accuracy in complex scenarios.

2.2. Application and Challenges of ViT in Food Image Classification

The core idea of ViT is to split the input image into fixed-size patches and map them into token sequences by linear projection. Assuming an input image resolution of H × W and a patch size of P × P, the number of tokens is N=(H×W)/P2. The token sequences are summed with learnable one-dimensional positional embeddings and fed into the Transformer encoder. The encoder consists of multiple stacked Transformer layers, each containing a MSA module and a FeedForward Network (FFN), with normalization layers and skip connections applied after both MSA and FFN. The output of each encoder layer can be denoted as:

AttentionQ,K,V=SoftmaxQKTCV (1)
FFNx=σxW1+b1W2+b2 (2)
y=LNx+FFNx,x=LNx+MSAx (3)

where Q, K, V are query, key, and value matrices, respectively; W1 and W2 are linear transformation weights; b1, b2 are biases; and σ is the Gaussian Error Linear Unit (GELU) nonlinear activation function. ViT captures long-range dependencies between image patches through global self-attention, but its computational complexity grows quadratically with the number of tokens, and it can be weak in modeling local details such as ingredient textures.

ViT provides a new idea for food recognition by modeling global dependencies between image patches through self-attention. SwinT [18], EfficientFormer [19], BiFormer [20], EfficientViT [21], and SHViT [22] reduce the computational overhead of ViT by optimizing the attention computation process, but high complexity and local detail loss problems still limit its practical application. For ViT lightweighting, MobileViTv2 [23,24], RepViT [25] and Fastvit [26] adopt parameter sharing and attention sparsification strategies to reduce computational overhead, but they still require a large number of floating-point operations (FLOPs). For the food recognition task, Sheng et al. [27] designed a parallel architecture of ViT and CNN to account for global context and local features, but the model parameter count and FLOPs remain high; Song et al. [28] introduced metric learning to strengthen local feature expression, improving fine-grained classification accuracy but failing to solve the computational efficiency bottleneck; Liu [29] proposed a multi-task learning framework to jointly optimize food category and ingredient recognition, but this further increases model complexity.

The above studies demonstrate that it is difficult to balance the accuracy–efficiency trade-off by relying solely on ViT or CNNs. Recent hybrid architecture studies [30,31] attempted to fuse ViT and CNN, but redundant computation between modules (e.g., feature fusion for parallel branches) resulted in efficiency loss. For instance, Sheng et al. proposed a location-preserving ViT within a hybrid network for food recognition [32]. In summary, existing methods have not yet reached an effective balance between lightweight design and fine-grained classification results, necessitating innovative designs that synergize the advantages of both model types.

In addition to standard architectures, recent research has extensively explored the synergy between multi-level feature fusion and attention mechanisms to optimize food recognition. Notably, Chen [33] proposed the Multi-level Attention Feature Fusion Network (MAF-Net), which integrates feature maps from distinct CNN stages utilizing a self-attention mechanism. By simultaneously capturing global contexts and local discriminative features, this method achieved superior performance across multiple food datasets. These findings validate the effectiveness of combining hierarchical feature representations with attention strategies for fine-grained food classification tasks.

Beyond the specific domain of food computing, the efficacy of hybrid deep learning architectures has been substantially validated in other challenging fine-grained classification tasks. Contemporary studies consistently demonstrate that merging the local feature extraction capabilities of CNNs with the global dependency modeling of Transformers yields robust performance in complex scenarios. For instance, in medical image analysis, Negi [34] and Tewari [35] successfully integrated ResNet backbones with Transformer modules, significantly enhancing feature representation for skin cancer and diabetic retinopathy detection by capturing both subtle textures and global contexts. Addressing the computational overhead of standard attention mechanisms, Alam [36] employed an improved Swin Transformer V2 for precise brain tumor classification, underscoring the necessity of efficient attention strategies in high-resolution processing. Similarly, in the realm of spatio-temporal modeling, Zheng [37] proposed a multi-scale framework (AST-GNNFormer) that effectively fuses features at different granularities. These cross-disciplinary advancements provide a strong theoretical foundation for the hybrid design philosophy adopted in our DKTransformer, motivating the integration of multi-scale dilated attention with KAN to tackle the intricate inter-class variations inherent in food imagery.

3. Methodology

3.1. Diagram of DKTransformer Network Structure

The proposed DKTransformer employs an improved DKBlock as the basic unit to construct the encoder. As illustrated in Figure 2, after patch embedding, the input image undergoes feature extraction through cascaded DKBlocks. Downsampling layers are inserted at different stages to construct a multi-scale feature pyramid [21,22], preserving high-frequency details (e.g., ingredient texture) at shallow layers and capturing abstract semantics (e.g., overall shape) at deeper layers. Specifically, we adopt an overlapping downsampling strategy to minimize information loss during dimension reduction. At the transition between stages (e.g., from Stage 1 to Stage 2), we employ a 3×3 convolution layer with a stride of 2. Since the kernel size (3) is larger than the stride (2), this creates overlapping receptive fields between adjacent sliding windows, ensuring better feature continuity compared to non-overlapping patch merging. This operation effectively halves the spatial dimensions of the feature map (H,W→H/2,W/2) while expanding the channel dimension, enabling the model to capture broader receptive fields and more abstract semantic information at deeper layers. Notably, this operation does not involve pooling layers, aiming to preserve more structural information through learnable convolution weights. Furthermore, normalization is not applied independently within the downsampling layer but is unified via Layer Normalization in the subsequent DKBlock. This design ensures the progressive abstraction and fusion of local features while avoiding extra parameters and computational overhead.

Figure 2.

Figure 2

Schematic diagram of the DKTransformer network and DKBlock. Different colors are used only to visually distinguish different components for clarity.

This strategy mimics the hierarchical processing mechanism of the human visual system while reducing computational redundancy. Compared to the standard ViT, the DKBlock incorporates three core improvements to achieve a balance between accuracy and efficiency: LDE module, MSDA module and EfficientKAN. Input features are progressively refined through these three modules, ultimately generating representations that incorporate both local details and global semantics. For fine-grained food classification tasks, the model first focuses on local ingredient textures via the LDE module, then suppresses background noise through multi-scale contextual modeling with MSDA, and finally compresses redundant computations using the parameter-efficient feedforward network. This process simulates the human visual system’s cognitive progression from local to global, while simultaneously balancing accuracy and efficiency.

3.2. Local Feature Extraction

Although ViT excels at capturing global dependencies, its standard tokenization—a linear projection of image patches—often fails to capture subtle local patterns crucial for fine-grained tasks like food recognition. Discriminative details, such as the texture of beef fibers or breadcrumb granularity, can be overlooked, hindering the distinction between visually similar categories. To bridge this gap between global understanding and local precision, this paper proposes LDE, which embeds depthwise separable convolution into the token sequence to efficiently capture fine-grained local features. As illustrated in Figure 3, depthwise separable convolution decomposes standard convolution into depthwise and pointwise operations, significantly reducing computational complexity while preserving local feature representation. As shown in Figure 4, LDE splits the input token sequence into patch tokens xp∈RB∗N∗C and the class token xc∈RB∗H∗W∗C. For notation, B denotes the batch size, N is the number of patch tokens, C is the token embedding dimension, and H×W is the spatial token grid with H×W=N. It then applies depthwise separable convolution with kernel size k to the spatially reshaped patch tokens to extract local features xfp:

xfp= DepthwiseConvReshape2Dxp, kernel=k∗k (4)

Figure 3.

Figure 3

Schematic diagram of depthwise separable convolution. Different colors are used only to visually distinguish different components for clarity.

Figure 4.

Figure 4

Schematic diagram of the LDE module. Arrows indicate the data flow, and different colors are used only to visually distinguish different components for clarity.

Finally, xfp is flattened back into a 1D sequence space and connected with xc, serving as the input to the next layer:

xout=ConcatReshape1D(xfp),xc∈RB∗(N+1)∗C (5)

Given N patch tokens with embedding dimension C, we prepend a learnable class token and add positional embeddings to all tokens. As a result, the input to the Transformer encoder is x∈RB∗(N+1)∗C.

3.3. MSDA Module for Attention to Differential Regions

Traditional multi-head self-attention incurs high computational and memory resource consumption due to dense computation and sensitivity to noise (as shown in the heatmap of Figure 5; the heatmap in Figure 5 is generated from the class-token attention of the last layer and upsampled to the input image size for visualization.), making it difficult to focus effectively on core ingredient regions. MSDA reduces computation while covering a larger receptive field and suppressing noise interference from cutlery, side dishes, etc., through a strategy combining sparse localized attention and multi-scale receptive fields.

Figure 5.

Figure 5

Visualization of attention heatmaps highlighting the core ingredient focus of standard ViT.

To further enhance the robustness and generalization capability of the MSDA module, the DropKey mechanism is incorporated. DropKey is a simple yet effective regularization method that randomly drops (i.e., sets the attention scores to negative infinity) a portion of the key vectors during training. This forces the model not to over-rely on specific local features, thereby improving its robustness to noise and background interference. This method is particularly suitable for complex backgrounds and diverse ingredient combinations commonly found in food images. The core difference between DropKey and standard Dropout lies in their underlying mechanisms and impact on probability distribution. Standard Dropout typically acts on the attention weights after Softmax, randomly zeroing out elements, which inevitably disrupts the unity of the probability distribution. In contrast, DropKey operates before the Softmax layer by masking a portion of the Key vectors (setting corresponding similarity scores to −∞). This allows the subsequent Softmax operation to re-normalize the attention map, ensuring the total probability remains 1. By suppressing over-concentrated attention peaks on local features, this ‘drop-and-renormalize’ mechanism encourages the model to capture more balanced and global semantic dependencies. Specifically, before computing the attention scores, we apply the DropKey operation to the key vectors in each attention head.

Let the drop rate be p; during the training phase, each key vector is randomly dropped with probability p. The modified attention computation in the MSDA module is defined as:

xij=Attentionqij,Kr,Vr=SoftmaxqijkrTdk+MVr (6)

In Equation (6), Q, K and V represent query, key, and value matrices, respectively; M represents the DropKey mask matrix, where elements corresponding to dropped keys are set to −∞ and others to 0. For the query vector qij at position i,j, r denotes the dilation rate. The keys and values are sparsely selected within a sliding window centered at i,j with size ω∗ω for self-attention computation. The set of locations i′,j′ utilized for keys and values is:

i′,j′i′=i+p∗r,j′=j+q∗r,−ω2≤p,q≤ω2 (7)

To exploit sparsity at different scales, the MSDA module extracts semantic information at multiple receptive fields. The structure of MSDA is shown in Figure 6:

Figure 6.

Figure 6

Schematic diagram of the MSDA module. The red cell denotes the center point, the dark blue rings represent different dilation coefficients ri, and the yellow cells indicate the DropKey positions. Arrows indicate the data flow.

In MSDA, the feature map is divided into multiple heads along the channel dimension. Patches around the query are sampled using different dilation coefficients ri in different heads for self-attention computation. The results from different heads are concatenated and passed through a linear layer:

hi=MSDAQi,Ki,Vi,ri, 1≤i≤n (8)
X=LinearConcath1,…,hn (9)

In the above equations, ri denotes the dilation rate of the i-th head, and Qi, Ki, Vi denote the features of the i-th head. By defining a localized window around each query and controlling r, MSDA ensures the attention mechanism focuses primarily on information within relevant regions, drastically reducing computation. A higher dilation rate r means sparser sampling, covering a larger receptive field but potentially sacrificing some local detail. Setting different dilation rates in different heads allows each head to focus on features at different scales. The outputs of all heads are then merged via linear projection. By integrating DropKey, the MSDA module not only retains the advantage of multi-scale receptive field modeling but also further enhances the model’s ability to suppress noise, allowing it to focus more on discriminative regions in complex food images. This leads to improved classification accuracy and generalization performance.

3.4. Efficient Kolmogorov-Arnold Network

FFN struggles to meet lightweight requirements due to parameter redundancy in fully connected layers. EfficientKAN, based on the Kolmogorov–Arnold representation theorem, replaces fixed activation functions with learnable piecewise basis functions to reduce parameters while maintaining expressive power.

The Kolmogorov–Arnold theorem suggests that any multivariate continuous function can be represented by a finite composition of univariate functions, where each neuron realizes this decomposition through a learnable activation function. Specifically, the output of the l-th layer of the KAN is:

xjl+1=∑i=1nlϕi,jl,xil (10)

where ϕi,jl a learnable activation function (e.g., parameterized B-spline basis functions) and nl is the number of neurons in l-th layer.

The core improvement of EfficientKAN is to simplify the parameterization of the activation function and optimize computation. The original KAN [11] utilized high-complexity basis functions (e.g., B-splines):

ϕx=∑k=1KωkBkx (11)

where Bk(x) is the B-spline basis and K is the number of basis functions.

EfficientKAN switches to low-order polynomials or piecewise linear basis functions, for example:

ϕx=αx+β+∑m=1Mcm∗Relux−tm (12)

where α, β and cm are the learnable parameters; tm are learnable or fixed segmentation points. Furthermore, EfficientKAN employs parameter sharing and grouping: sharing some parameters (e.g., segmentation points tm) across activation functions of neighboring neurons and grouping neurons to share basis functions within a group while learning independently between groups, balancing expressiveness and efficiency. L1 regularization is applied to the basis function coefficients (e.g., ωk, cm) to induce sparsity, allowing coefficients close to zero to be pruned, further compressing the model.

The parameter reduction is significant: assuming n neurons per layer and each activation function having K basis functions, the original KAN requires O(n2K) parameters (K parameters per neuron pair connection). EfficientKAN reduces this to O(n2+nK) (shared basis functions + sparse coefficients). Consequently, forward propagation computation decreases from O(n2K) to O(n2+nK).

The standard MLP layer in the Transformer block is replaced by EfficientKAN [38,39]. The block computation is modified as:

x=LNx+MSAx (13)

Within the EfficientKAN layer, a rational activation function is utilized as the basis function instead of B-splines, i.e., the function ϕ (x) on each edge is parameterized as a rational function defined by polynomials P (x) and Q (x) of orders m and n:

ϕx=ωFx=ωPxQx=ωa0+a1x+⋯+amxmb0+b1x+⋯+bnbn (14)

In summary, EfficientKAN offers key advantages: (1) More efficient computation: Low-order basis functions and parameter sharing reduce FLOPs. (2) Fewer parameters: Sparsification and grouping strategies reduce the memory footprint. (3) Maintained expressiveness: Flexible piecewise functions approximate complex nonlinearities. These improvements significantly enhance the computational efficiency and memory usage of KAN, enabling its application to large-scale food image classification tasks.

4. Experimental Analysis

4.1. Experimental Setup

This study evaluates the proposed DKTransformer model on three publicly available food image datasets: ETH Food-101 [40], Vireo-Food-172 [41], and ISIA Food-500 [42]. ETH Food-101 is a large-scale Western food dataset comprising 101 categories with 101,000 images. Each category includes 750 training and 250 test images. Vireo-Food-172 is a Chinese cuisine dataset released in 2016, containing 172 categories and 110,241 images, with 80,241 training and 30,000 test images. ISIA Food-500 is an extended version of Food-101, containing 500 categories and 399,726 images, providing a more challenging benchmark for large-scale food recognition. As illustrated in Figure 7, these datasets exhibit diverse and complex food images. For instance, steak images often contain mixed ingredients and side dishes, while ice cream and cake categories feature layered compositions. In many cases, ingredients are scattered throughout the image (e.g., in Chinese dishes where main ingredients are interspersed with side dishes), requiring models to effectively distinguish local details and suppress background noise. To ensure a fair and consistent evaluation, we adopt uniform hyperparameter settings across all experiments. Input images are resized to 224 × 224 resolution. We use a batch size of 32, the AdamW optimizer with a weight decay of 0.005, and an initial learning rate of 0.001. Training runs for a maximum of 500 epochs. For data preprocessing, we apply standard normalization using ImageNet statistics and employ random cropping and horizontal flipping for data augmentation during training.

Figure 7.

Figure 7

Samples from ETH Food-101, Vireo-Food-172 and ISIA Food-500.

The evaluation metrics include Top-1 classification accuracy, precision, recall, F1, model parameter count (in millions), and floating-point operations (FLOPs) to assess computational efficiency.

4.2. Experimental Results and Analysis

4.2.1. ETH Food-101 Results Comparison

As shown in Table 1, DKTransformer achieves 92.71% Top-1 accuracy on the Food-101 dataset with 47 M parameters and 7.21 G FLOPs, demonstrating strong performance under a moderate computational budget. Compared with other Transformer-based baselines, DKTransformer surpasses ViT-B (90.30%) and Swin-S (90.91%) while requiring substantially fewer FLOPs than ViT-B (17.6 G) and fewer parameters than ViT-B (86 M). In addition, DKTransformer achieves better accuracy than ViT-S (89.47%) with lower FLOPs (7.21 G vs. 9.41 G). These results indicate that DKTransformer provides an effective accuracy–efficiency trade-off for fine-grained food image classification on Food-101.

Table 1.

Experimental results on Food-101.

Model Acc. (%) Pre. (%) Rec. (%) F1. (%) Parameters (M) FLOPS (G)
ViT-B 90.30 90.33 90.30 90.30 86 17.6
Swin-S 90.91 90.92 90.92 90.91 50 8.7
ViT-S 89.47 89.52 89.47 89.46 48.60 9.41
ConvNext-S [43] 89.41 89.43 89.41 89.41 56.6 8.69
ResNet-101 89.87 89.89 89.87 89. 87 44 7.9
Dilate-B 91.01 91.04 91.01 91.01 51 10.0
Ours 92.71 92.74 92.71 92.71 47 7.21

4.2.2. Vireo Food-172 and ISIA Food-500 Results

On the Vireo-Food-172 dataset (Table 2), DKTransformer achieves a Top-1 accuracy of 90.70% with 47 M parameters and 7.21 G FLOPs, outperforming ViT-B (89.30%) and Swin-S (89.46%) while maintaining lower computational complexity. In addition, DKTransformer exceeds ViT-S (87.63%) by a clear margin, indicating its effectiveness in handling fine-grained Chinese cuisine categories with complex ingredient compositions.

Table 2.

Experimental results on Vireo-Food-172.

Model Acc. (%) Pre. (%) Rec. (%) F1. (%) Parameters (M) FLOPS (G)
ViT-B 89.30 89.37 89.30 89.30 86 17.6
Swin-S 89.46 89.48 89.46 89.46 50 8.7
ViT-S 87.63 87.66 87.63 87.63 48.60 9.41
ConvNext-S 87.34 87.38 87.35 87.34 56.6 8.69
ResNet-101 87.24 87.31 87.24 89.24 44 7.9
Dilate-B 89.97 90.04 89.97 89.98 51 10.0
Ours 90.70 90.74 90.71 90.72 47 7.21

On the large-scale ISIA Food-500 dataset (Table 3), DKTransformer obtains a Top-1 accuracy of 66.89% with 47 M parameters and 7.21 G FLOPs. Compared with ViT-B (65.50%) and Swin-S (63.84%), DKTransformer achieves higher classification accuracy while significantly reducing model size and computational cost, demonstrating strong scalability and efficiency for large-category food recognition.

Table 3.

Experimental results on Food-500.

Model Acc. (%) Pre. (%) Rec. (%) F1. (%) Parameters (M) FLOPS (G)
ViT-B 65.50 65.58 65.50 65.50 86 17.6
Swin-S 63.84 63.89 63.84 63.84 50 8.7
ViT-S 62.53 62.55 62.53 62.53 48.60 9.41
ConvNext-S 60.12 60.14 60.12 60.12 56.6 8.69
ResNet-101 60.18 60.22 60.18 60.19 44 7.9
Dilate-B 64.45 64.49 64.45 64.45 51 10.0
Ours 66.89 66.91 66.89 66.89 47 7.21

4.3. Ablation Experiments

To validate the effectiveness of each proposed module, we conduct a series of ablation experiments using the ETH Food-101 dataset, which is a representative and widely adopted benchmark in fine-grained food recognition due to its scale, diversity, and realistic challenges. The results are summarized in Table 4.

(1) Under the premise of keeping the original ViT block unchanged, we introduce a downsampling strategy (+dawnSample + ViTBlock). The parameter count remains nearly the same, while FLOPs are significantly reduced to 9.30 G, and accuracy improves to 90.95%. This demonstrates that downsampling effectively reduces token redundancy and improves feature representation efficiency, providing a more efficient baseline for subsequent improvements.

(2) When replacing the ViT block with DKBlock without downsampling (+DKBlock), the accuracy increases to 91.40%, and parameters drop to 58.0 M, validating the parameter-efficient design of DKBlock. However, FLOPs remain high (15.20 G) due to the unchanged token number, indicating that block-level optimization alone is insufficient for reducing computation.

(3) With downsampling enabled, we further analyze the contributions of different DKBlock components. Adding LDE improves accuracy to 91.55%, but slightly increases parameters and FLOPs to 86.5 M and 9.55 G, suggesting enhanced local detail modeling with additional cost. In contrast, MSDA yields the most notable single-module gain, achieving 92.20% accuracy while reducing parameters and FLOPs to 60.0 M and 8.10 G, highlighting its effectiveness in modeling complex spatial dependencies. KAN mainly improves efficiency, reducing parameters and FLOPs to 52.0 M and 7.70 G, while maintaining 91.85% accuracy.

(4) For module combinations, MSDA + KAN achieves 92.60% accuracy with 47.0 M parameters and 7.30 G FLOPs. The full model delivers the best overall trade-off, reaching 92.71% accuracy with only 47.0 M parameters and 7.21 G FLOPs, demonstrating strong potential for practical deployment.

Table 4.

Results of ablation studies.

Model Acc. (%) Parameters (M) FLOPS (G)
ViT-B 90.30 86 17.6
+downSample + ViTBlock 90.95 82 9.30
+DKBlock 91.40 58 15.20
+LDE 91.55 86.5 9.55
+MSDA 92.20 60.0 8.10
+KAN 91.85 52.0 7.70
+LDE + MSDA 92.45 61.0 8.25
+LDE + KAN 92.10 53.0 7.85
+MSDA + KAN 92.60 47 7.30
Ours 92.71 47 7.21

4.4. Visualization and Analysis

To further analyze the feature localization capability of DKTransformer in fine-grained food image classification, we randomly selected and manually annotated a subset of images from the ETH Food-101 dataset to conduct a comparative study. This study combines quantitative localization metrics with attention heatmap visualizations. The quantitative evaluation results are reported in Table 5, while the qualitative attention responses are illustrated in Figure 8.

Table 5.

Results of quantitative results.

Model Mean IoU. (%) Loc Acc. (%)
ResNet-101 38.7 52.4
ConvNeXt-S 41.2 55.8
+DKBlock 40.9 54.6
Dilate-B 45.6 59.6
Swin-S 45.6 59.7
ViT-Base 42.5 52.1
Ours 64.2 73.5

Figure 8.

Figure 8

Comparison of attention heatmaps between ViT and DKTransformer. The first row shows the original images, the second row shows the results of ViT, and the third row shows the results of DKTransformer. The heatmaps visualize the model attention, where warmer colors indicate regions with higher importance for classification.

As shown in Table 5, DKTransformer achieves the best overall localization performance among the compared methods, reaching 64.2% Mean IoU and 73.5% localization accuracy. Overall, it consistently outperforms the other baseline approaches across the reported localization metrics, indicating stronger discriminative region awareness in fine-grained food recognition. Figure 8 further provides qualitative comparisons of attention heatmaps between the standard ViT and DKTransformer. For images with relatively clean backgrounds, both models can attend to the main food area, but DKTransformer produces more compact and concentrated responses. For more challenging cases involving complex backgrounds or multiple ingredients, the attention maps of ViT are often distracted by irrelevant regions, whereas DKTransformer suppresses background interference and highlights the core food components more consistently. These observations are consistent with the quantitative gains reported in Table 5, demonstrating improved localization robustness and discriminative feature learning.

4.5. Robustness Analysis Under Real-World Variations

To evaluate the robustness of DKTransformer under practical conditions, we first examine its performance consistency across datasets with distinct visual characteristics, and then conduct controlled perturbation experiments to quantify its resilience to common real-world degradations. ETH Food-101 mainly consists of Western dishes with relatively structured appearances, whereas Vireo-Food-172 contains Chinese cuisine images with mixed ingredients and more complex presentations. As reported in Table 1 and Table 2, DKTransformer achieves 92.71% accuracy on ETH Food-101 and 90.70% on Vireo-Food-172, demonstrating stable performance across different culinary styles. This consistency suggests that the MSDA module can effectively suppress background noise while preserving discriminative semantic cues.

To assess robustness under synthetic perturbations, we randomly selected a subset of images from the ETH Food-101 test set and applied varying degrees of occlusion and image corruptions to them. Figure 9a shows the accuracy degradation curves under increasing ratios of random block occlusion. When the occlusion ratio reaches 50%, the CNN-based ConvNeXt drops to 45.2%, likely due to its reliance on contiguous local patterns. In contrast, DKTransformer maintains 77.1% accuracy, outperforming the ViT-Base baseline by 5.6%, indicating that multi-scale dilated attention helps aggregate global context and supports reliable recognition under partial observations. Figure 9b further reports results under Gaussian noise and motion blur. DKTransformer exhibits stronger stability than ConvNeXt, with accuracy decreasing by only 3.7% under noise, compared with 9.1% for the baseline. These results confirm that the joint design of LDE and MSDA improves robustness against environmental variations commonly encountered in real-world food images.

Figure 9.

Figure 9

Robustness evaluation under synthetic perturbations. (a) Occlusion sensitivity analysis showing accuracy degradation under increasing ratios of random block occlusion. (b) Robustness comparison under common image corruptions, including Gaussian noise and Gaussian blur.

5. Conclusions and Future Work

This paper presents DKTransformer, a lightweight hybrid architecture that integrates ViT and CNNs for fine-grained food image classification. By jointly modeling local discriminative details and global contextual information, DKTransformer effectively addresses the challenges posed by high inter-class similarity and large intra-class variation in food image datasets.

The proposed method incorporates three key components. First, the LDE module, based on depthwise separable convolution, enhances the modeling of fine-grained local textures and structural details. Second, the MSDA module captures long-range dependencies with reduced computational complexity while suppressing background interference from non-food regions. Third, EfficientKAN replaces the conventional feedforward network to significantly reduce parameter redundancy while maintaining strong representational capacity.

Extensive experiments conducted on three public food image datasets—ETH Food-101, Vireo-Food-172, and ISIA Food-500—demonstrate the effectiveness of DKTransformer. The proposed model achieves competitive or superior classification accuracy compared with existing lightweight methods while requiring substantially fewer parameters and lower computational cost. These results indicate that DKTransformer provides a favorable balance between accuracy and efficiency, making it suitable for practical fine-grained food image classification applications.

Despite its promising performance, there are several directions for future work. First, online or incremental learning strategies could be explored to adapt the model to dynamically expanding food categories. Second, integrating multimodal information, such as ingredient lists, cooking instructions, or nutritional data, may further enhance recognition robustness and enable richer food-related applications. Finally, additional model compression and deployment-oriented optimizations, including quantization and hardware-aware acceleration, could be investigated to facilitate efficient inference on resource-constrained edge devices.

Acknowledgments

The authors would like to thank all contributors who provided valuable discussions and technical support during this study.

Author Contributions

Conceptualization, H.W. and X.A.; methodology, H.W.; software, C.W.; validation, H.W. and C.W.; formal analysis, C.W.; investigation, H.W.; writing—original draft preparation, C.W.; writing—review and editing, H.W. and X.A.; visualization, C.W.; supervision, X.A. All authors have read and agreed to the published version of the manuscript.

Data Availability Statement

The datasets used in this study are publicly available, including ETH Food-101, Vireo-Food-172, and ISIA Food-500.

Conflicts of Interest

The authors declare no conflicts of interest.

Funding Statement

This research received no external funding.

Footnotes

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

References

  • 1.Min W., Jiang S., Liu L., Rui Y., Jain R.C. A Survey on Food Computing. ACM Comput. Surv. 2019;52:1–36. doi: 10.1145/3329168. [DOI] [Google Scholar]
  • 2.Rostami A., Nagesh N., Rahmani A., Jain R.C. MADiMa’ 22: Proceedings of the 7th International Workshop on Multimedia Assisted Dietary Management. ACM; New York, NY, USA: 2022. World Food Atlas for Food Navigation; pp. 39–47. [Google Scholar]
  • 3.Ishino A., Yamakata Y., Karasawa H., Aizawa K. MM’ 21: The 29th ACM Multimedia Conference. ACM; New York, NY, USA: 2021. RecipeLog: Recipe Authoring App for Accurate Food Recording; pp. 2798–2800. [Google Scholar]
  • 4.Ródenas J., Nagarajan B., Bolaños M., Radeva P. MADiMa’ 22: Proceedings of the 7th International Workshop on Multimedia Assisted Dietary Management. ACM; New York, NY, USA: 2022. Learning multi subset of classes for fine-grained food recognition; pp. 17–26. [DOI] [Google Scholar]
  • 5.Zhu F., Bosch M., Schap T., Khanna N., Ebert D.S., Boushey C.J., Delp E.J. Segmentation Assisted Food Classification for Dietary Assessment. Proc. SPIE Int. Soc. Opt. Eng. 2011;7873:78730B. doi: 10.1117/12.877036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Kawano Y., Yanai K. FoodCam: A real-time food recognition system on a smartphone. Multimed. Tools Appl. 2015;74:5263–5287. doi: 10.1007/s11042-014-2000-8. [DOI] [Google Scholar]
  • 7.Tan R.Z., Chew X., Khaw K.W. Neural architecture search for lightweight neural network in food recognition. Mathematics. 2021;9:1245. doi: 10.3390/math9111245. [DOI] [Google Scholar]
  • 8.Dosovitskiy A., Beyer L., Kolesnikov A., Weissenborn D., Zhai X., Unterthiner T., Dehghani M., Minderer M., Heigold G., Gelly S., et al. An Image is Worth 16 × 16 Words: Transformers for Image Recognition at Scale. arXiv. 20202010.11929 [Google Scholar]
  • 9.Howard A.G., Zhu M., Chen B. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv. 2017 doi: 10.48550/arXiv.1704.04861.1704.04861 [DOI] [Google Scholar]
  • 10.Jiao J., Tang Y.M., Lin K.Y., Gao Y., Ma A.J., Wang Y., Zheng W.S. Dilateformer: Multi-scale dilated transformer for visual recognition. IEEE Trans. Multimed. 2023;25:8906–8919. doi: 10.1109/TMM.2023.3243616. [DOI] [Google Scholar]
  • 11.Liu Z., Wang Y., Vaidya S., Ruehle F., Halverson J., Soljačić M., Hou T.Y., Tegmark M. KAN: Kolmogorov-Arnold Networks. arXiv. 20242404.19756 [Google Scholar]
  • 12.Szegedy C., Liu W., Jia Y., Sermanet P., Reed S., Anguelov D., Erhan D., Vanhoucke V., Rabinovich A. 2015 IEEE Conference on Computer Vision and Pattern Recognition. POD Publisher; Seattle, WA, USA: 2015. Going deeper with convolutions; pp. 1–9. [Google Scholar]
  • 13.Ming Z.Y., Chen J., Cao Y., Forde C., Ngo C.W., Chua T.S. International Conference on Multimedia Modeling. Springer International Publishing; Cham, Switzerland: Food photo recognition for dietary tracking: System and experiment; pp. 129–141. [Google Scholar]
  • 14.McAllister P., Zheng H., Bond R., Moorhead A. Combining deep residual neural network features with supervised machine learning algorithms to classify diverse food image datasets. Comput. Biol. Med. 2018;95:217–233. doi: 10.1016/j.compbiomed.2018.02.008. [DOI] [PubMed] [Google Scholar]
  • 15.Ma N., Zhang X., Zheng H.T., Sun J. Computer Vision—ECCV 2018. Springer; Cham, Switzerland: 2018. Shufflenet v2: Practical guidelines for efficient cnn architecture design; pp. 116–131. [Google Scholar]
  • 16.Sandler M., Howard A., Zhu M., Zhmoginov A., Chen L.C. IEEE Conference on Computer Vision and Pattern Recognition. IEEE; New York, NY, USA: 2018. Mobilenetv2: Inverted residuals and linear bottlenecks; pp. 4510–4520. [Google Scholar]
  • 17.Howard A., Sandler M., Chu G., Chen L.C., Chen B., Tan M., Wang W., Zhu Y., Pang R., Vasudevan V., et al. 2019 IEEE/CVF International Conference on Computer Vision (ICCV) IEEE; New York, NY, USA: 2019. Searching for mobilenetv3; pp. 1314–1324. [Google Scholar]
  • 18.Liu Z., Lin Y., Cao Y., Hu H., Wei Y., Zhang Z., Lin S., Guo B. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) IEEE; New York, NY, USA: 2021. Swin transformer: Hierarchical vision transformer using shifted windows; pp. 10012–10022. [Google Scholar]
  • 19.Li Y., Yuan G., Wen Y., Hu J., Evangelidis G., Tulyakov S., Wang Y., Ren J. Efficientformer: Vision transformers at mobilenet speed. Adv. Neural Inf. Process. Syst. 2022;35:12934–12949. [Google Scholar]
  • 20.Zhu L., Wang X., Ke Z., Zhang W., Lau R.W. IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; New York, NY, USA: 2023. Biformer: Vision transformer with bi-level routing attention; pp. 10323–10333. [Google Scholar]
  • 21.Cai H., Li J., Hu M., Gan C., Han S. IEEE/CVF International Conference on Computer Vision. IEEE; New York, NY, USA: 2023. Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction; pp. 17302–17313. [Google Scholar]
  • 22.Yun S., Ro Y. IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; New York, NY, USA: 2024. Shvit: Single-head vision transformer with memory efficient macro design; pp. 5756–5767. [Google Scholar]
  • 23.Mehta S., Rastegari M. Separable Self-attention for Mobile Vision Transformers. arXiv. 2022 doi: 10.48550/arXiv.2206.02680.2206.02680 [DOI] [Google Scholar]
  • 24.Vasu P.K.A., Gabriel J., Zhu J., Tuzel O., Ranjan A. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR ‘23) IEEE; New York, NY, USA: 2023. MobileOne: An Improved One millisecond Mobile Backbone; pp. 7907–7917. [Google Scholar]
  • 25.Wang A., Chen H., Lin Z., Han J., Ding G. IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; New York, NY, USA: 2024. Repvit: Revisiting mobile cnn from vit perspective; pp. 15909–15920. [Google Scholar]
  • 26.Hatamizadeh A., Heinrich G., Yin H., Tao A., Alvarez J.M., Kautz J., Molchanov P. Fastervit: Fast vision transformers with hierarchical attention. arXiv. 20232306.06189 [Google Scholar]
  • 27.Sheng G., Sun S., Liu C., Yang Y. Food Recognition via an Efficient Neural Network with Transformer Grouping. Int. J. Intell. Syst. 2022;37:11465–11481. doi: 10.1002/int.23050. [DOI] [Google Scholar]
  • 28.Song J., Min W., Liu Y., Li Z., Jiang S., Rui Y. 2022 IEEE 5th International Conference on Multimedia Information Processing and Retrieval (MIPR) IEEE; New York, NY, USA: 2022. A Noise-robust Locality Transformer for Fine-grained Food Image Retrieval; pp. 348–353. [DOI] [Google Scholar]
  • 29.Liu Y., Min W., Jiang S., Rui Y. Convolution-Enhanced Bi-Branch Adaptive Transformer With Cross-Task Interaction for Food Category and Ingredient Recognition. IEEE Trans. Image Process. 2024;33:2572–2586. doi: 10.1109/TIP.2024.3374211. [DOI] [PubMed] [Google Scholar]
  • 30.Zhang Y., Liu H., Hu Q. International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer International Publishing; Cham, Switzerland: 2021. Transfuse: Fusing transformers and cnns for medical image segmentation; pp. 14–24. [Google Scholar]
  • 31.Chen X., Pan J., Lu J., Fan Z., Li H. AAAI Conference on Artificial Intelligence. Volume 37. AAAI; Washington, DC, USA: 2023. Hybrid cnn-transformer feature fusion for single image deraining; pp. 378–386. [Google Scholar]
  • 32.Sheng G., Min W., Zhu X., Xu L., Sun Q., Yang Y., Wang L., Jiang S. A Lightweight Hybrid Model with Location-Preserving ViT for Efficient Food Recognition. Nutrients. 2024;16:200. doi: 10.3390/nu16020200. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Chen Z., Wang J., Wang Y. Enhancing Food Image Recognition by Multi-Level Fusion and the Attention Mechanism. Foods. 2025;14:461. doi: 10.3390/foods14030461. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Negi S., Joshi A. Transformer Enabled ResNet Based Automated Skin Cancer Detection System. Biomed. Inform. Smart Healthc. 2025;1:149–154. doi: 10.62762/BISH.2025.444143. [DOI] [Google Scholar]
  • 35.Tewari Y., Parihar N.S., Rautela K., Kaundal N., Diwakar M., Pandey N.K. Diabetic Retinopathy Detection and Analysis with Convolutional Neural Networks and Vision Transformer. Biomed. Inform. Smart Healthc. 2025;1:18–26. doi: 10.62762/BISH.2025.724307. [DOI] [Google Scholar]
  • 36.Alam N., Zhu Y., Shao J., Usman M., Fayaz M. A Novel Deep Learning Framework for Brain Tumor Classification Using Improved Swin Transformer V2. ICCK Trans. Adv. Comput. Syst. 2025;1:154–163. doi: 10.62762/TACS.2025.807755. [DOI] [Google Scholar]
  • 37.Zheng Y., Jin X. AST-GNNFormer: Adaptive Spatio-Temporal Graph Neural Network with Layer-Aware Preservation for Traffic Flow Prediction. ICCK Trans. Emerg. Top. Artif. Intell. 2025;2:203–219. doi: 10.62762/TETAI.2025.387543. [DOI] [Google Scholar]
  • 38.Blealtan Efficient-Kan. 2024. [(accessed on 20 May 2024)]. Available online: https://github.com/Blealtan/efficient-kan.
  • 39.Yang X., Wang X. Kolmogorov-arnold transformer. arXiv. 20242409.10594 [Google Scholar]
  • 40.Bossard L., Guillaumin M., Van Gool L. European Conference on Computer Vision. Springer International Publishing; Cham, Switzerland: 2014. Food 101mining discriminative components with random forests; pp. 446–461. [Google Scholar]
  • 41.Chen J., Ngo C. MM’ 16: Proceedings of the 24th ACM International Conference on Multime. ACM; New York, NY, USA: 2016. Deep-based ingredient recognition for cooking recipe retrieval; pp. 32–41. [Google Scholar]
  • 42.Min W., Liu L., Wang Z., Luo Z., Wei X., Wei X., Jiang S. 28th ACM International Conference on Multimedia. ACM; New York, NY, USA: 2020. Isia food-500: A dataset for large-scale food recognition via stacked global-local attention network; pp. 393–401. [Google Scholar]
  • 43.Liu Z., Mao H., Wu C.Y., Feichtenhofer C., Darrell T., Xie S. IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; New York, NY, USA: 2022. A convnet for the 2020s; pp. 11976–11986. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets used in this study are publicly available, including ETH Food-101, Vireo-Food-172, and ISIA Food-500.


Articles from Sensors (Basel, Switzerland) are provided here courtesy of Multidisciplinary Digital Publishing Institute (MDPI)

RESOURCES