Abstract
Foundation models, including language models, e.g., GPT, and vision models, e.g., CLIP, have significantly advanced numerous biomedical tasks. Despite these advancements, the high inference latency and the “overthinking” issues in model inference impair the efficiency and effectiveness of foundation models, thus limiting their application in real-time clinical settings. To address these challenges, we proposed EPEE (Entropy- and Patience-based Early Exiting), a novel hybrid strategy designed to improve the inference efficiency of foundation models. The core idea was to leverage the strengths of entropy-based and patience-based early exiting methods to overcome their respective weaknesses. To evaluate EPEE, we conducted experiments on three core biomedical tasks-classification, relation extraction, and event extraction-using eight foundation models (i.e., BERT, ALBERT, GPT-2, ViT, Qwen, GPT-oss, BioMistral, and Meditron3) across twelve datasets, including clinical notes and medical images. The results showed that EPEE significantly reduced inference time while maintaining or improving accuracy, demonstrating its adaptability to diverse datasets and tasks. EPEE addressed critical barriers to deploying foundation models in healthcare by balancing efficiency and effectiveness. It potentially provided a practical solution for real-time clinical decision-making with foundation models, supporting reliable and efficient workflows.
Subject terms: Computational biology and bioinformatics, Health care, Mathematics and computing
Introduction
Foundation models, including language models, e.g., BERT1 and GPT series2, and vision models, e.g., Vision Transformers (ViTs)3 and CLIP4, have become increasingly popular in artificial intelligence, setting new benchmarks across diverse tasks5,6. Their impact is particularly significant in healthcare5,7, where language models excel in analyzing biomedical text and electronic health records (EHRs)8, including clinical note classification9,10, complex reasoning11, and information extraction12,13. Similarly, vision models have demonstrated exceptional performance in medical image analysis14,15, enabling tasks such as disease detection16, segmentation17, and classification18.
Despite these advancements, several challenges hinder the effective application of foundation models in healthcare. First, prediction accuracy is critical, as errors can pose risks to clinicians and patients. A major challenge is addressing “overthinking”19–22, where deeper model layers add unnecessary complexity without improving outcomes, wasting computational resources and potentially degrading performance23,24. Second, inference efficiency is paramount, especially in urgent settings like intensive care units (ICUs), where real-time, accurate decision-making is vital25. Delays in processing medical information can result in suboptimal treatment, prolonged hospital stays, or worse outcomes. Thus, models must provide rapid and reliable assessments to support timely clinical decisions25. However, as models grow larger, computational inefficiency and increased latency become significant barriers, impacting real-time applications and patient care workflows26–28.
While techniques such as network pruning29,30, knowledge distillation31,32, and weight quantization33,34 have been employed to improve inference efficiency, they fail to address the overthinking issue, as all layers are still used for predictions. This limits their effectiveness in biomedical contexts. A promising alternative is early exiting, a form of adaptive inference23,35–37, which introduces intermediate “exit” points within the model. This mechanism enables simpler cases to bypass deeper layers, significantly reducing inference time while maintaining or even improving performance. By dynamically adjusting computational depth based on input complexity, early exiting not only enhances efficiency but also mitigates overthinking. Furthermore, compared to other efficiency-enhancing techniques, early exiting offers greater flexibility, allowing adjustments to meet specific real-world performance requirements, such as latency or energy consumption. These attributes make early exiting an ideal approach to improve the inference efficiency of foundation models in healthcare.
Early exiting strategies are primarily categorized into entropy-based38 and patience-based methods23, both of which have limitations that can lead to suboptimal performance in biomedical tasks. Specifically, entropy-based methods are efficient and robust under high entropy thresholds but suffer from poor accuracy (Fig. 1), limiting their reliability in clinical applications. In contrast, patience-based methods, including PABEE23, PCEE27, and F-PABEE39, allow early exiting when classifiers produce the same prediction consecutively. While these methods generally achieve high accuracy (Fig. 1), their efficiency is highly sensitive to the hyper-parameter “patience”. Consequently, determining the optimal “patience” value to balance effectiveness and efficiency is challenging, particularly for specific clinical needs. To address these challenges, an ideal method for clinical use should be robust to variations in its hyper-parameter, enabling easy adjustment across diverse clinical tasks while ensuring reliable and efficient decision-making40,41.
Fig. 1. Performance comparison of entropy-based and patience-based methods on the Dietary Supplements Usage clas- sification task.
The speed-up ratio, as defined in Section 4.3, reflects reduced computational requirements and latency, with higher values indicating greater efficiency. The entropy-based method (green color) demonstrates notable efficiency and robustness at high thresholds but sacrifices accuracy. In contrast, the patience-based method (red color) lacks robustness to the varying hyper-parameter values (i.e., patience) yet generally achieves high accuracy.
Motivated by the potential to leverage the strengths of these approaches to overcome their respective weaknesses, we proposed EPEE (Entropy- and Patience-based Early Exiting), a novel method designed to enhance the inference efficiency of foundation models, as shown in Fig. 2. EPEE incorporated intermediate classifiers at each transformer block, allowing early exits when prediction entropy was sufficiently low or predictions remained consistent across a predefined number of layers. This hybrid approach offered greater flexibility for adjusting speed-up ratios by setting both entropy and patience thresholds and was adaptable to any foundation model. We conducted a comprehensive evaluation of EPEE on three core biomedical tasks-classification, relation extraction, and event extraction-using eight commonly employed foundation models: BERT1, ALBERT42, GPT-22, ViT3, Qwen2.543, GPT-oss44, BioMistral45, and Meditron346. Our experiments spanned eleven publicly available datasets, including clinical notes (e.g., MIMIC-ICU47), medical images (e.g., PneumoniaMNIST48,49 and PathMNIST49,50), and one private dataset of clinical notes51. Results demonstrated that EPEE significantly improved inference efficiency while maintaining the effectiveness of foundation models. Our contributions are summarized as follows:
We proposed EPEE, a novel hybrid early exiting method that enhanced the inference efficiency of foundation models in biomedical applications and was applicable to any foundation model.
To the best of our knowledge, EPEE is the first method specifically designed to address the “overthinking” issue in foundation model inference, ensuring both efficiency and effectiveness.
Extensive experiments verified the performance of EPEE across twelve biomedical datasets, demonstrating its ability to boost the efficiency and effectiveness of four foundation models on three critical tasks.
Fig. 2. Overview of the proposed EPEE method.
The method can be applied to both language and vision models. The entropy-based methods exit when the entropy criterion is satisfied, and the patience-based methods exit when it reaches the pre-set patience threshold. Our EPEE method uses both criteria for a more general and flexible early exiting strategy.
Results
In this section, we first demonstrated the efficiency and effectiveness of EPEE on language models under two operational modes: budgeted mode and dynamic mode. Next, we validated the effectiveness of our approach on visual foundation models. Finally, we showed the proposed EPEE method could be equivalent to the entropy-based method and the patience-based method respectively by changing the hyper-parameters, which demonstrated its superior generalization.
Budgeted Mode
After the LLMs and classifiers were trained jointly, we used each exit (i.e., classifiers after each transformer layer) to perform the classification, relation extraction, and event extraction tasks, and the accuracy and prediction entropy for each layer was shown in Fig. 3. It showed that the prediction entropy decreased with the layer increased, which meant the prediction was more and more confident as the transformer layer increased. In contrast, the classification accuracy presented an increasing trend as it delved deeper. The EPEE method worked well with BERT on the MIMIC-ICU dataset. The performance of the first layer was close to the highest layer which was the fourth layer. Most of the layers outperformed the last layer but the majority of current work uses the last layer directly. Similarly, for the private dietary supplement use dataset, the accuracy got above 0.8 only after two layers, and it achieved the highest accuracy at the fourth layer while the entropy reached the lowest at the last layer. In addition, our proposed method EPEE demonstrated similar results for PHEE, DDI, and GIT datasets as they achieved the highest accuracy before the last layer. The most exciting finding was that the first exit of models trained for Drug review, PHEE, DDI, and Medical health advice datasets exhibited high performance, especially for DDI and PHEE datasets; their first exit performance was close to the highest performance, which indicated huge efficient rewards from a little sacrifice in performance.
Fig. 3. Entropy and accuracy across each layer of BERT in budgeted mode.
The results indicate that the intermediate layer achieves performance that is comparable to or exceeds existing alternatives. Additionally, the decreasing entropy values suggest increased confidence in the predictions.Panel-dataset correspondence (to Table 1): a Intensive Care (MIMIC-III); b Dietary Supplement (Dietary Supplement Usage); c Drug–Drug Interaction (DDI); d Pharmacovigilance Event (PHEE); e Medical Health Advice (Medical Health Advice); f General Biomedical Health (GIT); g Drug Review (Drug Review).
To show the robustness of EPEE, hyper-parameters analysis was performed on ALBERT and GPT-2 models. They are both transformer-based models, but ALBERT is the encoder-based and GPT-2 is the decoder-based architecture. The results were shown in Fig. 5a–d and Fig. 6a–d separately. More experiment results on GPT-oss, Qwen2.5, Biomistral, and Meditron3 are in Section 2.5. We could observe similar results: 1) the best performance may not happen in the last layer; 2) the first few layers perform well with huge efficiency, which shows the overthinking issue is more severe in the GPT-2 model than in the BERT model.
Fig. 5. Hyper-parameter analysis with ALBERT on five biomedical text datasets.
Panels correspond to two evaluation modes. Budget mode reports the accuracy at each exit, while dynamic mode reports both the accuracy and the speed-up ratio. Specifically, panels a–e show the budget mode results (accuracy at each exit), panels f–j show the dynamic mode accuracy, and panels k–o show the dynamic mode speed-up ratio. Panels correspond to different biomedical datasets (listed in Table 1): a, f, k General Biomedical Health (GIT); b, g, l Intensive Care (MIMIC-III); c, h, m Dietary Supplement (Dietary Supplement Usage); d, i, n Pharmacovigilance Event (PHEE); e, j, o Drug–Drug Interaction (DDI). The results demonstrated the impact of key factors of EPEE on its efficiency and effectiveness.
Fig. 6. Hyper-parameter analysis with GPT-2 on five biomedical text datasets.
Panels correspond to two evaluation modes. Budget mode reports the accuracy at each exit, while dynamic mode reports both the accuracy and the speed-up ratio. Specifically, panels a–e show the budget mode results (accuracy at each exit), panels f–j show the dynamic mode accuracy, and panels k–o show the dynamic mode speed-up ratio. Panels correspond to different biomedical datasets (listed in Table 1): a, f, k Intensive Care (MIMIC-III); b, g, l Medical Health Advice (Medical Health Advice); c, h, m General Biomedical Health (GIT); d, i, n Drug–Drug Interaction (DDI); e, j, o Pharmacovigilance Event (PHEE). The results demonstrated the impact of key factors of EPEE on its efficiency and effectiveness.
Dynamic Mode
Unlike the budgeted mode, which sets a fixed computational depth for all inputs, the dynamic mode enables adaptive inference depth, allowing models to determine the optimal exit point for each input dynamically. Simpler cases exit earlier to minimize computational overhead, while more complex cases continue processing through deeper layers to ensure accuracy.
To comprehensively evaluate the flexibility of our method, we conducted a grid search over different entropy and patience settings. The results for the BERT model, presented in Fig. 4, illustrated the trade-off between accuracy and inference speed across varying configurations. Notably, for the dietary supplement usage classification dataset, the highest accuracy was achieved with a patience value of 2 or 3 across all entropy thresholds. This aligned with the budgeted mode findings, where the fourth layer yielded the highest classification accuracy.
Fig. 4. Dynamic mode performance of EPEE method compared with entropy-based and the patience-based methods.
For comparison with the entropy-based method and the patience-based method, their results were plotted as color green and red with one parameter considered to be maximum or minimum. Our method presented superior flexibility over the baselines. Panels are grouped by dataset, with the first panel in each pair showing accuracy and the second panel showing the speed-up ratio. The datasets are listed in Table 1: a–b Intensive Care (MIMIC-III); c–d Dietary Supplement (Dietary Supplement Usage); e–f Drug–Drug Interaction (DDI); g–h Pharmacovigilance Event (PHEE); i–j Medical Health Advice (Medical Health Advice); k–l General Biomedical Health (GIT); m–n Drug Review (Drug Review).
The broader experimental results across multiple datasets reveal two consistent trends: (1) Increasing the patience value generally improves accuracy, as it allows the model to confirm predictions over multiple layers. (2) Higher entropy thresholds and lower patience values result in greater inference speed-up, reducing computational cost but potentially impacting accuracy.
To further validate the robustness of the dynamic mode across different model architectures, we extended the experiments to ALBERT and GPT-2, as shown in Fig. 5e–l and 6e–l. The results confirmed that EPEE maintained its flexibility across transformer encoder and decoder architectures. Notably, for GPT-2, the first few exits already achieved near-peak accuracy, indicating that the overthinking issue was more pronounced in decoder-based architectures. Moreover, the dynamic mode effectively balanced speed and accuracy, demonstrating that even with different backbone models, the optimal exit layer varied based on input complexity. More experiment results on GPT-oss, Qwen2.5, Biomistral, and Meditron3 are in subsection??.
These findings highlight the adaptability of EPEE in dynamic mode, enabling fine-grained control over computational efficiency and predictive performance. The dynamic mode allows users to select the best entropy and patience combination to achieve high performance and low latency at the same time. Compared to existing early exiting strategies, EPEE provides a more flexible mechanism to balance speed and accuracy, making it particularly suitable for real-world biomedical applications where inference latency and decision reliability are critical.
EPEE for Biomedical Vision
To demonstrate the generalizability of our proposed EPEE method beyond language models, we extended our experiments to vision foundation models. Specifically, we evaluated EPEE on ViTs, which have emerged as a powerful architecture for medical image analysis.
In the budgeted mode, as shown in Fig. 7a–e, we assessed how the accuracy and entropy of predictions evolve across different transformer layers in ViTs. Similar to our findings in language models, we observed that predictions become more confident (i.e., entropy decreases) as layer depth increases. However, classification accuracy tends to stabilize after a few intermediate layers, suggesting that deeper layers may not always be necessary for optimal performance.
Fig. 7. Hyper-parameter analysis with ViT on five medical image datasets.
Panels correspond to two evaluation modes. Budget mode reports the accuracy at each exit, while dynamic mode reports both the accuracy and the speed-up ratio. Specifically, panels a–e show the budget mode results (accuracy at each exit), panels f–j show the dynamic mode accuracy, and panels k–o show the dynamic mode speed-up ratio. Panels correspond to different biomedical datasets (listed in Table 1): a, f, k Dermatoscope (DermaMNIST); b, g, l BreastUltrasound (BreastMNIST); c, h, m ChestX-Ray (PneumoniaMNIST); d, i, n BloodCellMicroscope (BloodMNIST); e, j, o ColonPathology (PathMNIST). The results demonstrated the impact of key factors of EPEE on its efficiency and effectiveness.
For most datasets, high classification accuracy was achieved at early exits, demonstrating that overthinking also exists in vision models. For example, in the PathMNIST dataset, intermediate layers produced comparable accuracy to the final layer, while significantly reducing computational cost. This finding reinforces the importance of early exiting in medical imaging, where efficiency is crucial for real-time diagnosis and clinical decision-making.
In addition, Fig. 7f–o illustrates how varying entropy and patience thresholds impact the accuracy and speed-up ratio across different datasets. The dynamic mode gives users a chance to improve performance and efficiency at the same time by refining the entropy and patience settings. For example on the BloodMNIST dataset, accuracy could reach 0.97 with the speed-up ratio of 0.7 by setting the entropy threshold to 0.1 and patience to 6.
Method Generalization and Special Cases
To illustrate that the EPEE method not only is flexible but also covers the function of the entropy-based method or the patience-based method, we examine its behavior in specific limiting cases. As shown in Fig. 8a, b, if we set the patience to M (i.e., the number of transformer layers of PLM), then the patience counter would only possibly be satisfied if it reaches to the last layer. However, when it reaches the last layer, it predicts the output regardless of whether any criteria are satisfied. In this setting, the patience criterion loses control of decisions and thus our EPEE method reduces to the entropy-based method. On the other hand, as shown in Fig. 8c, d, if the entropy threshold is set to 0, the entropy criterion is prohibited because the entropy criterion would never be satisfied. Then the patience criterion would be the only way to early exit. Therefore, the EPEE method is more general, and the two previous methods are the special cases of the EPEE method.
Fig. 8. Degeneracy analysis of EPEE.
Our method could be simplified into the entropy-based or the patience-based methods by invalidating a parameter. a entropy-based method; b our method with patience = 12; c patience-based method; d our method with entropy threshold = 0.
Additional Foundation Model Benchmarks
To reveal the generalization of our method via benchmarking more recent foundation models and comparing general-domain versus biomedical models, we additionally evaluated EPEE on four modern open-source LLMs: GPT-oss-20B and Qwen2.5-7B (general-domain), as well as BioMistral-7B and Meditron3-8B (biomedical-domain). Figs. 9–12 report the hyperparameter analysis (entropy threshold vs. patience) with these four new foundational models on three representative biomedical text datasets. Across all four backbones, EPEE remains effective and consistently provides favorable accuracy-efficiency trade-offs, indicating that the observed benefits generalize to newer and larger foundation models.
Fig. 10. Ablation study on Qwen model.
Hyper-parameter analysis with Qwen2.5 on three biomedical text datasets. Panels correspond to two evaluation modes. Budget mode reports the accuracy at each exit, while dynamic mode reports both the accuracy and the speed-up ratio. Specifically, panels a–c show the budget mode results (accuracy at each exit), panels d–f show the dynamic mode accuracy, and panels g–i show the dynamic mode speed-up ratio. Panels correspond to different biomedical datasets (listed in Table 1): a, d, g Pharmacovigilance Event (PHEE); b, e, h Medical Health Advice (Medical Health Advice); d, f, i Drug–Drug Interaction (DDI).
Fig. 11. Ablation study on BioMistral model.
Hyper-parameter analysis with BioMistral on three biomedical text datasets. Panels correspond to two evaluation modes. Budget mode reports the accuracy at each exit, while dynamic mode reports both the accuracy and the speed-up ratio. Specifically, panels a–c show the budget mode results (accuracy at each exit), panels d–f show the dynamic mode accuracy, and panels g–i show the dynamic mode speed-up ratio. Panels correspond to different biomedical datasets (listed in Table 1): a, d, g Pharmacovigilance Event (PHEE); b, e, h Medical Health Advice (Medical Health Advice); d, f, i Drug–Drug Interaction (DDI).
Fig. 9. Ablation study on GPT-oss model.
Hyper-parameter analysis with GPT-oss-20B on three biomedical text datasets. Panels correspond to two evaluation modes. Budget mode reports the accuracy at each exit, while dynamic mode reports both the accuracy and the speed-up ratio. Specifically, panels a–c show the budget mode results (accuracy at each exit), panels d–f show the dynamic mode accuracy, and panels g–i show the dynamic mode speed-up ratio. Panels correspond to different biomedical datasets (listed in Table 1): a, d, g Pharmacovigilance Event (PHEE); b, e, h Medical Health Advice (Medical Health Advice); d, f, i Drug–Drug Interaction (DDI).
Fig. 12. Ablation study on Meditron model.
Hyper-parameter analysis with Meditron3 on three biomedical text datasets. Panels correspond to two evaluation modes. Budget mode reports the accuracy at each exit, while dynamic mode reports both the accuracy and the speed-up ratio. Specifically, panels a–c show the budget mode results (accuracy at each exit), panels d–f show the dynamic mode accuracy, and panels g–i show the dynamic mode speed-up ratio. Panels correspond to different biomedical datasets (listed in Table 1): a, d, g Pharmacovigilance Event (PHEE); b, e, h Medical Health Advice (Medical Health Advice); d, f, i Drug–Drug Interaction (DDI).
Statically analysis
Finally, we statistically evaluate the inference efficiency of our method (EPEE) against three baselines: standard inference, an entropy-based early-exit strategy, and a patience-based early-exit strategy. As shown in Fig. 13, we report normalized inference time by setting the standard inference time to 1.0 for each task and scaling the remaining methods accordingly. All results are obtained with a BERT-based model; the bars denote the mean inference time over repeated runs, and the error bars indicate the standard deviation. Overall, EPEE consistently achieves lower (or at least comparable) inference time than the competing early-exit baselines across all evaluated tasks, and it yields the largest speedup on the Intensive Care task. These results demonstrate that EPEE provides a more efficient inference procedure while maintaining stable performance under statistical variation.
Fig. 13. Normalized inference time (mean ± std) of a BERT-based model across six biomedical tasks.
Standard inference is normalized to 1.0 per task. We compare standard inference, an entropy-based method, a patience-based method, and our approach (EPEE). Error bars represent the standard deviation over repeated runs.
Training Budget and Sensitivity Analysis
Although our method primarily improves inference efficiency, the training cost is also an important criterion for evaluating the overall practicality of the approach. In Fig. 14a, we report the training time, parameter size, and test-time inference latency of a BERT model and a GPT-oss-20B model across different datasets. The results show that the early-exit classifier introduces only a negligible overhead (< 3% of total parameters). Moreover, when comparing training cost with inference-time latency, our method yields substantial efficiency gains during inference while adding minimal training burden.
Fig. 14. Training overhead and sensitivity analysis of EPEE.
a Training-time cost, parameter overhead, and test-time per-instance inference latency (batch size = 1) for backbone models equipped with early-exit heads, comparing BERT and GPT-oss-20B across representative biomedical datasets. The additional early-exit heads introduce only a negligible increase in parameters relative to the backbone. b, c Sensitivity analysis of EPEE on b Dietary Supplement Usage and c DDI, reporting accuracy and speed-up ratio under different combinations of entropy threshold and patience.
We further present two sets of sensitivity experiments using the BERT model on the Dietary Supplement and DDI datasets, as shown in Fig. 14b, c, where we vary the entropy threshold and patience threshold. Across a wide range of settings, the performance remains highly stable: the accuracy does not exhibit any noticeable degradation under any configuration. These observations indicate that our method is robust to hyperparameter choices.
Discussion
In this study, we introduced the EPEE method, specifically designed to address accuracy and latency challenges in the biomedical and healthcare domains. We evaluated EPEE across three tasks-classification, relation extraction, and event extraction-using twelve diverse datasets: MIMIC-ICU, dietary supplement usage, drug review, PHEE, DDI, GIT, and medical health advice, pathMNIST, PneumoniaMNIST, BloodMNIST, DermaMNIST, BreastMNIST. The method was evaluated against two existing approaches, entropy-based and patience-based early exiting, and was implemented in three main structures: transformer encoder, transformer decoder, and vision transformer via eight pre-trained models: BERT, ALBERT, GPT-2, ViT, GPT-oss, Qwen2.5, BioMistral, and Meditron3 under two operational modes: budgeted and dynamic. Our findings highlight both the necessity of early exiting in this domain and the advantages offered by our EPEE method in improving both performance and computational efficiency.
The budgeted mode results clearly emphasize the importance of early exiting in the biomedical domain. Traditional inference approaches, which rely on the classifier in the final layer of the model, often lead to the “overthinking” issue-where additional processing in deeper layers does not contribute to better predictions and may even degrade performance. This phenomenon was evident in all models we used- BERT, ALBERT, GPT-2, ViT (and GPT-oss, Qwen2.5, BioMistral, and Meditron3 in Supplementary Information), where exiting earlier at intermediate layers resulted in better performance compared to utilizing the final layer. The severity of this overthinking issue is expected to increase with larger models, as all the first exits in the GPT-2 model achieved high accuracy, making early exiting an essential strategy for future model deployments.
Beyond improving performance, early exiting also delivers significant efficiency gains in inference time. For example, in the budgeted mode, we observed that the fourth exit of the BERT model trained on the dietary supplement usage dataset achieved the highest F1 score while consuming only one-third of the inference time required by the full model. This finding underscores the practical value of early exiting, particularly in scenarios such as high-traffic healthcare environments, where both performance and response times are critical.
Interestingly, our experiments revealed consistent patterns across datasets. For each dataset, there is typically a specific transformer layer or a small range of transitional layers (2-3 layers) after which the performance stabilizes and closely approximates that of the final layer. For instance on the BERT model, in the Drug Review, PHEE, GIT, and Medical Health Advice datasets, models achieved near-optimal performance around the 9th transformer layer. These layers may have already captured most of the key features, making further computations in deeper layers potentially redundant. This trend is consistently observed across different tasks and datasets, indicating that the computational demand of the model does not grow linearly but instead exhibits diminishing returns beyond a certain depth, which illustrates the importance of the early exiting method.
In the dynamic mode, our proposed EPEE method demonstrated exceptional flexibility and adaptability compared to the entropy-based and patience-based methods. While the latter two approaches are governed by a single parameter, EPEE leverages a combination of entropy and patience thresholds to provide a finer level of control over the speed-up ratio. As illustrated in Figs. 4 and 5, the EPEE method allows users to dynamically adjust the trade-off between accuracy and inference speed by tuning these thresholds. This capability is particularly valuable in real-world applications, where different tasks or operational environments may demand varying levels of precision and computational efficiency.
A key advantage of EPEE in dynamic mode is that it enables the discovery of optimal configurations where models achieve both high accuracy and significant speed-up, striking a balance that neither entropy-based nor patience-based methods can achieve individually. Specifically, our experiments reveal that by carefully selecting the entropy and patience parameters, EPEE consistently identifies configurations where inference is accelerated while maintaining performance comparable to full-depth execution. On multiple datasets, we observe speed-up gains and performance gains could be achieved at the same time.
In addition, through grid-search optimization of entropy and patience thresholds, the EPEE method can achieve any desired speed-up ratio to satisfy user needs. This flexibility enables practitioners to explore a wider range of trade-off configurations, allowing for better alignment with specific application requirements. Notably, our method addresses a key limitation of entropy-based methods, where the speed-up ratio often exhibits abrupt changes over a narrow parameter range, making fine-tuning challenging. By pairing specific entropy values with corresponding patience values, EPEE provides a smoother, more controlled adjustment, ensuring both efficiency and reliability in the model’s performance.
The experiments on medical computer vision reveal the robustness of the EPEE method for different architectures, including transformer encoder, decoder, and vision transformer, which guarantees that the EPEE method is compatible with all foundation models. In the medical computer vision domain, overthinking and latency issues also exist, as shown in Fig. 7. The ViT can always provide higher or comparable accuracy with a better speed-up ratio, which lays the ground for efficient and accurate vision foundation models in biomedical and healthcare domains.
The final notable strength of the EPEE method lies in its generality and inclusivity. By appropriately setting one of its parameters to an extreme value, the EPEE framework can effectively reduce to either the entropy-based or patience-based methods, demonstrating its generalization on covering main existing approaches. This capability highlights the versatility of EPEE, as it not only enhances flexibility but also integrates the strengths of prior methods under a unified framework.
Despite the encouraging results, this study has several limitations. First, EPEE requires tuning two hyper-parameters (entropy threshold and patience), and while this improves controllability, it may introduce additional calibration and selection effort when deploying to new tasks or operating points. In addition, our efficiency evaluation focuses on per-instance inference with a batch size of 1 and reports wall-clock time and speed-up ratios under a specific hardware/software stack; latency characteristics can differ across devices, batching regimes, and system-level optimizations.
Overall, the findings of this study underscored the critical role of early exiting strategies in optimizing the performance and efficiency of foundation models in the biomedical and healthcare domains. The results suggested that incorporating the advanced EPEE method could address the overthinking issue inherent in deep transformer architectures while enabling models to operate effectively within the constraints of real-world applications. Future work could explore extending the EPEE framework to even larger and more diverse datasets, as well as investigating its applicability to other domains beyond healthcare. Additionally, integrating EPEE with emerging multi-modal foundation models that process both text and image inputs simultaneously, could further reveal its capabilities and widen its scope of impact.
Methods
Preliminaries
In this subsection, we present the essential background and define the mathematical notations for the early exiting method. This study focuses on a multi-class classification setting, where the dataset is denoted as (X, Y), with individual samples represented by (xi, yi) for i = 1, 2, …, N. Here, xi ∈ X represents the input sentence, and yi ∈ Y corresponds to its associated label. The classification task involves a class space denoted by K. We define M as the total number of Transformer layers, d as the hidden layer dimension, and sm as the hidden state obtained after the m-th layer, where m ∈ {1, 2, …, M}.
Entropy-based Early-Exiting Method
As illustrated in Fig. 2, early exiting architectures incorporate exit points at each Transformer layer. For a model with M Transformer layers, M classifiers fm(sm; θm): sm → K (m = 1, 2, …, M) are designated at these layers. Each classifier maps the hidden state sm of its respective layer to a probability distribution pm(sm; θm) over ∣K∣ classes using the softmax function. The confidence level of each layer m is quantified using the entropy of the predicted class distribution pm. Normalized entropy, which serves as a measure of confidence, is calculated as follows:
| 1 |
where represents the probability assigned to the k-th class by the m-th Transformer layer. A lower entropy value Hm signifies higher confidence in the prediction. If Hm is less than a predefined threshold τ, the prediction at layer m is considered confident. Otherwise, the process proceeds to the next Transformer layer. If none of the classifiers generate a confident prediction, the final classifier at the last layer outputs the prediction, irrespective of its confidence level.
Patience-based Early-Exiting Method
As illustrated in Fig. 2, a patience counter P is maintained to track how many consecutive classifiers predict the same class. If two consecutive predictions differ, the patience counter is reset to 1. At each layer m, the patience counter is updated as follows:
| 2 |
When Pm reaches a predefined threshold Pt (the patience parameter), the model exits early at layer m. If this condition is never satisfied, the final classifier at the last layer M produces the prediction. This approach allows the model to exit early when multiple classifiers consistently predict the same result, ensuring high confidence in the prediction.
EPEE Method
The entropy-based acceleration method has garnered significant attention due to its simplicity, efficiency, and flexibility. However, it suffers from reduced accuracy when the entropy threshold is increased to accelerate inference. Conversely, the patience-based acceleration method achieves state-of-the-art performance but encounters inefficiencies: it tends to over-process simple inputs when the patience parameter is set too high, while under-processing certain inputs when the parameter is set too low. Furthermore, both methods rely on a single parameter to control the speed-up ratio, which can be inconvenient when aiming for a specific budgeted speed-up ratio. Consequently, there is a pressing need for a novel method to address these limitations.
To bridge this gap, we propose a novel early exiting method, termed EPEE (Entropy- and Patience-based Early Exiting), as illustrated in Fig. 2. Similar to the existing approaches, our method evaluates whether to exit at each attention layer during inference. However, EPEE simultaneously employs two exiting criteria, thereby inheriting the strengths of both methods.
For an input sentence x, at the m-th layer, the model exits directly if the entropy score is below a predefined threshold or if the patience counter reaches a predefined value. Otherwise, the patience counter is updated: it increments by 1 if the current layer’s prediction matches that of the previous layer; otherwise, it resets to 1. If no stopping criterion is met, the classifier at the final transformer layer generates the prediction. The mechanism is summarized as:
| 3 |
The primary advantage of the proposed EPEE method lies in its flexibility. While the entropy-based and patience-based methods rely solely on entropy and patience counters, respectively, to determine the exit, EPEE leverages both, enabling a more versatile control over the speed-up ratio. Moreover, EPEE encompasses both existing methods as special cases: setting the entropy threshold to 0 reduces EPEE to the patience-based method, while defining the patience parameter as M (i.e., all layers) reduces EPEE to the entropy-based method.
Additionally, EPEE resolves the limitations of the individual methods by effectively combining their strengths. The entropy-based approach efficiently handles simple input sentences with high confidence but may falter with complex inputs. In contrast, the patience-based method can waste computational resources on simple inputs due to its fixed patience threshold. By adopting a small entropy threshold, EPEE ensures that simple sentences exit quickly with high confidence, while more complex inputs utilize the patience counter to exit appropriately. This ensures that complex sentences are processed efficiently without waiting until the final layer, while simple inputs are resolved expeditiously.
Study Design
Datasets
We evaluated EPEE against alternative methods using seven biomedical text datasets across three core tasks: Classification (3 datasets), Relation Extraction (2 datasets), and Event Extraction (1 dataset). Additionally, we conducted experiments on five medical image classification datasets.
Among the selected datasets, the Dietary Supplements Usage Status dataset10,51,52 is a private dataset developed using clinical notes from the University of Minnesota. This dataset specifically targets mentions of dietary supplements, comprising a total of 3,000 annotated sentences categorized into four use status classes: Continuing (C), Discontinued (D), Started (S), and Uncertain (U). The dataset captures mentions of dietary supplements frequently used by patients in clinical settings. The 25 dietary supplements included in the dataset are: Alfalfa, Biotin, Black Cohosh, Coenzyme Q10, Cranberry, Dandelion, Echinacea, Fish Oil, Flax Seed, Folic Acid, Garlic, Ginger, Ginkgo, Ginseng, Glucosamine, Glutamine, Kava Kava, Lecithin, Melatonin, Milk Thistle, Saw Palmetto, St. John’s Wort, Turmeric, Valerian, Vitamin E. All other datasets used in this study are publicly available. A summary of the datasets is presented in Table 1.
Table 1.
Overview of the data statistics
| Dataset | Data Status | Data Source | Data Type | Task | Train | Dev | Test | Classes |
|---|---|---|---|---|---|---|---|---|
| MIMIC-III47 | Public | Beth Israel Deaconess Medical Center | Intensive Care Unit Record | Classification | 3861 | 483 | 483 | 4 |
| Dietary Supplement Usage51 | Private | University of Minnesota | Electrical Health Record | Classification | 2000 | 230 | 230 | 4 |
| Drug Review57 | Public | Patient Reviews | Patient-generated Text | Classification | 161,297 | 53,766 | 53,766 | 2 |
| Medical Health Advice58 | Public | PubMed | Medical Literature | Classification | 6940 | 868 | 868 | 3 |
| DDI59 | Public | Medline Abstract | Medical Literature | Relation Extraction | 11,556 | 1285 | 3020 | 5 |
| GIT60 | Public | PubMed Abstract | Medical Literature | Relation Extraction | 3734 | 465 | 492 | 22 |
| PHEE61 | Public | Literature and Reports | Medical Literature | Event Extraction | 2898 | 961 | 968 | 2 |
| PathMNIST49,50 | Public | NCT Biobank and the UMM Pathology Archive | Colon Pathology | Classification | 89,996 | 10,004 | 7180 | 9 |
| PneumoniaMNIST48,49 | Public | Children Hospital | Chest X-Ray | Classification | 4708 | 524 | 624 | 2 |
| BreastMNIST49,62 | Public | Baheya Hospital | Breast Ultrasound | Classification | 546 | 78 | 156 | 2 |
| BloodMNIST49,63 | Public | Hospital Clinic of Barcelona | Blood Cell Microscope | Classification | 11,959 | 1712 | 3421 | 8 |
| DermaMNIST49,64,65 | Public | Medical University of Vienna and Cliff Rosendahl | Dermatoscope | Classification | 7007 | 1003 | 2005 | 7 |
Backbone Models
Considering the efficiency and widespread impact, this study mainly used BERT1, ALBERT42, GPT-22, ViT3, Qwen2.543, GPT-oss44, BioMistral45, and Meditron346 as the backbone models. As shown in Table 2, they consist of three different structures: Transformer encoder53, Transformer decoder53, and Vision transformer54 .
Table 2.
Overview of the foundational models
| Model | Parameter | Backbone | Pre-trained Data |
|---|---|---|---|
| BERT | 109M | Transformer encoder | Wikipedia + BooksCorpus |
| ALBERT | 12M | Transformer encoder | Wikipedia + BooksCorpus |
| GPT-2 | 124M | Transformer decoder | WebText |
| ViT | 86M | Patch embedding + Transformer encoder | JFT-300M (pre-train), ImageNet (fine-tune) |
| Qwen2.5-7B | 7.61B | Decoder-only Transformer | Large-scale corpus (up to 18T tokens) |
| GPT-oss-20b | 20B | MoE Transformer | general knowledge |
| Meditron3-8B | 8B | Decoder-only Transformer | Guidelines, publications, synthetic DDx, replay, MCQ |
| BioMistral-7B | 7B | Decoder-only Transformer | PubMed Central Open Access subset |
Training
During training, all exiting classifiers are jointly optimized using a weighted cross-entropy loss function. Following previous works23,27,55, the loss function is formulated as a weighted average of the cross-entropy losses, given by:
| 4 |
Here, denotes the cross-entropy loss for the m-th exit, where y represents the ground truth, pm is the predicted probability distribution at the m-th exit, and N is the number of classes. The weight wm corresponds to the relative inference cost associated with the m-th exit.
Speed-up Ratio
The efficiency of the early exiting method is quantified using the speed-up ratio23,27, which is computed as follows. Let M denote the total number of layers in the backbone model, and for each test sample xi (where i = 1, 2, …, N), let indicate whether the m-th transformer layer is utilized during inference for input xi. The average speed-up ratio over the test set is defined as:
| 5 |
This metric is chosen for its linear relationship with actual computational cost. Based on our experiments, it demonstrates a strong correlation with wall-clock runtime while maintaining stability across runs, even in the presence of potential randomness introduced by other processes on the same machine. Additionally, the FLOPS-based computational analysis is in supplementary material 1.
Inference Modes
During inference, the model with multiple exits can employ two early exiting strategies based on whether the computational budget is predefined.
Budgeted exiting: if the computational budget is set, a specific exit m* can be chosen, where m*-th exiting classifier is used to predict all queries.
Dynamic exiting: in this mode, after receiving an input x, the model sequentially predicts using classifiers from beginning to end, reusing computations where feasible. The process continues until an exit criterion is met at layer m* < M or the final exit M is reached. The final prediction combines the current and previous predictions, allowing different samples to exit at varying layers.
Experimental Settings
The foundational models and classifier at every layer were trained once and we reloaded the coefficients for different early exiting strategies. During training, a grid search was conducted to determine optimal hyper-parameters, with batch sizes set at 16, 32, and 128, and learning rates tested at 1e-5, 2e-5, 3e-5, and 5e-5 using the Adam optimizer for 15 epochs. All implementations were built on Hugging Face’s Transformers library 56, and experiments were conducted on a single Nvidia A100 GPU with 40GB memory. For inference, we adopted a per-instance approach, setting the batch size to 1 to simulate real-world usage, where individual requests may come from different users at varying times.
Supplementary information
Acknowledgements
This work was supported by the National Institutes of Health’s National Center for Complementary and Integrative Healthunder grant number R01AT009457 and U01AT012871, the National Institute on Aging under grant number R01AG078154, the National Cancer Institute under grant number R01CA287413, the National Institute of Diabetes and Digestive and Kidney Diseases under grant number R01DK115629, the National Institute on Minority Health and Health Disparities under grant number 1R21MD019134, and the Food and Drug Administration under grant number U01FD008720. Many thanks to Yuqian Chen for her help with the drawing.
Author contributions
Z.Z. and R.Z. contributed to the concept and design of the study. Z.Z. and S.Z. performed data collection. Z.Z. implementedthe code and conducted the experiments. Z.Z. and S.Z. drafted the manuscript. R.Z. supervised the study. Z.Z., S.Z., H.Z., Z.L., and R.Z. contributed to the research discussion and to the review and revision of the manuscript. Z.Z., S.Z., H.Z., Z.L., and R.Z. read and approved the final manuscript.
Data Availability
Eleven datasets involved in this study are publicly available from the following links:Drug Review: https://archive.ics.uci.edu/dataset/462/drug+review+dataset+drugs+comMedical Health Advice: https://huggingface.co/datasets/medalpaca/medical_meadow_health_adviceDDI: https://github.com/isegura/DDICorpusGIT: https://github.com/ToneLi/BIoMedRAG/tree/main/dataset/0_GM-CIHTPHEE: https://github.com/zhaoyuesun/pheePathMNIST, PneumoniaMNIST, DermaMNIST, BloodMNIST and BreastMNIST: https://medmnist.com/.
Code availability
The code is publicly available on Github: https://github.com/Learner4everrr/EPEE.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary information
The online version contains supplementary material available at 10.1038/s44401-026-00083-2.
References
- 1.Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Burstein, J., Doran, C. & Solorio, T. (eds.) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186 (Association for Computational Linguistics, 2019). https://aclanthology.org/N19-1423.
- 2.Radford, A. et al. Language models are unsupervised multitask learners. Meta AI (2019).
- 3.Dosovitskiy, A. et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR, 2021). https://openreview.net/forum?id=YicbFdNTTy.
- 4.Radford, A. et al. Learning transferable visual models from natural language supervision. In Meila, M. & Zhang, T. (eds.) Proceedings of the 38th International Conference on Machine Learning, vol. 139 of Proceedings of Machine Learning Research, 8748–8763 (PMLR, 2021). https://proceedings.mlr.press/v139/radford21a.html.
- 5.Azad, B. et al. Foundational models in medical imaging: a comprehensive survey and future vision. arXiv preprint arXiv:2310.18689 (2023).
- 6.Lu, M. Y. et al. A visual-language foundation model for computational pathology. Nat. Med.30, 863–874 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Zhou, S. et al. Large language models for disease diagnosis: a scoping review. npj Artif. Intell.1, 9 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Shickel, B., Tighe, P. J., Bihorac, A. & Rashidi, P. Deep EHR: a survey of recent advances in deep learning techniques for electronic health record (ehr) analysis. IEEE J. Biomed. Health Inform.22, 1589–1604 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Zhan, Z. & Zhang, R. Towards better multi-task learning: A framework for optimizing dataset combinations in large language models. In Findings of the Association for Computational Linguistics: NAACL 2025, 5373–5386 (ACL, 2025).
- 10.Zhan, Z., Zhou, S., Li, M. & Zhang, R. Ramie: retrieval-augmented multi-task information extraction with large language models on dietary supplements. Journal of the American Medical Informatics Association ocaf002 10.1093/jamia/ocaf002. https://academic.oup.com/jamia/advance-article-pdf/doi/10.1093/jamia/ocaf002/61415205/ocaf002.pdf (2025). [DOI] [PMC free article] [PubMed]
- 11.Zhou, S. et al. Explainable differential diagnosis with dual-inference large language models. npj Health Syst.2, 12 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Zhan, Z., Wang, J., Zhou, S., Deng, J. & Zhang, R. Mmrag: multi-mode retrieval-augmented generation with large language models for biomedical in-context learning. J. Am. Med. Inform. Assoc.32, 1505–1516 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Yue, X. & Zhou, S. PHICON: Improving generalization of clinical text de-identification models via data augmentation. In Proceedings of the 3rd Clinical Natural Language Processing Workshop, 209–214 (Association for Computational Linguistics, 2020).
- 14.Chen, X. et al. Recent advances and clinical applications of deep learning in medical image analysis. Med. Image Anal.79, 102444 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.He, K. et al. Transformers in medical image analysis. Intell. Med.3, 59–78 (2023). [Google Scholar]
- 16.Zhou, Y. et al. A foundation model for generalizable disease detection from retinal images. Nature622, 156–163 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Asgari Taghanaki, S., Abhishek, K., Cohen, J. P., Cohen-Adad, J. & Hamarneh, G. Deep semantic segmentation of natural and medical images: a review. Artif. Intell. Rev.54, 137–178 (2021). [Google Scholar]
- 18.Zhang, K. et al. A generalist vision-language foundation model for diverse biomedical tasks.Nat. Med30, 3129–3141 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Chen, S. et al. Don’t shoot butterfly with rifles: Multi-channel continuous speech separation with early exit transformer. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6139–6143 (IEEE, 2021).
- 20.Bajpai, D. J. & Hanawal, M. K. A survey of early exit deep neural networks in NLP. arXiv preprint arXiv:2501.07670 (2025).
- 21.Xie, K., Lu, S., Wang, M. & Wang, Z. Elbert: Fast albert with confidence-window based early exit. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 7713–7717 (IEEE, 2021).
- 22.Gao, X., Liu, Y., Huang, T. & Hou, Z. Pf-berxit: Early exiting for bert with parameter-efficient fine-tuning and flexible early exiting strategy. Neurocomputing558, 126690 (2023). [Google Scholar]
- 23.Zhou, W. et al. Bert loses patience: Fast and robust inference with early exit. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M. & Lin, H. (eds.) Advances in Neural Information Processing Systems, vol. 33, 18330–18341 (Curran Associates, Inc., 2020).
- 24.Zhu, W., Wang, X., Ni, Y. & Xie, G. GAML-BERT: Improving BERT early exiting by gradient aligned mutual learning. In Moens, M.-F., Huang, X., Specia, L. & Yih, S. W.-t. (eds.) Proc. Conference on Empirical Methods in Natural Language Processing, 3033–3044 (Association for Computational Linguistics, 2021). https://aclanthology.org/2021.emnlp-main.242.
- 25.Morrow, D. A. et al. Evolution of critical care cardiology: transformation of the cardiovascular intensive care unit and the emerging need for new medical staffing and training models: a scientific statement from the american heart association. Circulation126, 1408–1428 (2012). [DOI] [PubMed] [Google Scholar]
- 26.Zheng, Y. et al. Large language models for medicine: a survey. International Journal of Machine Learning and Cybernetics 1–26 (Springer, 2024).
- 27.Zhang, Z. et al. PCEE-BERT: Accelerating BERT inference via patient and confident early exiting. In Carpuat, M., de Marneffe, M.-C. & Meza Ruiz, I. V. (eds.) Findings of the Association for Computational Linguistics: NAACL 2022, 327–338 (Association for Computational Linguistics, 2022). https://aclanthology.org/2022.findings-naacl.25.
- 28.Gueziri, H.-E., McGuffin, M. J. & Laporte, C. Latency management in scribble-based interactive segmentation of medical images. IEEE Trans. Biomed. Eng.65, 1140–1150 (2018). [DOI] [PubMed] [Google Scholar]
- 29.Zhu, M. & Gupta, S. To prune, or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878 (2017).
- 30.Xu, C., Zhou, W., Ge, T., Wei, F. & Zhou, M. BERT-of-theseus: Compressing BERT by progressive module replacing. In Webber, B., Cohn, T., He, Y. & Liu, Y. (eds.) Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 7859–7869 (Association for Computational Linguistics, 2020). https://aclanthology.org/2020.emnlp-main.633.
- 31.Jiao, X. et al. TinyBERT: Distilling BERT for natural language understanding. In Cohn, T., He, Y. & Liu, Y. (eds.) Findings of the Association for Computational Linguistics: EMNLP 2020, 4163–4174 (Association for Computational Linguistics, 2020). https://aclanthology.org/2020.findings-emnlp.372.
- 32.Sun, S., Cheng, Y., Gan, Z. & Liu, J. Patient knowledge distillation for BERT model compression. In Inui, K., Jiang, J., Ng, V. & Wan, X. (eds.) Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 4323–4332 (Association for Computational Linguistics, 2019). https://aclanthology.org/D19-1441.
- 33.Zhang, W. et al. TernaryBERT: Distillation-aware ultra-low bit BERT. In Webber, B., Cohn, T., He, Y. & Liu, Y. (eds.) Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 509–521 (Association for Computational Linguistics, Online, 2020). https://aclanthology.org/2020.emnlp-main.37.
- 34.Kim, S., Gholami, A., Yao, Z., Mahoney, M. W. & Keutzer, K. I-bert: Integer-only bert quantization. In International conference on machine learning, 5506–5518 (PMLR, 2021).
- 35.Scardapane, S., Scarpiniti, M., Baccarelli, E. & Uncini, A. Why should we add early exits to neural networks? Cogn. Comput.12, 954–966 (2020). [Google Scholar]
- 36.Teerapittayanon, S., McDanel, B. & Kung, H.-T. Branchynet: Fast inference via early exiting from deep neural networks. In 2016 23rd international conference on pattern recognition (ICPR), 2464–2469 (IEEE, 2016).
- 37.Xin, J., Tang, R., Lee, J., Yu, Y. & Lin, J. DeeBERT: Dynamic early exiting for accelerating BERT inference. In Jurafsky, D., Chai, J., Schluter, N. & Tetreault, J. (eds.) Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2246–2251 (Association for Computational Linguistics, Online, 2020). https://aclanthology.org/2020.acl-main.204.
- 38.Kaya, Y., Hong, S. & Dumitras, T. Shallow-deep networks: Understanding and mitigating network overthinking. In Chaudhuri, K. & Salakhutdinov, R. (eds.) Proceedings of the 36th International Conference on Machine Learning, vol. 97 of Proceedings of Machine Learning Research, 3301–3310 (PMLR, 2019). https://proceedings.mlr.press/v97/kaya19a.html.
- 39.Gao, X., Zhu, W., Gao, J. & Yin, C. F-pabee: Flexible-patience-based early exiting for single-label and multi-label text classification tasks. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5 (2023).
- 40.Yuan, D. et al. μ-net: Medical image segmentation using efficient and effective deep supervision. Computers Biol. Med.160, 106963 (2023). [DOI] [PubMed] [Google Scholar]
- 41.Kang, S. et al. An efficient and effective ensemble of support vector machines for anti-diabetic drug failure prediction. Expert Syst. Appl.42, 4265–4273 (2015). [Google Scholar]
- 42.Lan, Z. et al. ALBERT: A lite BERT for self-supervised learning of language representations. CoRRabs/1909.11942 (2019). arXiv: 1909.11942
- 43.Team, Q. et al. Qwen2 technical report. arXiv preprint arXiv:2407.106712 (2024).
- 44.Agarwal, S. et al. gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925 (2025).
- 45.Labrak, Y. et al. Biomistral: A collection of open-source pretrained large language models for medical domains. In Findings of the Association for Computational Linguistics: ACL 2024, 5848–5864 (ACL, 2024).
- 46.Sallinen, A. et al. Llama-3-meditron: an open-weight suite of medical llms based on llama-3.1. In Workshop on Large Language Models and Generative AI for Health at AAAI 2025 (AAAI, 2025).
- 47.Hager, P. et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat. Med.30, 2613–2622 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Kermany, D. S. et al. Identifying medical diagnoses and treatable diseases by image-based deep learning. Cell172, 1122–1131 (2018). [DOI] [PubMed] [Google Scholar]
- 49.Yang, J. et al. Medmnist v2-a large-scale lightweight benchmark for 2D and 3D biomedical image classification. Sci. Data10, 41 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Kather, J. N. et al. Predicting survival from colorectal cancer histology slides using deep learning: a retrospective multicenter study. PLoS Med.16, e1002730 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Fan, Y., He, L., Pakhomov, S. V., Melton, G. B. & Zhang, R. Classifying supplement use status in clinical notes. AMIA Summits Transl. Sci. Proc.2017, 493 (2017). [PMC free article] [PubMed] [Google Scholar]
- 52.Fan, Y. & Zhang, R. Using natural language processing methods to classify use status of dietary supplements in clinical notes. BMC Med. Inform. Decis. Mak.18, 15–22 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Vaswani, A. et al. Attention is all you need. In Guyon, I.et al. (eds.) Advances in Neural Information Processing Systems, vol. 30 (Curran Associates, Inc., 2017). https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
- 54.Krizhevsky, A., Sutskever, I. & Hinton, G. E. Imagenet classification with deep convolutional neural networks. In Pereira, F., Burges, C., Bottou, L. & Weinberger, K. (eds.) Advances in Neural Information Processing Systems, vol. 25 (Curran Associates, Inc., 2012). https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf.
- 55.Huang, G. et al. Multi-scale dense convolutional networks for efficient prediction. arXiv preprint arXiv:1703.098442 (2017).
- 56.Wolf, T. et al. Transformers: State-of-the-art natural language processing. In Liu, Q. & Schlangen, D. (eds.) Proc. Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 38–45 (Association for Computational Linguistics, Online, 2020). https://aclanthology.org/2020.emnlp-demos.6.
- 57.Gräßer, F., Kallumadi, S., Malberg, H. & Zaunseder, S. Aspect-based sentiment analysis of drug reviews applying cross-domain and cross-data learning. Proc. International Conference on Digital Healthhttps://api.semanticscholar.org/CorpusID:5040048 (2018).
- 58.Yu, B., Li, Y. & Wang, J. Detecting causal language use in science findings. In Proc. Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 4664–4674 (Association for Computational Linguistics, 2019). https://aclanthology.org/D19-1473.
- 59.Segura-Bedmar, I., Martínez, P. & Herrero-Zazo, M. SemEval-2013 task 9 : Extraction of drug-drug interactions from biomedical texts (DDIExtraction 2013). In Manandhar, S. & Yuret, D. (eds.) Second Joint Conference on Lexical and Computational Semantics (*SEM), Volume 2: Proceedings of the Seventh International Workshop on Semantic Evaluation (SemEval 2013), 341–350 (Association for Computational Linguistics, Atlanta, Georgia, USA, 2013). https://aclanthology.org/S13-2056.
- 60.Li, M., Chen, M., Zhou, H. & Zhang, R. Petailor: Improving large language model by tailored chunk scorer in biomedical triple extraction. arXiv preprint arXiv:2310.18463 (2023).
- 61.Sun, Z. et al. PHEE: a dataset for pharmacovigilance event extraction from text. In Goldberg, Y., Kozareva, Z. & Zhang, Y. (eds.) Proc. Conference on Empirical Methods in Natural Language Processing, 5571–5587 (Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 2022). https://aclanthology.org/2022.emnlp-main.376/.
- 62.Al-Dhabyani, W., Gomaa, M., Khaled, H. & Fahmy, A. Dataset of breast ultrasound images. Data Brief.28, 104863 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Acevedo, A. et al. A dataset of microscopic peripheral blood cell images for development of automatic recognition systems. Data Brief.30, 105474 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Tschandl, P., Rosendahl, C. & Kittler, H. The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Sci. Data5, 1–9 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Codella, N. et al. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic). arXiv preprint arXiv:1902.03368 (2019).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Eleven datasets involved in this study are publicly available from the following links:Drug Review: https://archive.ics.uci.edu/dataset/462/drug+review+dataset+drugs+comMedical Health Advice: https://huggingface.co/datasets/medalpaca/medical_meadow_health_adviceDDI: https://github.com/isegura/DDICorpusGIT: https://github.com/ToneLi/BIoMedRAG/tree/main/dataset/0_GM-CIHTPHEE: https://github.com/zhaoyuesun/pheePathMNIST, PneumoniaMNIST, DermaMNIST, BloodMNIST and BreastMNIST: https://medmnist.com/.
The code is publicly available on Github: https://github.com/Learner4everrr/EPEE.














