Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 May 21;31(9):912–921. doi: 10.1002/resp.70266

Few‐Shot Lung Cancer Classification via Electronic Nose Using Large Language Models: A Multicentre Prospective Study

Meng‐Rui Lee 1,2, Chien‐Chi Huang 3, Joyce Yue Sun 4, Chang‐Ru Lin 1,5, Wen‐Yuan Lin 1, Nai‐Hui Chi 6, Kea‐Tiong Tang 3,✉, Jann‐Yuan Wang 1, Chao‐Chi Ho 1, Jin‐Yuan Shih 1, Chong‐Jen Yu 1
PMCID: PMC13535749  PMID: 42167356

ABSTRACT

Background and Objective

Electronic Nose (eNose) breathprints are promising non‐invasive lung cancer diagnostic tools, but cross‐site validation and adaptation remain barriers to clinical applications. It remains unknown whether a natural language processing‐pretrained large language model (LLM) can enable few‐shot, site‐specific classification of lung cancer using eNose breathprints.

Methods

We collected eNose breathprints of lung cancer and non‐lung cancer patients from two medical centres in Taiwan. A GPT‐2–backbone LLM with parameter‐efficient adaptation was compared with convolutional neural networks (CNN) trained from scratch or pretrained on CIFAR‐100. Few‐shot protocols (2–6 shots per class) and full‐data training were evaluated.

Results

We collected 432 eNose breathprints from two sites (S1 and S2). With 6 labelled samples per class (6 shots), LLM achieved an area under the curve (AUC) of 0.79 (95% CI: 0.71–0.87), sensitivity of 0.74 (0.63–0.83), and specificity of 0.77 (0.67–0.87) on S1. On S2, it achieved an AUC of 0.76 (0.69–0.82), sensitivity of 0.77 (0.69–0.84), and specificity of 0.61 (0.51–0.70). LLM outperforms scratch CNN models (S1; AUC: 0.44, p = 0.0002) (S2; AUC: 0.63, p = 0.0198) and CNN pretrained on CIFAR‐100 images (S1; AUC: 0.57, p = 0.0100) and (S2; AUC: 0.61, p = 0.0248). LLM or a CNN model trained on the source site fails to improve performance after transferring to the target site for fine‐tuning; for the LLM, performance even deteriorates.

Conclusion

Our study demonstrates the potential of pretrained LLMs for few‐shot lung cancer classification in a real‐world mixed clinical cohort, reducing dependence on large training datasets.

graphic file with name RESP-31-912-g003.webp

Keywords: breathprint, diagnosis, electronic nose, large language model, lung cancer


Few‐shot adaptation of pretrained large language models for lung cancer classification using electronic nose breathprints. LLM achieved AUC 0.76–0.79 with only 6 samples per class, significantly outperforming CNN models in few‐shot scenarios. LLM facilitates practical eNose‐based lung cancer classification by reducing dependence on large training datasets.

graphic file with name RESP-31-912-g001.webp

1. Introduction

Lung cancer remains the leading cause of cancer‐related mortality worldwide, with an estimated 2.5 million new cases and 1.8 million deaths in 2022, accounting for nearly 18.7% of all global cancer deaths [1]. Although early‐stage detection significantly improves prognosis, current imaging‐based screening methods such as low‐dose chest computed tomography (LDCT) are still limited by high costs, radiation exposure, and limited accessibility [2, 3].

Electronic nose (eNose), which detects volatile organic compounds (VOCs) in exhaled breath, offers a promising non‐invasive, rapid, and cost‐effective alternative for lung cancer detection [4]. Numerous studies have demonstrated that machine learning and convolutional neural network (CNN)‐based models can classify eNose breathprints with high in‐cohort accuracy [5, 6, 7]. A systematic review of 35 trials reported a pooled sensitivity of 0.90, a specificity of 0.89, and a median area under the curve (AUC) of 0.91 for breath‐based lung cancer detection [8]. In one study, researchers proposed a breath‐analysis system with a 1‐D CNN, achieving 92% accuracy on a test set of 181 breath samples [9]. Meanwhile, one study analysed 199 breath samples and achieved 79% accuracy for lung cancer detection with XGBoost‐based ensemble [10].

Despite these encouraging results, problems exist regarding eNose real‐world application. First, one of the major challenges impeding clinical translation is its poor generalizability across sites. This drop in cross‐site performance is largely attributed to variability in environmental VOC backgrounds, sensor calibration drift, and population heterogeneity [11, 12]. Also, most existing models rely on access to large quantities of labelled data to enhance performance [13]. This may fail to address practical deployment scenarios in which only a small number of labelled data points may be available from a new site. Consequently, few‐shot adaptation for eNose‐based lung cancer screening remains an open and largely unexplored challenge. Unlike CT imaging for lung cancer, where large‐scale annotated datasets collected from multiple scanners and sites are increasingly available, most publicly accessible eNose datasets are limited to basic gas classification tasks, and large‐scale, multi‐site datasets specifically focused on lung cancer detection remain scarce [14].

Large language models (LLM) are deep learning‐based systems trained on massive text datasets and originally developed to perform diverse natural language processing tasks. Built on transformer architectures, they are capable of capturing long‐range dependencies, which enables them to excel across a wide spectrum of natural language processing (NLP) applications [15]. Recently the applications of LLM have extended beyond language and texts. LLMs have attracted attention in few‐shot learning scenarios since these models have shown remarkable cross‐domain knowledge transfer capabilities, offering a promising path forward in data‐scarce environments [16]. Also, there has been growing interest in adapting LLMs to time‐series domains such as eNose breathprints, where data scarcity and heterogeneity are common challenges [17, 18].

It, therefore, is intriguing whether LLMs could serve as effective representations, enabling few‐shot adaptation for lung cancer classification through eNose. To address this knowledge gap, we initiated this prospective multicentre study to evaluate the application of LLMs in lung cancer classification through eNose breathprints.

2. Methods

2.1. Ethics Statement

This study was prospectively conducted in accordance with established ethical principles and standards. The institutional review boards (IRB) of participating hospitals approved this study (IRB no. 202112057RINB, 108‐011‐E). Informed consent was obtained from all participants who agreed to participate in this study.

2.2. Study Participants and Setting

This study was conducted in two tertiary referral medical centres in Taiwan, including National Taiwan University Hospital Hsin‐chu branch hospital (NTUH‐HC; Site 1 [S1]) and National Taiwan University Hospital (NTUH; Site 2 [S2]) during 2019 and 2023. We recruited participants who are either lung cancer or non‐lung cancer patients. In S1, we recruited participants who were lung cancer patients admitted to chest wards and patients who received pulmonary function testing in examination rooms. In S2, we recruited participants who were admitted to the chest medicine ward due to either lung cancer or other pulmonary diseases and patients who visited the outpatient department due to structural lung diseases. Additionally, we also recruited a cohort of chronic obstructive pulmonary disease (COPD) patients in S2. In both hospitals, we recruited healthy control patients and, in this way, we hope to better approximate real‐world lung cancer screening settings among which healthy and diseased controls were both present.

The absence of lung cancer was ascertained among the non‐lung cancer group through chest CT imaging and follow‐up evaluations. During a two‐year follow‐up period, all control participants, encompassing both healthy and diseased controls, remained free from lung cancer.

2.3. Definition and Data Collection

For lung cancer cases, diagnosis required pathological confirmation, and staging followed the American Joint Committee on Cancer (AJCC) 9th edition lung cancer system [19]. Data were obtained from a prospectively maintained database and electronic medical records. Recorded comorbidities included COPD, asthma, diabetes mellitus (DM), and end‐stage renal disease (ESRD). Healthy participants underwent a screening interview to exclude prior lung disease and tobacco use; when available, chest radiographs (CXRs) were reviewed to rule out structural lung abnormalities.

2.4. Breath Sampling Technique

The breath‐sampling apparatus comprised a single‐use one‐way bacterial/viral filter (VBMax Pulmonary Function Testing Filter, A‐M Systems Inc.) and two 1‐L multilayer foil gas‐sampling bags. A disposable mouthpiece was attached to the filter inlet, and the outlet split into two parallel tubes equipped with Robert clamps. One tube (Clamp 1) permitted venting of the initial dead‐space air; the other (Clamp 2) directed end‐tidal breath into a gas‐sampling bag.

All participants fasted for at least 4 h, abstained from smoking and alcohol, and refrained from drinking water for at least 1 h before sampling. All equipment was cleaned and purged with nitrogen to minimize carryover and VOC contamination. During sampling, participants took a deep inspiration and exhaled into the system. Clamp 1 was opened for 3–5 s to discard dead‐space air, then closed; Clamp 2 was subsequently opened for 90 s to collect the end‐tidal breath sample. Each participant contributed one single eNose for analysis. Breath sampling and eNose measurement were performed within the same day.

2.5. eNose Breathprints

All samples were analysed with the SEXTANT eNose (Enosim Bio‐Tech Co. Ltd., Hsinchu City, Taiwan), which integrates 14 metal‐oxide semiconductor gas sensors together with flow, temperature and humidity sensors. For each sample, the time‐series signal of the electrical resistance from each sensor is recorded and pre‐processed to a fixed‐length sequence of 256 data points [6, 7, 20]. Specifically, to standardized the input for the model, a fixed temporal window of 768 s is extracted. The window begins 15 s prior to the gas entering the eNose to establish a stable baseline. Given the hardware's sampling rate of 1/3 Hz, this 768‐s signal is down‐sampled to 256 data points per sensor.

2.6. eNose Model Architecture

For the analysis of eNose breathprints, we used LLM and CNN as two approaches. For the LLM approach, to adapt the continuous eNose signals for the LLM framework, we convert the multichannel time series signal into a sequence of patches. First, the raw signal from 14 sensors is concatenated into a 1D signal. This signal is then divided into N non‐overlapping patches, each patch is projected into a learnable embedding, effectively translating the raw signal into a representation the model can process. After appending positional information to preserve the chronological order of the original signals, the patch embeddings are passed through Transformer blocks, where the self‐attention mechanism analyzes all patches to find critical relationships across different time points and produces the final diagnostic. The workflow of LLM is also illustrated in Figure 1. As detailed in Figure S1, we employ an efficient fine‐tuning strategy to leverage ability of the pre‐trained model while mitigating overfitting under few‐shot conditions. We freeze the core Multi‐head Attention and Feed Forward Layers, exclusively updating the Input and Output Layers, Positional Embeddings, and Layer Norms. For CNN, the input is a 2D eNose signals (14 channels × 256 timepoints; reshaped to 14 × 16 × 16) Figure S2.

FIGURE 1.

FIGURE 1

Schematic illustration of large language model workflow on electronic nose signals.

2.7. Few‐Shot Scenario Performance

To simulate real‐world clinical scenarios with extremely limited labelled data, we conduct few‐shot fine‐tuning on the target site (S1 or S2). In the few‐shot setting, we primarily adopt an n‐shot protocol, that is, n labelled samples per class from the training set of the target site; 3‐shot, for instance, means 3 lung cancer and 3 non‐lung cancer cases for training.

We designed a series of experiments to evaluate model performance under both few‐shot and full‐data conditions, using three random seeds. To ensure consistency across all experiments, the validation and test sets were kept identical within each seed, with only the number of training samples varying between the few‐shot and full‐data settings. Specifically, for S1, the validation/testing samples were fixed at the same 43 and 45 patients (testing: 22 lung cancer/21 controls; validation: 23 lung cancer/22 controls) while for S2, the validation/testing samples were fixed at the same 72 and 72 patients (testing: 38 lung cancer/34 controls; validation: 41 lung cancer/31 controls). Two sites share the maximal training samples (full data, if used) to be 100 patients (S1 training: 48 lung cancer/52 controls; S2 training: 65 lung cancer/35 controls).

2.8. Comparison of Pretraining Strategies in Few‐Shot Scenarios

We compared the following three main model variants (LLM, CNN‐scratch, CNN‐CIFAR100) to systemically investigate whether pretraining could enhance the classification performance of CNN and LLM models on eNose data. CNN‐scratch refers to a baseline model trained entirely from randomly initialized weights without any form of pretraining. CNN‐CIFAR100 was pretrained the CNN on a large‐scale general image dataset (CIFAR‐100) [21], then fine‐tuned it on the eNose dataset. LLM retained the GPT‐2 weights obtained from large‐scale natural language pretraining and was fine‐tuned on eNose data with only a minimal number of trainable parameters (∼1.8 M).

2.9. Cross‐Site Generalization

To investigate whether cross‐site generalization could be achieved, models were first trained to convergence using full labelled data from a source site (e.g., S1), obtaining source‐site‐specific pretrained weights. These source‐pretrained models were then transferred to a target site (e.g., S2), fine‐tuned on limited target‐site data, and subsequently evaluated.

By comparing the performance of cross‐site pretrained models (e.g., CNN‐S1‐pt (pretrained), LLM‐S1‐pt(pretrained)) against baseline models without any source‐site‐specific pretraining (CNN‐scratch, LLM), we aimed to clarify whether directly transferring eNose pretrained weights across sites could enhance few‐shot adaptation at the target site. The analytic approach is summarized in Figure S3.

2.10. Statistical Analysis

Model performance was evaluated by accuracy, sensitivity, specificity, and AUC. For the sensitivity analysis, we performed subgroup analyses stratified by staging and histology in lung cancer patients, and by disease subtype (COPD and non‐lung cancers) in the control group under 6‐shot training. A covariate‐matched analysis was also conducted by matching on age, COPD status, smoking history, and histology in the lung cancer group, and on age, COPD status, and smoking history in the control group, using 1:1 nearest‐neighbour propensity score matching. 95% confidence intervals (CI) were estimated by bootstrap resampling. All models were implemented with the PyTorch package (version 2.5.1) in Python (version 3.11.0) and trained on a single NVIDIA GeForce GTX 1080 Ti GPU.

3. Results

3.1. Clinical Characteristics of Participants

A total of 432 participants were enrolled between 2019 and 2023: Figure S3. S1 contributed 188 participants (93 lung cancer, 95 non–lung cancer controls), and S2 contributed 244 (144 lung cancer, 100 controls). Baseline characteristics are summarized in Table 1. The mean age was approximately 60 years, and the sex distribution was roughly balanced across sites. Among lung cancer cases (n = 237), adenocarcinoma predominated (n = 167, 70.5%), followed by small‐cell lung carcinoma and squamous cell carcinoma. Most lung cancers were stage IV at diagnosis, whereas 8.9% were early‐stage (I–II). The detailed demographic data of participants are described in Table 1. Lung cancer histology classification and staging are described in detail in Table S1.

TABLE 1.

Clinical characteristics of 432 participants.

Site 1 Site 2
All (n = 432) All (n = 188) Lung cancer (n = 93) Control (n = 95) p All (n = 244) Lung cancer (n = 144) Control (n = 100) p p* (lung cancer vs. control) p** (Site 1 vs. Site 2)
Male/Female 232 (53.7)/200 (46.3) 96 (51.1)/92 (48.9) 48 (51.6)/45 (48.4) 48 (50.5)/47 (49.5) 0.8816 136 (55.7)/108 (44.3) 85 (59.0)/59 (41.0) 51 (51.0)/49 (49.0) 0.2144 0.2672 0.3341
Age (mean ± SD) 62.6 ± 14.2 61.0 ± 15.9 63.9 ± 12.0 58.1 ± 18.6 0.0110 63.8 ± 12.7 64.7 ± 10.4 62.7 ± 15.4 0.2614 0.0059 0.0420
Smoking
Current 56 (13.0) 20 (10.6) 8 (8.6) 12 (12.6) 0.1200 36 (14.8) 25 (17.4) 11 (11.0) 0.3077 0.5759 0.0104
Ex‐Smoker 124 (28.7) 43 (22.9) 27 (29.0) 16 (16.8) 81 (33.2) 44 (30.6) 37 (37.0)
Never 252 (58.3) 125 (66.5) 58 (62.4) 67 (70.5) 127 (52.0) 75 (52.1) 52 (52.0)
DM 62 (14.4) 22 (11.7) 11 (11.8) 11 (11.6) 0.9577 40 (16.4) 32 (22.2) 8 (8.0) 0.0032 0.0132 0.1679
Hypertension 135 (31.3) 58 (30.9) 31 (33.3) 27 (28.4) 0.4660 77 (31.6) 59 (41.0) 18 (18.0) 0.0002 0.0009 0.8752
Asthma 21 (4.9) 6 (3.2) 0 6 (6.3) 0.0288 15 (6.1) 4 (2.8) 11 (11.0) 0.0128 0.0011 0.1810
COPD 107 (24.8) 32 (17.0) 11 (11.8) 21 (22.1) 0.0608 75 (30.7) 21 (14.6) 54 (54.0) < 0.0001 < 0.001 0.0011
Cancer 256 (59.3) 95 (50.5) 93 (100.0) 2 (2.1) < 0.0001 161 (66.0) 144 (100.0) 17 (17) < 0.0001 < 0.001 0.0010

Abbreviations: COPD, chronic obstructive pulmonary disease; DM, diabetes mellitus; SD, standard deviation.

3.2. Few‐Shot Classification Performance

The few‐shot performance of our model was illustrated in Table 2 and Figure 2. With only 3 shots (3 labelled samples per class), the AUC of NLP‐pretrained LLM achieved an AUC of 0.77 (95% CI: 0.70–0.86) in S1 while CNN‐scratch and CNN‐CIFAR100 performed only 0.47 (0.37–0.57, p = 0.0042) and 0.58 (0.48–0.67, p = 0.023). In S2 with 3 shots, LLM achieved an AUC of 0.63 (0.55–0.70) compared with CNN‐scratch of 0.55 (0.47–0.63, p = 0.2691) and CNN‐CIFAR100 of 0.57 (0.48–0.65, p = 0.4237). NLP‐pretrained LLM achieved sensitivity of 0.78 (95% CI: 0.68–0.87) on S1 and 0.68 (0.59–0.75) on S2 and specificity of 0.71 (0.60–0.82) on S1 and 0.49 (0.39–0.58) on S2. NLP‐pretrained LLM had positive predictive value (PPV) of 0.75 (95% CI: 0.64–0.84) on S1 and 0.58 (0.50–0.67) on S2 and negative predictive value (NPV) of 0.75 (0.63–0.85) on S1 and 0.59 (0.49–0.69) on S2.

TABLE 2.

Few‐shot performance of tested models at single site.

S1 S2
LLM (GPT‐2) CNN‐Scratch CNN‐CIFAR100 LLM (GPT‐2) CNN‐Scratch CNN‐CIFAR100
2 shots AUC (95% CI) 0.73 (0.65–0.82) 0.55 (0.44–0.64) 0.48 (0.38–0.58) 0.59 (0.52–0.67) 0.59 (0.51–0.68) 0.61 (0.53–0.69)
ACC (95% CI) 0.72 (0.65–0.80) 0.58 (0.50–0.66) 0.63 (0.55–0.71) 0.60 (0.53–0.67) 0.59 (0.51–0.68) 0.58 (0.51–0.65)
Sensitivity (95% CI) 0.76 (0.67–0.87) 0.93 (0.87–0.99) 0.96 (0.91–1.00) 0.69 (0.61–0.78) 0.86 (0.79–0.93) 0.81 (0.74–0.88)
Specificity (95% CI) 0.68 (0.56–0.79) 0.20 (0.11–0.29) 0.96 (0.91–1.00) 0.49 (0.39–0.60) 0.29 (0.21–0.39) 0.33 (0.24–0.43)
PPV (95% CI) 0.72 (0.62–0.82) 0.56 (0.47–0.65) 0.59 (0.50–0.68) 0.59 (0.50–0.68) 0.56 (0.48–0.64) 0.56 (0.48–0.64)
NPV (95% CI) 0.73 (0.61–0.84) 0.72 (0.50–0.93) 0.86 (0.69–1.00) 0.60 (0.50–0.70) 0.67 (0.53–0.81) 0.63 (0.49–0.76)
3 shots AUC (95% CI) 0.77 (0.70–0.86) 0.47 (0.37–0.57) 0.58 (0.48–0.67) 0.63 (0.55–0.70) 0.55 (0.47–0.63) 0.57 (0.48–0.65)
ACC (95% CI) 0.75 (0.67–0.82) 0.47 (0.37–0.57) 0.61 (0.53–0.70) 0.63 (0.55–0.70) 0.54 (0.47–0.61) 0.52 (0.45–0.60)
Sensitivity (95% CI) 0.78 (0.68–0.87) 0.47 (0.37–0.57) 0.96 (0.91–1.00) 0.68 (0.59–0.75) 0.81 (0.73–0.87) 0.72 (0.64–0.81)
Specificity (95% CI) 0.71 (0.60–0.82) 0.26 (0.16–0.37) 0.23 (0.13–0.33) 0.49 (0.39–0.58) 0.26 (0.17–0.36) 0.31 (0.23–0.40)
PPV (95% CI) 0.75 (0.64–0.84) 0.57 (0.48–0.66) 0.57 (0.48–0.66) 0.58 (0.50–0.67) 0.54 (0.46–0.61) 0.53 (0.44–0.61)
NPV (95% CI) 0.75 (0.63–0.85) 0.71 (0.52–0.90) 0.83 (0.64–1.00) 0.59 (0.49–0.69) 0.56 (0.40–0.69) 0.52 (0.39–0.65)
4 shots AUC (95% CI) 0.77 (0.70–0.85) 0.56 (0.46–0.65) 0.55 (0.45–0.65) 0.73 (0.65–0.79) 0.65 (0.58–0.73) 0.63 (0.54–0.71)
ACC (95% CI) 0.73 (0.66–0.80) 0.59 (0.52–0.67) 0.57 (0.49–0.64) 0.70 (0.63–0.76) 0.67 (0.60–0.73) 0.62 (0.55–0.69)
Sensitivity (95% CI) 0.75 (0.64–0.85) 0.93 (0.87–0.99) 0.76 (0.66–0.86) 0.69 (0.61–0.78) 0.71 (0.61–0.79) 0.78 (0.69–0.86)
Specificity (95% CI) 0.75 (0.64–0.85) 0.23 (0.13–0.33) 0.35 (0.24–0.47) 0.70 (0.60–0.79) 0.63 (0.53–0.72) 0.45 (0.35–0.54)
PPV (95% CI) 0.74 (0.64–0.84) 0.57 (0.48–0.66) 0.56 (0.46–0.66) 0.71 (0.62–0.79) 0.67 (0.58–0.75) 0.60 (0.51–0.68)
NPV (95% CI) 0.72 (0.61–0.83) 0.75 (0.54–0.93) 0.57 (0.42–0.73) 0.68 (0.59–0.77) 0.67 (0.58–0.76) 0.66 (0.54–0.77)
5 shots AUC (95% CI) 0.76 (0.68–0.84) 0.46 (0.37–0.56) 0.56 (0.45–0.65) 0.72 (0.65–0.79) 0.63 (0.56–0.71) 0.66 (0.58–0.73)
ACC (95% CI) 0.72 (0.65–0.80) 0.48 (0.40–0.56) 0.62 (0.54–0.70) 0.68 (0.61–0.74) 0.62 (0.56–0.69) 0.57 (0.50–0.63)
Sensitivity (95% CI) 0.72 (0.62–0.82) 0.54 (0.43–0.65) 0.92 (0.85–0.97) 0.69 (0.61–0.78) 0.77 (0.68–0.84) 0.74 (0.66–0.82)
Specificity (95% CI) 0.73 (0.62–0.84) 0.41 (0.29–0.53) 0.29 (0.18–0.39) 0.66 (0.56–0.75) 0.47 (0.37–0.57) 0.38 (0.29–0.48)
PPV (95% CI) 0.74 (0.64–0.85) 0.50 (0.40–0.61) 0.58 (0.50–0.68) 0.68 (0.59–0.77) 0.61 (0.52–0.69) 0.56 (0.48–0.64)
NPV (95% CI) 0.71 (0.59–0.81) 0.45 (0.32–0.57) 0.76 (0.58–0.93) 0.67 (0.58–0.76) 0.66 (0.54–0.76) 0.58 (0.46–0.70)
6 shots AUC (95% CI) 0.79 (0.71–0.87) 0.44 (0.35–0.54) 0.57 (0.47–0.67) 0.76 (0.69–0.82) 0.63 (0.55–0.71) 0.61 (0.53–0.68)
ACC (95% CI) 0.75 (0.68–0.83) 0.47 (0.38–0.55) 0.61 (0.53–0.69) 0.69 (0.63–0.75) 0.63 (0.56–0.69) 0.56 (0.50–0.63)
Sensitivity (95% CI) 0.74 (0.63–0.83) 0.60 (0.48–0.71) 0.60 (0.48–0.71) 0.77 (0.69–0.84) 0.77 (0.69–0.84) 0.78 (0.70–0.85)
Specificity (95% CI) 0.77 (0.67–0.87) 0.33 (0.21–0.45) 0.62 (0.51–0.74) 0.61 (0.51–0.70) 0.48 (0.38–0.57) 0.33 (0.24–0.43)
PPV (95% CI) 0.78 (0.68–0.87) 0.49 (0.39–0.60) 0.63 (0.52–0.74) 0.67 (0.59–0.76) 0.61 (0.52–0.69) 0.55 (0.47–0.63)
NPV (95% CI) 0.73 (0.61–0.83) 0.43 (0.30–0.56) 0.59 (0.46–0.70) 0.71 (0.62–0.80) 0.66 (0.55–0.76) 0.59 (0.45–0.71)

FIGURE 2.

FIGURE 2

Receiver operator characteristics curves of different models at single site. ROC: receiver operator characteristics.

With only 6 labelled samples per class, the NLP‐pretrained LLM achieved the highest performance on both S1 and S2, reaching AUCs of 0.79 (95% CI: 0.71–0.87) and 0.76 (0.69–0.82), respectively. This outperformed CNN‐scratch (S1: AUC 0.44, 95% CI: 0.35–0.54, p = 0.0002; S2: AUC 0.63, 0.55–0.71, p = 0.0198) and CNN‐CIFAR100 (S1: AUC 0.57, 0.47–0.67, p = 0.0100; S2: AUC 0.61, 0.53–0.68, p = 0.0248). NLP‐pretrained LLM also achieved sensitivity of 0.74 (95% CI: 0.63–0.83) on S1 and 0.77 (0.69–0.84) on S2 and specificity of 0.77 (0.67–0.87) on S1 and 0.61 (0.51–0.70) on S2. NLP‐pretrained LLM had PPV of 0.78 (95% CI: 0.68–0.87) on S1 and 0.67 (0.59–0.76) on S2 and NPV of 0.73 (0.61–0.83) on S1 and 0.71 (0.62–0.80) on S2. The detailed few‐shot performance with different number of training data is illustrated in Figure 3.

FIGURE 3.

FIGURE 3

Single‐site performance as the number of training data increases.

3.3. Cross‐Site Generalization

On S1, the LLM pretrained on S2 (LLM‐S2‐pretrained) achieved an AUC of 0.33 (95% CI: 0.24–0.43, p = 0.0002) on 3‐shot and 0.35 (0.26–0.46, p = 0.0002) on 6‐shot. On S2, the LLM pretrained on S1 (LLM‐S1‐pretrained) achieved an AUC of 0.50 (0.42–0.58, p = 0.0056) on 3‐shot and 0.64 (0.57–0.71, p = 0.0174) on 6‐shot, both lower than the LLM with only NLP pretraining.

For CNNs, the cross‐site results closely mirrored those observed for the LLMs. On S1, CNN‐S2‐pretrained achieved AUC of 0.37 (95% CI: 0.28–0.47, p = 0.0002) on 3‐shot and 0.34 (0.25–0.43, p = 0.0002) on 6‐shot. On S2, CNN‐S1‐pretrained achieved AUC of 0.55 (0.48–0.63, p = 0.1744) on 3‐shot and 0.64 (0.56–0.71, p = 0.0324) on 6‐shot, which are also lower than LLM with only NLP pretraining. The detailed cross‐site performance with different number of training data was illustrated in Figure 4. The detailed few‐shot performance of tested models at cross‐site is described in Table S2.

FIGURE 4.

FIGURE 4

Cross‐site performance as the number of training data increases.

3.4. Full‐Data Performance

On S1, the LLM achieved an AUC of 0.93 (95% CI: 0.87–0.97), CNN‐scratch achieved 0.95 (0.90–0.98), and CNN‐CIFAR100 achieved 0.93 (0.88–0.98). The specificity was 0.83 (0.74–0.92) for the LLM, 0.91 (0.83–0.97) for CNN‐scratch, and 0.91 (0.83–0.98) for CNN‐CIFAR100. The sensitivity was 0.96 (0.90–1.00) for the LLM, 0.92 (0.85–0.97) for CNN‐scratch, and 0.89 (0.81–0.96) for CNN‐CIFAR100.

On S2, LLM achieved AUC 0.82 (0.76–0.87), a specificity of 0.82 (0.75–0.89), and a sensitivity of 0.72 (0.63–0.80). In comparison, CNN‐scratch achieved AUC 0.77 (0.70–0.83), a specificity of 0.72 (0.63–0.80), and a sensitivity of 0.70 (0.61–0.79). CNN‐CIFAR100 achieved AUC 0.76 (0.70–0.82), a specificity of 0.65 (0.56–0.70), and a sensitivity of 0.81 (0.73–0.88).

3.5. Subgroup Analysis and Covariates Matched Analysis

The results of subgroup analysis by lung cancer histology and staging are presented in Table S3. Notably, the LLM model correctly identified early‐stage lung cancer across both sites (stage I: accuracy 0.50 on S1 and 0.80 on S2; stage II: accuracy 1.00 on S2). The model also performed well in patients with COPD (accuracy 1.00 on S1 and 0.71 on S2) and in patients with cancers other than lung cancer (accuracy 1.00 on S1 and 0.63 on S2).

In the covariate‐matched analysis, the AUCs of the LLM model were 0.81 (95% CI: 0.68–0.92) on S1 and 0.71 (95% CI: 0.57–0.84) on S2.

4. Discussion

Our study demonstrates that LLM pretrained on natural language corpora can be successfully applied to eNose breathprints, for few‐shot classification task in the multi‐centre settings. Requiring only few training samples, LLM could achieve satisfactory performance in lung cancer classification and LLM significantly outperformed CNN‐based models.

Traditional serum tumour markers—such as cytokeratin 19 fragment (CYFRA 21‐1) and carcinoembryonic antigen—are widely used during routine health examinations but have limited sensitivity and specificity for lung cancer screening [22]. More recently, circulating biomarkers, including cell‐free DNA and exosomes, have gained attention for lung cancer detection [23]. At present, LDCT remains the mainstay of screening [24]. For comparison, LDCT demonstrated a sensitivity of 93.8% and a specificity of 73.4% in the initial screening round of the National Lung Screening Trial (NLST) [25]. In one meta‐analysis, the sensitivity of LDCT exceeded 80% in most studies and the specificity exceeded 75% in most studies, despite considerable variation across cohorts [26]. While the diagnostic accuracy of our few‐shot LLM model does not yet match that of LDCT, eNose offers complementary advantages, including non‐invasiveness, the absence of radiation, low cost, and suitability for point‐of‐care deployment without specialised imaging infrastructure. The clinical utility of eNose‐based classification is therefore most plausible as a triage or pre‐screening tool to identify individuals who warrant further workup, rather than as a standalone diagnostic replacement.

eNose technology has been explored for lung cancer detection for over a decade. In a meta‐analysis of 35 studies including 4483 participants, the pooled sensitivity and specificity were 0.90 (95% CI: 0.87–0.93) and 0.89 (95% CI: 0.85–0.93), respectively [8]. Despite these encouraging results, clinical translation is limited by sensor instability, nonstandardized breath‐collection protocols, demographic variability, and the need for large training datasets [27]. Environmental conditions such as humidity and temperature can induce sensor drift [5]. Sample‐size demands are substantial: a simulation study estimated that approximately 400 data points per eNose device at 50% cancer prevalence, and about 2100 at 5% prevalence, are required to achieve stable performance [28]. Taken together, the need for large training cohorts and the common drop in performance during external validation remain key barriers to routine clinical adoption [13].

Under conditions with only a small number of labelled samples, LLMs tend to adapt well to new tasks. In our study, CNN model pretrained on image datasets showed limited performance when transferred to the eNose domain. Notably, this limitation may not be solely due to differences in data modality. Study shows that models trained on huge datasets—whether from language or vision—often transfer better to new problems than smaller, task‐specific models [17]. By contrast, image‐pretrained CNNs are relatively small in capacity which may constrain transfer performance [29]. Therefore, the performance bottleneck observed for CNNs in our study is likely due to both the relatively small scale of the pretraining dataset (e.g., CIFAR‐100, ~50 k images) and the limited capacity of the small‐size model, rather than from modality mismatches.

The fact that bigger, better‐trained language models usually generalize better also support our findings. LLMs that scored higher on standard language tests also did better on brand‐new tasks without any extra training (“zero‐shot”), which suggests that model size and the quality of pretraining help them learn broad, reusable skills [18]. In our case, the LLM benefited from large‐scale pretraining on approximately 40 GB of text data, equipping them with robust, universal representations. Combined with the flexibility of the transformer architecture and universal tokenization, these models can effectively handle input from different modalities.

The cross‐site generalization results indicate that directly transferring eNose pretrained weights to another site provides limited benefits for LLM models. This suggests that eNose source‐site pretraining introduces biases that adversely affect adaptation to the target site, with the magnitude of this effect varying across sites and in LLM may overwrite universal sequence representations learned from large‐scale NLP corpora, thereby weakening few‐shot adaptability to unseen domains. These findings imply that, under domain shift and data‐scarce conditions, the more effective strategy is to retain the original NLP‐pretrained weights and applies site‐specific fine‐tuning on the target site using only a handful of labelled samples.

This study has limitations. First, the lung cancer group in our cohort consists predominantly of late‐stage (stage IV) patients. Although our subgroup analysis demonstrated comparable effectiveness in early‐stage lung cancer patients, the small sample size limits the interpretability of these findings. It should therefore be acknowledged that the high AUCs achieved in this study may be lower in a true screening population, where lung cancer prevalence is lower, disease is detected at earlier stages, and VOC signals are more subtle. Nonetheless, our results suggest the feasibility and potential of few‐shot adaptation with a large language model for lung cancer detection. Second, the cohort comprised exclusively Taiwanese patients; consequently, external validity to other populations remains uncertain.

In conclusion, LLM applied to eNose time‐series can achieve few‐shot lung cancer classification: with as few as 3–6 labelled samples per class, the model attained clinically meaningful discrimination. By reducing dependence on largetraining cohorts, this approach could make eNose‐based screening more practicable. If applicable to other diseases, our findings would significantly ameliorate the demanding task of recruiting large training cohorts.

Author Contributions

Meng‐Rui Lee: conceptualization, data curation, formal analysis, funding acquisition, investigation, methodology, writing – original draft, writing – review and editing. Chien‐Chi Huang: conceptualization, data curation, formal analysis, investigation, methodology, writing – original draft. Joyce Yue Sun: formal analysis, methodology, writing – original draft. Chang‐Ru Lin: data curation, writing – original draft. Wen‐Yuan Lin: data curation. Nai‐Hui Chi: data curation, writing – original draft. Kea‐Tiong Tang: conceptualization, funding acquisition, methodology, supervision, writing – original draft, writing – review and editing. Jann‐Yuan Wang: conceptualization, investigation, methodology, supervision, writing – review and editing. Chao‐Chi Ho: supervision, writing – review and editing. Jin‐Yuan Shih: supervision, writing – review and editing. Chong‐Jen Yu: supervision, writing – review and editing.

Funding

This study was funded by the Taiwan Ministry of Science and Technology (113‐2628‐B‐002‐017‐MY3). The funders had no role in the study design, data analysis, and manuscript writing.

Ethics Statement

This study was prospectively conducted at the National Taiwan University Hospital and National Taiwan University Hospital Hsin‐Chu Branch in accordance with established ethical principles and standards. The institutional review boards (IRB) of participating hospitals approved this study (IRB no. 202112057RINB, 108‐011‐E). Informed consent was obtained from all participants who agreed to participate in this study.

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Figure S1: Model Structural of Large Language Model.

Figure S2: Model Structural of Convolutional Neural Network.

Figure S3: Study flow and design.

Table S1: Lung Cancer Demographic data.

Table S2: Few‐shot Performance of Tested Models at Cross‐Site.

Table S3: Subgroup analysis by lung cancer histology and staging.

RESP-31-912-s001.docx (473.3KB, docx)

Acknowledgements

We would like to thank all the participants who agreed to participate in this study.

Lee M.‐R., Huang C.‐C., Sun J. Y., et al., “Few‐Shot Lung Cancer Classification via Electronic Nose Using Large Language Models: A Multicentre Prospective Study,” Respirology 31, no. 9 (2026): 912–921, 10.1002/resp.70266.

Associate Editor: Hidenori KageSenior Editor: Phan Nguyen

Data Availability Statement

Data are available upon reasonable request from the corresponding author.

References

  • 1. Bray F., Laversanne M., Sung H., et al., “Global Cancer Statistics 2022: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries,” CA: A Cancer Journal for Clinicians 74, no. 3 (2024): 229–263, 10.3322/caac.21834. [DOI] [PubMed] [Google Scholar]
  • 2. Jonas D. E., Reuland D. S., Reddy S. M., et al., “Screening for Lung Cancer With Low‐Dose Computed Tomography: Updated Evidence Report and Systematic Review for the US Preventive Services Task Force,” JAMA 325, no. 10 (2021): 971–987, 10.1001/jama.2021.0377. [DOI] [PubMed] [Google Scholar]
  • 3. Lancaster H. L., Heuvelmans M. A., and Oudkerk M., “Low‐Dose Computed Tomography Lung Cancer Screening: Clinical Evidence and Implementation Research,” Journal of Internal Medicine 292, no. 1 (2022): 68–80, 10.1111/joim.13480. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Rocco G., Pennazza G., Tan K. S., et al., “A Real‐World Assessment of Stage I Lung Cancer Through Electronic Nose Technology,” Journal of Thoracic Oncology 19, no. 9 (2024): 1272–1283, 10.1016/j.jtho.2024.05.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Ye Z., Liu Y., and Li Q., “Recent Progress in Smart Electronic Nose Technologies Enabled With Machine Learning Methods,” Sensors 21, no. 22 (2021): 7620, 10.3390/s21227620. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Yu K.‐L., Yang H.‐C., Lee C.‐F., et al., “Exhaled Breath Analysis Using a Novel Electronic Nose for Different Respiratory Disease Entities,” Lung 203 (2025): 14, 10.1007/s00408-024-00776-1. [DOI] [PubMed] [Google Scholar]
  • 7. Lee M.‐R., Kao M.‐H., Hsieh Y.‐C., et al., “Cross‐Site Validation of Lung Cancer Diagnosis by Electronic Nose With Deep Learning: A Multicenter Prospective Study,” Respiratory Research 25 (2024): 203, 10.1186/s12931-024-02840-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Steenhuis E. G. M., Asmara O. D., Kort S., et al., “The Electronic Nose in Lung Cancer Diagnostics: A Systematic Review and Meta‐Analysis,” ERJ Open Research 11, no. 3 (2025): 00723‐2024, 10.1183/23120541.00723-2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Lee B., Lee J., Lee J.‐O., et al., “Breath Analysis System With Convolutional Neural Network (CNN) for Early Detection of Lung Cancer,” Sensors and Actuators B: Chemical 409 (2024): 135578, 10.1016/j.snb.2024.135578. [DOI] [Google Scholar]
  • 10. Binson V. A., Subramoniam M., and Mathew L., “Detection of COPD and Lung Cancer With Electronic Nose Using Ensemble Learning Methods,” Clinica Chimica Acta 523 (2021): 231–238, 10.1016/j.cca.2021.10.005. [DOI] [PubMed] [Google Scholar]
  • 11. Sun F., Sun R., and Yan J., “Cross‐Domain Active Learning for Electronic Nose Drift Compensation,” Micromachines 13, no. 8 (2022): 1260, 10.3390/mi13081260. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Tsou P.‐H., Lin Z.‐L., Pan Y.‐C., et al., “Exploring Volatile Organic Compounds in Breath for High‐Accuracy Prediction of Lung Cancer,” Cancers 13, no. 6 (2021): 1431, 10.3390/cancers13061431. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Yang C.‐Y., Lee M.‐R., and Yang P.‐C., “Beyond Low‐Dose Computed Tomography: Emerging Diagnostic Tools for Early Lung Cancer Detection,” Journal of Thoracic Oncology 19, no. 9 (2024): 1261–1264, 10.1016/j.jtho.2024.06.013. [DOI] [PubMed] [Google Scholar]
  • 14. Qu C., Liu C., Gu Y., Chai S., Feng C., and Chen B., “Open‐Set Gas Recognition: A Case‐Study Based on an Electronic Nose Dataset,” Sensors and Actuators B: Chemical 360 (2022): 131652, 10.1016/j.snb.2022.131652. [DOI] [Google Scholar]
  • 15. Brown T. B., Mann B., Ryder N., et al., “Language Models Are Few‐Shot Learners,” [Preprint], arXiv (2020), 10.48550/arXiv.2005.14165. [DOI]
  • 16. Lu K., Grover A., Abbeel P., and Mordatch I., “Frozen Pretrained Transformers as Universal Computation Engines,” Proceedings of the AAAI Conference on Artificial Intelligence 36, no. 7 (2022): 7628–7636, 10.1609/aaai.v36i7.20729. [DOI] [Google Scholar]
  • 17. Zhou T., Niu P., Wang X., Sun L., and Jin R., “One Fits All: Power General Time Series Analysis by Pretrained LM,” Advances in Neural Information Processing Systems 36, no. 1877 (2023): 43322–43355, 10.5555/3666122.3667999. [DOI] [Google Scholar]
  • 18. Gruver N., Finzi M., Qiu S., and Wilson A. G., “Large Language Models Are Zero‐Shot Time Series Forecasters,” Advances in Neural Information Processing Systems 36, no. 861 (2023): 19622–19635, 10.5555/3666122.3666983. [DOI] [Google Scholar]
  • 19. Detterbeck F. C., Woodard G. A., Bader A. S., et al., “The Proposed Ninth Edition TNM Classification of Lung Cancer,” Chest 166, no. 4 (2024): 882–895, 10.1016/j.chest.2024.05.026. [DOI] [PubMed] [Google Scholar]
  • 20. Lee M.‐R., Huang H.‐L., Huang W.‐C., et al., “Electronic Nose in Differentiating and Ascertaining Clinical Status Among Patients With Pulmonary Nontuberculous Mycobacteria: A Prospective Multicenter Study,” Journal of Infection 87, no. 3 (2023): 255–258, 10.1016/j.jinf.2023.06.012. [DOI] [PubMed] [Google Scholar]
  • 21. Krizhevsky A., Nair V., and Hinton G., “CIFAR‐10 and CIFAR‐100 Datasets,” University of Toronto (2009), https://www.cs.toronto.edu/~kriz/cifar.html.
  • 22. Okamura K., Takayama K., Izumi M., Harada T., Furuyama K., and Nakanishi Y., “Diagnostic Value of CEA and CYFRA 21‐1 Tumor Markers in Primary Lung Cancer,” Lung Cancer 80, no. 1 (2013): 45–49, 10.1016/j.lungcan.2013.01.002. [DOI] [PubMed] [Google Scholar]
  • 23. Ren F., Fei Q., Qiu K., Zhang Y., Zhang H., and Sun L., “Liquid Biopsy Techniques and Lung Cancer: Diagnosis, Monitoring and Evaluation,” Journal of Experimental & Clinical Cancer Research 43 (2024): 96, 10.1186/s13046-024-03026-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Yang C.‐Y., Lin Y.‐T., Lin L.‐J., et al., “Stage Shift Improves Lung Cancer Survival: Real‐World Evidence,” Journal of Thoracic Oncology 18, no. 1 (2023): 47–56, 10.1016/j.jtho.2022.09.005. [DOI] [PubMed] [Google Scholar]
  • 25. National Lung Screening Trial Research Team , Church T.‐R., Black W.‐C., Aberle D.‐R., et al., “Results of Initial Low‐Dose Computed Tomographic Screening for Lung Cancer,” New England Journal of Medicine 368 (2013): 1980–1991, 10.1056/NEJMoa1209120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Jonas D.‐E., Reuland D.‐S., Reddy S.‐M., et al., “Screening for Lung Cancer With Low‐Dose Computed Tomography: An Evidence Review for the U.S. Preventive Services Task Force. [Internet],” Rockville (MD): Agency for Healthcare Research and Quality (US) Report No.: 20‐05266‐EF‐1 (2021). [PubMed]
  • 27. Dhanush Gowda A. M., Dessai A. D., and Nayak U. Y., “Electronic‐Nose Technology for Lung Cancer Detection: A Non‐Invasive Diagnostic Revolution,” Lung 203 (2025): 76, 10.1007/s00408-025-00828-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. van Riswijk M. L. M., van Tintelen B. F. M., Lucas R. H., van der Palen J., and Siersema P. D., “Overcoming Methodological Barriers in Electronic Nose Clinical Studies, a Simulation Data‐Based Approach,” Journal of Breath Research 19, no. 3 (2025): 036006, 10.1088/1752-7163/add291. [DOI] [PubMed] [Google Scholar]
  • 29. Zhao Z., Shen C., Tong H., et al., “From Images to Signals: Are Large Vision Models Useful for Time Series Analysis?” [Preprint], arXiv (2025), 10.48550/arXiv.2505.24030. [DOI]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S1: Model Structural of Large Language Model.

Figure S2: Model Structural of Convolutional Neural Network.

Figure S3: Study flow and design.

Table S1: Lung Cancer Demographic data.

Table S2: Few‐shot Performance of Tested Models at Cross‐Site.

Table S3: Subgroup analysis by lung cancer histology and staging.

RESP-31-912-s001.docx (473.3KB, docx)

Data Availability Statement

Data are available upon reasonable request from the corresponding author.


Articles from Respirology (Carlton, Vic.) are provided here courtesy of Wiley

RESOURCES