Abstract
Recent advances in cough sound analysis using deep learning techniques enable smartphone-based respiratory disease screening suitable for self-management care in a home setting, yet their utility is limited by device heterogeneity, population diversity, and challenges in multimodal integration. We propose a device-invariant, multimodal deep learning framework that jointly models cough acoustics, demographic data, and symptom descriptions for multi-label classification of adult respiratory diseases. To address the issues of device effect, an adversarial branch is embedded in the audio encoder to enforce device-invariant feature learning, while an invariant risk minimization-augmented loss enhances robustness to non-structural shifts. To evaluate the effectiveness of our proposed method, a real-world, multi-center dataset containing over 10,000 cases spanning seven major respiratory conditions was curated. On the tasks of individual respiratory disease identification for chronic obstructive pulmonary disease (COPD), lower respiratory tract infection (LRTI) and pulmonary shadows (PS), our method achieves superior performance with the area under the receiver operating characteristic curve (AUROC) of 0.9698, 0.8483 and 0.8720, respectively. It also shows promising results in identifying the presence of comorbidities for 7 respiratory diseases with an overall AUROC of 0.8907. More importantly, extensive experimental results demonstrate our method mitigates the issues of device effect and facilitates the cross-device generalization for cough-based respiratory disease diagnoses. This work demonstrates a scalable and transferable AI-based approach for cough-driven respiratory screening, emphasizing the importance of multimodal fusion and robust representation learning in advancing clinical applicability.
Subject terms: Computational biology and bioinformatics, Diseases, Engineering, Health care, Mathematics and computing
Introduction
Respiratory diseases represent a major global public health concern, contributing significantly to morbidity and mortality, particularly among adults and the elderly1. Conditions such as chronic obstructive pulmonary disease (COPD), interstitial lung disease (ILD), chronic bronchitis (CB), and various infectious diseases (e.g., lower respiratory tract infection, LRTI) often present with non-specific early symptoms—such as cough, sputum production, and shortness of breath—with coughs being one of the most common and earliest clinical indicators. Traditional methods for respiratory disease screening typically rely on chest imaging, pulmonary function tests2, or auscultation by trained clinicians. These approaches, however, often require specialized equipment and personnel3,4, making them difficult to scale in community settings or resource-limited environments5. In recent years, advances in artificial intelligence (AI) and acoustic modeling have facilitated the development of automated cough-based diagnostic tools. These systems offer a non-invasive, low-cost, and accessible alternative with the potential for remote deployment, positioning them as a promising solution for early respiratory disease screening and intelligent diagnostic support6,7.
Despite encouraging progress in cough sound analysis8,9, several limitations still hinder the real-world deployment of such AI models: first, distributional shifts in audio data caused by device variability significantly compromise model stability and generalization10. In the home setting, cough sounds may be recorded using a wide range of devices that differ in brand, model, and microphone placement11. Without explicit mechanisms for device-invariant learning, models often suffer performance degradation or even complete failure when deployed on unseen hardware. Second, most existing approaches rely solely on unimodal audio features, neglecting the rich diagnostic information embedded in patient demographics (e.g., age, sex, body composition, etc.) and symptom descriptions (e.g., sputum production, fever, dyspnea, etc.). This limited scope restricts the model’s ability to develop a holistic understanding of disease presentations. Furthermore, in real-world clinical settings, patients frequently present with multiple co-occurring respiratory conditions. However, most prior models are designed for single-label classification, making them ill-suited to capture the complex, multi-label pathologies seen in clinical practice12.
To address the aforementioned challenges, we propose a device-invariant, multi-modal deep learning framework for the identification and classification of adult respiratory diseases. Our approach integrates cough audio, demographic data, and symptom descriptions to systematically exploit the complementary information across modalities. We introduce an adversarial training mechanism into the audio encoder13, employing a gradient reversal strategy to counteract a device classifier and learn representations invariant to recording device differences. To further enhance robustness to distributional shifts between training and deployment environments, we incorporate invariant risk minimization (IRM)14 into a joint loss function15. Moreover, we design a unified multi-label learning framework capable of recognizing multiple respiratory diseases simultaneously, reflecting the real-world prevalence of comorbid conditions. Finally, we conduct a comprehensive evaluation on a multi-center real-world dataset comprising over 10,000 adult patient samples, demonstrating the model’s superiority in terms of accuracy, robustness, and generalization.
The proposed framework bridges the gap between experimental research and real-world clinical deployment of cough-based AI systems. It shows strong potential for application in frontline healthcare, remote health monitoring, and chronic disease screening among elderly populations.
Results
We conducted a comprehensive evaluation of the proposed multimodal deep learning framework for respiratory disease classification, with a particular focus on its effectiveness in multimodal integration, robustness to device-related biases, and adaptability to real-world deployment scenarios.
Data collection
We established a large-scale, multicenter cohort comprising 12,378 adult outpatients (≥18 years) recruited from respiratory clinics at four independent clinical centers. Written consent was obtained from all participants. The study was conducted in accordance with relevant ethical guidelines and approved by the ethics committees of Ruijin Hospital Shanghai Jiaotong University School of Medicine Ethics Committee (No. 2023199), Ruijin Hospital Luwan Branch Ethics Committee (2023HXK-V1), Shanghai Jing’an District Central Hospital Ethics Committee (No. 2023-33), and Shanghai Zhabei Central Hospital Ethics Committee (ZBLL2024030401001). The trial is registered at ClinicalTrials.gov (NCT06082791).
For each participant, at least 10 s of voluntary cough audio were recorded in a relatively quiet environment. All recordings underwent an internally validated quality control (QC) pipeline comprising event segmentation, cough detection, and validity assessment (as detailed in the Cough Sound Quality Control Process section). From each recording, multiple standardized 3-s cough segments were extracted, anchored at the cough onset and spanning 1 s before and 2 s after the burst, with recording device metadata documented for each segment.
To capture a comprehensive clinical context, participants completed structured questionnaires encompassing three domains: demographic information (e.g., sex, age, height, weight), smoking history (current and former status), and respiratory symptom profiles, including cough frequency and duration, sputum characteristics, dyspnea, and other relevant signs. Detailed information on subject questionnaires is provided in Supplementary Fig. 1. Disease labels were assigned based on final diagnoses by attending physicians and further verified through dual review by two independent senior pulmonologists, with multiple rounds of QC to ensure annotation accuracy. Detailed definitions and diagnostic criteria for the seven disease categories are provided in Supplementary Table 1, and the pulmonary shadows (PS) label is further described in Supplementary Table 2.
Partitioning the collected data in different ways resulted in several datasets designed to support distinct evaluation tasks. For the key binary classification analyses, three representative conditions were selected to enable independent positive-negative recognition: COPD (representing chronic airway disease, Fig. 1a), LRTI (representing infectious disease, Fig. 1b), and PS (representing potential tumors or structural abnormalities, Fig. 1c). For the COPD and LRTI tasks, we conducted five-fold cross-validation. Recordings originating from the same device and clinical center were confined to a single fold to strictly prevent information leakage across folds. All three binary classification tasks exhibited considerable class imbalance. In the COPD task, the dataset comprised 813 positive cases (COPD) and 4933 negative cases (Other), with positives accounting for 14.1% of the total. In the LRTI task, 720 subjects were labeled as positive (LRTI) and 4894 as negative (Other), corresponding to a positive proportion of 12.8%. The PS task exhibited the greatest imbalance, with only 432 positive cases (PS) compared to 6379 negatives (Other), corresponding to a positive rate of 6.3%. These distributions indicate that the PS classification problem poses the most challenging imbalance, while COPD and LRTI tasks remain moderately skewed toward negative samples.
Fig. 1. Data distribution across three binary classification tasks.
a COPD vs. other. (1) Disease label distribution shows that COPD cases account for 14.1% (n = 813) of the cohort; (2) Age distribution by label demonstrates that COPD patients are generally older than the non-COPD group; (3) Scatter plot of BMI versus age, colored by label, indicates overlapping but distinguishable patterns between COPD and other cases; (4) Device usage by label illustrates heterogeneous recording sources, with iPhone and external microphones as the most frequently used devices; (5) Smoking history, stacked with counts, reveals a higher proportion of smokers among COPD patients; (6) Missing data rate per feature highlights device type and smoking history as the most incomplete variables. b LRTI vs. other. (1) Disease label distribution indicates that LRTI cases account for 12.8% (n = 720); (2) Age distribution by label suggests that LRTI cases are more common among younger participants compared to non-LRTI individuals; (3) Gender distribution shows a slightly higher prevalence of LRTI among males; (4) Device usage by label demonstrates similar heterogeneity as in COPD, with substantial representation from iPhone devices and external microphones; (5) Smoking history, stacked with counts, shows a higher proportion of smokers in the LRTI group than in the other group; (6) Missing data rate per feature again emphasizes device type and smoking history as primary sources of missingness. c PS vs. other. (1) Disease label distribution indicates that PS cases account for 6.3% (n = 432); (2) Age distribution by label shows PS cases to be concentrated in middle-to-older age ranges compared to non-PS individuals; (3) Gender distribution highlights a slight male predominance among PS cases; (4) Device usage by label shows broad representation across devices, with iPhone and external microphones being the most common; (5) Smoking history, stacked with counts, reveals relatively balanced proportions of smokers and non-smokers in the PS group; (6) Missing data rate per feature again identifies device type and smoking history as the most incomplete variables.
In parallel, a multi-label respiratory disease recognition task was constructed to capture disease co-occurrence, allowing each subject to be annotated with the presence (1) or absence (0) of 7 conditions (i.e., COPD, Asthma, ILD, upper respiratory tract infection (URTI), LRTI, CB and Bronchiectasis). Models were required to estimate the probability of presence for each disease, with task-specific data distributions presented in Fig. 2, with asthma (10.3%) and COPD (9.4%) being most prevalent, followed by LRTI (7.8%) and multi-disease cases (4.1%). Recordings were obtained across heterogeneous devices, with iPhone series and external microphones dominating. The cohort spanned a wide age range (Fig. 2, panel 4) with expected gender differences in height, weight, and smoking history, the latter showing markedly higher prevalence among males.
Fig. 2. Data distribution of the multi-label respiratory disease classification task.
(1) Disease label distribution: Pie chart showing the proportion of different respiratory diseases, with the majority labeled as OTHER (63.21%), followed by ASTHMA (10.29%), COPD (9.36%), LRTI (7.76%), Multi-disease (4.11%), URTI (1.78%), CB (1.55%), BRONCHIECTASIS (0.95%), and ILD (0.95%); (2) Device-disease heatmap: Heatmap illustrating the cross-distribution of recording devices (e.g., iPhone models, Huawei, OnePlus, external microphone, etc.) and annotated disease labels. Notably, recordings from “Unknown” devices and iPhone 11 Series dominate across multiple disease categories; (3) Smoking history by gender: Stacked bar plot comparing smoking history distributions between Female (n = 3892), Male (n = 2908), and Unknown (n = 15) groups. Categories include Smoking, Clean, Never, and Unknown, with females showing a higher proportion of Clean history and males higher in Smoking and Never; (4) Age distribution: Histogram showing the age spread of participants stratified by gender. Both males and females are broadly distributed across ages 18–90, with peaks around 40 and 65 years; (5) Height distribution: Histogram of participants' height (cm), stratified by gender. Females cluster around 155–165 cm, while males center around 165–175 cm; (6) Weight distribution: Histogram of participants' weight (kg), stratified by gender. Females are concentrated in the 45–70 kg range, while males center around 60–85 kg.
Binary classification results for chronic obstructive pulmonary disease
For the binary classification task of COPD, we employed a five-fold cross-validation strategy for performance evaluation. Table 1 summarizes the mean and standard deviation of area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve (AUPRC), and reports the mean and standard deviation of the F1-score, precision, sensitivity, specificity, and overall accuracy calculated under two thresholding strategies: a manually defined threshold and the Youden Index method. The AUROC, AUPRC, and their corresponding confidence intervals (AUROC_CI) for each fold, as well as detailed metric reports under different thresholds, are provided in Supplementary Tables 3 and 4. The results indicate that our proposed model achieved mean AUROC values exceeding 0.96 on both the validation and test sets, with standard deviations below 0.007, demonstrating high stability and superior discriminative capability for this task.
Table 1.
Indicators of the COPD binary classification task
| Validation | Test | |||
|---|---|---|---|---|
| Mean | Std | Mean | Std | |
| AUROC | 0.9674 | 0.0068 | 0.9698 | 0.0066 |
| AUPRC | 0.8785 | 0.0266 | 0.8851 | 0.0130 |
| F1a | 0.8146 | 0.0100 | 0.8228 | 0.0244 |
| Precisiona | 0.7779 | 0.0389 | 0.7882 | 0.0426 |
| Sensitivitya | 0.8572 | 0.0278 | 0.8623 | 0.0290 |
| Specificitya | 0.961 | 0.0095 | 0.9629 | 0.0104 |
| Accuracya | 0.9468 | 0.0051 | 0.9491 | 0.0092 |
| F1b | 0.8015 | 0.0204 | 0.7918 | 0.0472 |
| Precisionb | 0.7269 | 0.0394 | 0.7021 | 0.0773 |
| Sensitivityb | 0.8951 | 0.0194 | 0.9143 | 0.0265 |
| Specificityb | 0.9463 | 0.0127 | 0.936 | 0.0266 |
| Accuracyb | 0.9393 | 0.0094 | 0.9331 | 0.0216 |
For detailed results of each fold, please refer to Supplementary Tables 3 and 4.
aCut off = 0.15. Manual cutoff thresholds were empirically determined based on validation set distributions to balance sensitivity and specificity for clinically relevant detection.
bCut off value selected by Youden Index.
Binary classification results for lower respiratory tract infection
For the LRTI binary classification task, we employed the same five-fold cross-validation protocol used in the COPD experiments. Table 2 presents the primary performance metrics, while the complete set of results is provided in Supplementary Tables 5 and 6. Despite the etiological complexity and substantial phenotypic heterogeneity of LRTI in clinical practice, our model consistently achieved mean AUROC values above 0.84 on both the validation and test sets, with standard deviations below 0.03, highlighting its robustness in handling clinically complex and heterogeneous disease presentations.
Table 2.
Indicators of the LRTI binary classification task
| Validation | Test | |||
|---|---|---|---|---|
| Mean | Std | Mean | Std | |
| AUROC | 0.8491 | 0.0293 | 0.8483 | 0.0198 |
| AUPRC | 0.6084 | 0.0484 | 0.5858 | 0.0339 |
| F1a | 0.6245 | 0.0177 | 0.5883 | 0.0352 |
| Precisiona | 0.6630 | 0.0593 | 0.5757 | 0.0979 |
| Sensitivitya | 0.5969 | 0.0553 | 0.6218 | 0.0736 |
| Specificitya | 0.9538 | 0.0144 | 0.9259 | 0.0367 |
| Accuracya | 0.9078 | 0.0056 | 0.8868 | 0.0252 |
| F1b | 0.5826 | 0.0464 | 0.5465 | 0.0512 |
| Precisionb | 0.5013 | 0.0899 | 0.443 | 0.0744 |
| Sensitivityb | 0.7181 | 0.0713 | 0.7293 | 0.0499 |
| Specificityb | 0.8872 | 0.0441 | 0.8575 | 0.048 |
| Accuracyb | 0.8654 | 0.0309 | 0.8411 | 0.0379 |
For detailed results of each fold, please refer to Supplementary Tables 5 and 6.
aCut off = 0.10. Manual cutoff thresholds were empirically determined based on validation set distributions to balance sensitivity and specificity for clinically relevant detection.
bCut off value selected by Youden Index.
Binary classification results for pulmonary shadows
To evaluate the potential of our model for early screening of pulmonary tumors, we defined the inclusion criteria for PS in consultation with clinical experts (see Supplementary Table 2) and excluded cases attributable to infectious etiologies. Given the relatively limited number of such cases, we randomly partitioned the dataset into training, validation, and test sets at a ratio of 3:1:1, ensuring comparable demographic distributions across splits. The evaluation metrics used were consistent with those applied in the COPD and LRTI experiments (Table 3).
Table 3.
Indicators of the PS binary classification task
| Type | Val | Test |
|---|---|---|
| AUROC | 0.8704 | 0.8720 |
| AUROC 95% CI | [0.8110, 0.9178] | [0.8066, 0.9241] |
| AUPRC | 0.4427 | 0.4785 |
| F1a | 0.3556 | 0.3418 |
| Precisiona | 0.2353 | 0.2207 |
| Sensitivitya | 0.7273 | 0.7581 |
| Specificitya | 0.8549 | 0.8207 |
| Accuracya | 0.8475 | 0.8168 |
| F1b | 0.3605 | 0.3125 |
| Precisionb | 0.2360 | 0.1897 |
| Sensitivityb | 0.7636 | 0.8871 |
| Specificityb | 0.8482 | 0.7462 |
| Accuracyb | 0.8433 | 0.7551 |
aCut off = 1 × 10−5. Manual cutoff thresholds were empirically determined based on validation set distributions to balance sensitivity and specificity for clinically relevant detection.
bCut off value selected by Youden Index.
The model achieved AUROC values above 0.87 across validation and test sets, indicating robust ranking capability despite the extreme class imbalance (positive rate 6.3%). However, the F1-score (val: 0.3556; test: 0.3418), precision (val: 0.2353; test: 0.2207), and an AUPRC of 0.4785—only moderately above the prevalence-based baseline—highlight the limited precision attainable for this rare-event outcome.
This discrepancy between AUROC and AUPRC reflects the heterogeneous and frequently clinically silent nature of PS, which ranges from incidental ground-glass nodules to more advanced tumors. Early-stage or peripheral lesions typically produce minimal or non-specific symptoms, whereas the model relies primarily on demographics, symptom questionnaires, and voluntary cough acoustics—features that are more indicative of obstructive airway disease or diffuse parenchymal involvement than of small or asymptomatic nodules. In view of these characteristics and the low prevalence of PS, we do not position the model as a stand-alone screening tool for pulmonary nodules or primary lung cancer. Instead, PS scores should be interpreted as a triage or risk-stratification aid to help prioritize individuals for confirmatory imaging, while a low predicted probability should not be used to defer further diagnostic evaluation.
Performance on multilabel respiratory disease classification
In the multilabel classification task involving seven distinct respiratory diseases, our proposed framework was benchmarked against a range of established models (Fig. 3a). Leveraging a large-scale, multi-center, multi-device dataset, the model achieved the highest performance in both AUROC (Fig. 3b) and AUPRC (Fig. 3c) metrics. These results significantly outperform baseline approaches, underscoring the framework’s strong generalization ability and its potential for deployment in complex clinical environments involving heterogeneous data modalities (Table 4).
Fig. 3. Comparison of different models in multi-label tasks.
a Radar chart of AUC results for 7 diseases using different models. b Bar chart of AUROC results for different models. c Bar chart of AUPRC results for different models.
Table 4.
AUC results of different models for each label in multi-label tasks
| Disease | EAT38 | BTS39 | FrameATST40 + Roberta26a | BEATs41 + Roberta26a | Ours |
|---|---|---|---|---|---|
| URTI | 0.7222 | 0.7507 | 0.8047 | 0.8392 | 0.8432 |
| LRTI | 0.7877 | 0.7903 | 0.8532 | 0.8140 | 0.8316 |
| ILD | 0.7485 | 0.7054 | 0.6916 | 0.8011 | 0.7732 |
| COPD | 0.9195 | 0.9288 | 0.9382 | 0.9299 | 0.9430 |
| CB | 0.7264 | 0.7633 | 0.8123 | 0.7878 | 0.8641 |
| BRONCHb | 0.7778 | 0.8090 | 0.7043 | 0.7757 | 0.8005 |
| ASTHMA | 0.7570 | 0.7687 | 0.7824 | 0.7726 | 0.7993 |
| ALLc | 0.7770 | 0.7880 | 0.7981 | 0.8172 | 0.8364 |
| ALLd,e | 0.8346 | 0.8463 | 0.8667 | 0.8712 | 0.8907 |
aThe same architecture as Fig. 1a (without adding the device adversarial module) is used.
bBRONCH is an abbreviation for BRONCHIECTASIS.
cROC AUC Macro of all labels.
dROC AUC Micro of all labels.
eAn additional stratified multi-label evaluation by introducing background subgroups can be seen in Supplementary Table 7.
Ablation studies and the impact of adversarial training
We conducted an ablation study on the LRTI binary classification task to systematically evaluate the contribution of different modalities and their combinations to disease prediction (Fig. 4). We assessed the independent and combined predictive utility of audio recordings, demographic information, and symptom descriptions. The results revealed substantial performance discrepancies across single modalities: models relying solely on demographic and symptom information (text modality) achieved the lowest AUROC (0.6392), whereas the use of audio alone markedly improved performance (AUROC = 0.7914). Dual-modality integration further enhanced classification accuracy, with audio combined with demographics (AUROC = 0.7966) or with symptoms (AUROC = 0.8061), both outperforming single-modality models, highlighting the complementary value of cross-modal information. Incorporating all three modalities—audio, demographics, and symptoms—yielded an additional performance gain (AUROC = 0.8393). This shows that audio-only models demonstrated strong baseline performance, while demographic and symptom features provided complementary benefits, particularly in multi-label recognition tasks. Notably, the inclusion of adversarial training in this multi-modal setting achieved the best performance (AUROC = 0.8571) and exhibiting superior generalizability and clinical utility, underscoring its effectiveness in mitigating device-related bias and improving model generalizability. Collectively, these findings emphasize the critical role of multimodal integration and adversarial optimization in advancing disease classification, and the synergistic value of multimodal medical data in respiratory disease diagnosis.
Fig. 4. Modality ablation study.
Performance comparison of the multi-modal model when selectively removing individual modalities. Each bar chart corresponds to one removal configuration and a corresponding ROC curve. The full-modal model achieved the highest diagnostic performance, indicating that the complementarity of cough sounds, demographic information, and symptom information collectively enhances the learning of disease-related features. Removing the audio modality resulted in the most significant performance degradation, highlighting its dominant role in capturing respiratory phenotypes, while demographic and symptom information provided additional, non-redundant contributions.
Additionally, to assess the effectiveness of adversarial training in mitigating device-related biases, we systematically compared model variants trained with and without adversarial mechanisms across all classification tasks (Fig. 5). Consistent improvements in AUC were observed with adversarial training, indicating its efficacy in suppressing device-induced variability and enhancing model robustness in heterogeneous deployment settings.
Fig. 5. Comparison of adversarial and non-adversarial modules in different tasks.
Performance comparison across multiple downstream tasks when integrating a device-invariant adversarial module. The adversarial version consistently enhances model generalization and reduces device-induced bias.
Adversarial training enables device-invariant feature learning
To visualize the impact of adversarial training on learned audio representations, we extracted CLS token features from the audio encoder16 and projected them into a 2D space using t-SNE17 (Fig. 6). Without adversarial training, samples from different recording devices exhibited distinct clustering patterns, revealing clear device bias. In contrast, models trained with the adversarial mechanism produced tighter clusters by disease category and exhibited markedly reduced device-based separation. These findings suggest that our approach enables the learning of device-invariant, yet discriminative, audio embeddings.
Fig. 6. Feature space distribution of disease labels and device labels in adversarial/non-adversarial training tasks.
t-SNE projections of latent embeddings trained with (top) and without (bottom) the device adversarial module. Left panels: embeddings colored by disease label (COPD vs. Other). Right panels: embeddings colored by recording device (e.g., iPhone series, Huawei P30, Samsung Galaxy S10). Without adversarial training, device-specific clusters dominate the feature space, indicating strong hardware-driven confounding. Incorporating adversarial learning effectively collapses device clusters while preserving disease-relevant structure, demonstrating successful removal of device biases during representation learning.
The adversarial model demonstrates superior cross-device generalizability
To further assess the adaptability of our framework to deployment across heterogeneous devices in real-world scenarios, we retrained both the adversarial and non-adversarial models on the multi-label dataset after excluding all samples collected from “Unknown” devices. We then conducted evaluation on an independent external test cohort consisting of cough recordings acquired from multiple smartphone models that were never used during training or validation (Fig. 7).
Fig. 7. Disease and label statistics in adaptive experiments for cross-device deployment.
Distribution of diseases and corresponding device used in external testing experiments.
The results revealed that the non-adversarial model exhibited a marked decline in classification performance on unseen devices, with AUROC notably lower than that observed on the training devices, suggesting that its learned representations were device-dependent (Table 5). In contrast, the adversarially trained model maintained stable performance under the same testing conditions, with substantially attenuated AUROC degradation.
Table 5.
Comparison of results between adversarial and non-adversarial models on external testsets
| ROC AUC Micro | ROC AUC Macro | |||||
|---|---|---|---|---|---|---|
| Internal Validation | External Test | Δ | Internal Validation | External Test | Δ | |
| Adv | 0.9189 | 0.9056 | 0.0133 | 0.8917 | 0.8903 | 0.0014 |
| None-Adv | 0.8862 | 0.8684 | 0.0178 | 0.8322 | 0.7649 | 0.0673 |
The training and validation datasets used for model comparison were derived from the multi-label task dataset after excluding samples labeled as “Unknown" devices. The external test set consisted of a fully independent collection of samples acquired from multiple smartphone models that were not present in either the training or validation sets.
These findings highlight that the device-adversarial module effectively mitigates distributional shifts induced by variations in audio acquisition hardware, thereby preserving discriminative power and robustness in cross-device deployment. This property is critical for ensuring reliable application of the proposed framework in multi-center, cross-regional, and resource-limited healthcare settings.
Invariant risk minimization enhances device robustness
To further address distributional shifts introduced by diverse recording hardware, we incorporated both device-adversarial learning and IRM strategies into our framework. In cross-device transfer tests, models without adversarial components experienced substantial performance degradation. However, models employing the combined adversarial + IRM strategy demonstrated significantly improved classification accuracy on unseen target devices, evidencing enhanced invariance and stability across deployment conditions (Table 6). Although the absolute performance gain in terms of AUC under identical experimental settings is numerically modest, repeated experiments with different random seeds on a relatively small COPD dataset revealed a consistent improvement trend (8 out of 10 cases). Specifically, the adversarial model augmented with IRM outperformed its non-IRM counterpart in the majority of runs, while exhibiting reduced performance variance, indicating improved training stability rather than stochastic fluctuation. Due to the limited dataset size, we did not perform formal statistical significance testing; nevertheless, the observed consistency across repeated trials suggests that the performance gain is systematic rather than incidental.
Table 6.
Comparisons of adversarial with IRM, adversarial without IRM, and non-adversarial tasks
| Task Name | Loss function | Average training duration per epoch / H | Training time to achieve the best AUC result / H | AUC |
|---|---|---|---|---|
| Adv with IRM | Fig. 10a (1) | 1.03 | 53.56 (52 epochs) | 0.8571 |
| Adv without IRM | Fig. 10a (2) | 0.52 | 49.40 (95 epochs) | 0.8529 |
| Non-Adv | Fig. 10a (3) | 0.48 | 44.64 (93 epochs) | 0.8393 |
Timeliness comparison based on the same batch size and computing resources.
From an optimization perspective, incorporating IRM into the device-adversarial framework increases the computational cost per iteration. However, this is offset by substantially faster convergence, as the adversarial + IRM model reaches its optimal performance with significantly fewer training epochs compared to the adversarial-only baseline. This improved convergence behavior further supports the role of IRM in stabilizing training dynamics under distributional shifts. Collectively, these findings indicate that the combined adversarial + IRM strategy offers a favorable trade-off between computational overhead and robustness, and holds strong potential for enabling reliable and scalable deployment of intelligent respiratory disease recognition models in multi-center clinical environments.
Tradeoff between model performance and efficiency
To investigate the trade-off between model complexity and deployment efficiency, we compared two versions of our framework: a base-size model and a large-size model with approximately double the parameter count (Table 7). While the large-size variant achieved a higher mean AUROC of 0.8571 in the multilabel task—slightly outperforming the base-size model (AUROC = 0.8405)—the lightweight base-size version maintained competitive accuracy while offering substantial advantages in inference latency and memory usage. These attributes make it particularly well-suited for deployment in resource-constrained environments such as mobile devices, edge hardware, or primary care settings.
Table 7.
Comparison of model parameters
| Model size level | Parameters of model components | Total parameters | AUC | |||
|---|---|---|---|---|---|---|
| Encoder Module | Fusion + Disease Classifier Module | Device Adversarial Module | Uncertainty Weighting (Determined by the gradient) | |||
| Base | 215M | 10.6M | 1.2M | 3 | 226.2M | 0.8405 |
| Large | 440M | 18.6M | 1.2M | 3 | 459.8M | 0.8571 |
This comparison highlights the critical importance of model compression and efficient deployment strategies in real-world medical AI applications, offering valuable insights for the future design of edge-intelligent diagnostic systems.
Unless otherwise specified, all experiments were conducted using the large-sized model (460M parameters) in this study. The base variant (226M parameters) was evaluated only for ablation and scalability analyses.
Discussion
Clinical significance and deployment feasibility
Cough is a hallmark symptom across a wide spectrum of respiratory diseases, and its acoustic characteristics carry rich pathological information. Compared with conventional diagnostic approaches relying on imaging, laboratory tests, or auscultation, cough-based acoustic analysis offers several compelling advantages: it is non-invasive, low-cost, easily scalable, and suitable for remote deployment. These properties make it particularly advantageous for primary care, telehealth platforms, and mobile health applications in resource-constrained settings18,19.
The proposed multimodal deep learning framework achieved high diagnostic accuracy across diverse respiratory conditions, even without relying on structured clinical inputs or specific device brands. The model demonstrated robust generalization across heterogeneous devices and multi-center data sources, confirming its potential for real-world clinical adaptation. These findings lay a strong foundation for its future deployment in community health centers, telemedicine platforms, and edge-intelligent medical devices.
Model scalability and cross-scenario adaptability
The architecture of the proposed framework is inherently modular and extensible, facilitating its adaptation to broader medical applications:
(1) Modal expansion: Beyond integrating audio, demographic features, and symptom questionnaire data, the model is readily extendable to incorporate additional clinical modalities such as chest X-rays, body temperature, heart rate, and peripheral oxygen saturation (SpO2), thereby enhancing both disease detection sensitivity and clinical interpretability.
(2) Task transferability: The audio encoder and device-adversarial modules can potentially be transferred to other healthcare-related acoustic domains, including cardiac auscultation analysis, snore detection, asthma exacerbation prediction, or screening for sleep-related breathing disorders.
(3) Acquisition flexibility: The model’s ability to handle device variability allows it to operate effectively across multiple audio input sources, including smartphones, remote microphones, and wearable sensors, thereby supporting deployment across a wide range of technical infrastructures and environmental conditions.
Looking ahead, we plan to leverage our ongoing longitudinal data collection efforts to integrate medical imaging and other data streams, further advancing the model’s diagnostic accuracy, cross-domain generalizability, and clinical transparency.
Limitations and future directions
Despite the promising results achieved across multiple diagnostic tasks, several limitations remain and warrant further investigation:
(1) Objectivity and consistency of diagnostic labeling: The current ground-truth labels are primarily derived from initial clinical diagnoses, which may be affected by physician experience, documentation quality, and institutional practices4. Such subjectivity may introduce potential bias and variability in label assignment. To mitigate this, future work will incorporate multi-round diagnostic consensus, cross-validation with imaging and laboratory findings, and long-term follow-up data to establish a more objective and standardized labeling framework.
(2) Class imbalance and rare disease detection: The dataset exhibits considerable label imbalance, with certain respiratory diseases—especially rare conditions or early-stage pulmonary lesions—being severely underrepresented. This limited data availability compromises model performance on these classes, as reflected in lower AUC and recall scores compared to more prevalent conditions. This issue is particularly pronounced in the multilabel setting. To address this, future work will incorporate strategies such as class reweighting, oversampling, synthetic data augmentation, and transfer learning to mitigate imbalance-related performance bottlenecks and improve small-sample generalization.
(3) Limited model interpretability: Although attention mechanisms and adversarial components enhance model transparency to some extent, the decision-making process remains largely opaque. Incorporating explainable AI techniques20 in future iterations will be critical for elucidating the model’s rationale in a clinically meaningful way.
(4) Lack of temporal disease modeling: The current framework is designed for static snapshot-based prediction and does not yet incorporate temporal patterns in disease progression. Future work will focus on longitudinal modeling and patient trajectory analysis to support disease monitoring and individualized risk prediction.
(5) Environmental constraints and noise robustness: All audio recordings were collected in relatively quiet clinical environments to ensure data consistency. However, this controlled setting may not fully reflect the acoustic variability encountered in real-world scenarios such as home or community settings. Moreover, each audio event was pre-segmented into 3-s clips to facilitate model input standardization, which assumes clean segmentation—a condition that may not always hold in uncontrolled environments. Future studies will expand data collection to diverse acoustic contexts and develop segmentation-free or noise-robust learning paradigms to improve model generalization under realistic conditions.
(6) Handling weakly labeled and noisy data: Incomplete or noisy annotations are inherent in large-scale clinical datasets. To address this, future work will explore robust and data-efficient learning strategies such as causal inference21, active learning22, and semi-supervised learning23, thereby improving adaptability under weak supervision and enhancing model reliability in real-world applications.
(7) Population homogeneity and ethnic generalizability: Although data were collected from multiple clinical sites, the enrolled population is relatively homogeneous in ethnicity and geographic distribution. Such population uniformity may limit the model’s generalizability to other ethnic groups, healthcare systems, and socio-environmental contexts. Future work will emphasize dataset diversification through international collaborations and multi-regional data acquisition, as well as the integration of domain generalization and fairness-aware learning strategies to enhance model robustness, equity, and clinical transferability across diverse populations.
(8) Clinical applicability in complex comorbid conditions: While the present work demonstrates promising generalization across devices and recording settings, its performance in patients with complex comorbidities remains to be further validated. Conditions such as heart failure, endocrine disorders, and gastroesophageal reflux disease (GERD) may produce cough sounds or respiratory manifestations that overlap with pulmonary pathologies, thereby challenging model specificity in differential diagnosis. Future studies will focus on stratified validation across such comorbid subgroups and integrate multimodal clinical information—including cardiovascular and metabolic parameters–to improve diagnostic discrimination and real-world clinical applicability, particularly in primary care and community health scenarios. Although the current study demonstrates promising cross-device generalization in clinical datasets, prospective evaluations in real-world settings—including home monitoring and primary care environments—are currently ongoing. Results from these studies will be reported in future work to further validate the model’s clinical utility, generalizability, and practical deployment.
Methods
Pipeline overview
The proposed framework comprises a complete end-to-end pipeline that integrates audio signal processing, text encoding, and multimodal learning to identify respiratory diseases from cough sounds and associated demographic information.
At the foundation of the pipeline lies a rigorous cough audio QC process that ensures the reliability of the raw recordings. Each audio sample undergoes segmentation, cough event detection, and cough validity verification to discriminate valid cough segments from background noise, whisper, throat clearance, and speech. The resulting high-quality cough clips form the basis for subsequent acoustic analysis.
Each validated cough segment is transformed into a two-dimensional mel-spectrogram representation, capturing both spectral and temporal variations relevant to disease-specific acoustic signatures. These representations are then processed by an audio encoder, based on the Efficient Audio Transformer (EAT) architecture, which generates compact and discriminative latent embeddings. In parallel, patient demographic data and textual attributes are encoded through a text encoder, built upon a pretrained Robustly Optimized BERT Pretraining Approach (RoBERTa) model fine-tuned for structured medical text.
The encoded audio and text embeddings are aligned and fused within the multimodal fusion classifier, which learns a shared latent space for cross-modal reasoning. This component integrates heterogeneous information to capture complex audio-text correlations that underpin clinical disease patterns. The final classifier makes a robust multimodal diagnostic decision.
To further enhance the model’s robustness by reducing the impact of device variance, the training process incorporates adversarial and IRM strategies, as illustrated in Figs. 9 and 10. The loss design enables the model to disentangle disease-related cues from device or environment-specific variations, thereby improving generalization across recording conditions and data sources.
Fig. 9. All model frameworks.
a Multimodal device adversarial model framework. b Single audio modality device adversarial model framework. c Single audio modality non-device adversarial model framework. d Single text modality model framework.
Fig. 10. Loss function calculation flow chart.
a (1) IRM-based adversarial strategies. (2) Adversarial strategies without IRM. (3) Non-adversarial strategies. b (1) Saturation dynamic adjustment. (2) Sigmoid dynamic adjustment.
Despite the relatively large total parameter count of the proposed multi-modal model, several deliberate design choices were adopted to ensure effective learning and to mitigate the risk of overfitting under the limited data regime. Extensive regularization strategies were employed throughout training. Specifically, dropout was applied across multiple layers, mixup-based data augmentation was used to improve robustness, and early stopping based on validation performance was adopted to prevent overfitting. And more importantly, the model heavily relies on pretrained initialization rather than learning from scratch. Both the audio encoder and the text encoder were initialized using publicly released pretrained weights provided by their original authors, which were trained on large-scale unimodal datasets. This pretrained initialization substantially reduces the effective degrees of freedom during optimization and enables stable convergence even with limited task-specific data.
Crucially, the entire pretrained model was not fully fine-tuned. Instead, a partial fine-tuning strategy was adopted to further control model capacity. In particular, the text encoder was entirely frozen during training, and no layers within the text encoder were updated. This design choice ensures that the rich semantic representations learned from large-scale corpora are preserved, while completely eliminating the risk of overfitting within the text branch. In contrast, the remaining components of the model, including the audio encoder and the downstream fusion and classification modules, were fine-tuned to adapt modality-specific and cross-modal representations to the classification tasks. By freezing the text encoder and restricting gradient updates to only a subset of the overall architecture, the number of trainable parameters was substantially reduced relative to the full model size.
Overall, this unified pipeline seamlessly bridges signal-level processing and high-level multimodal inference, ensuring both clinical scalability across diverse real-world settings.
Cough sound quality control process
The Fig. 8 illustrates a multi-level QC workflow for audio data applied during both training and validation/testing stages.
Fig. 8. Audio data quality control workflow.
a Quality control process for training. b Quality control process for validation, testing/inference.
During training, audio data first undergoes an environmental noise assessment (Front Noise Assessment Module). Audio segments are considered usable if the background noise is below 50 dB and at least 10 seconds of continuous cough activity is detected. Segments passing the noise filter are then processed through cough segmentation (as detailed in the Cough Segmentation section), which uses phase detection to identify the onset, peak, and intermediate phases of each cough event and to determine the split points used to extract individual cough segments. If valid segments are obtained, they are retained for model training; otherwise, the data is marked as failed QC.
For validation, testing, and inference, a more refined single-cough segment QC process is applied. Each cough segment is subsequently evaluated for validity using both a cough detection model and a cough effectiveness assessment model (as detailed in the Cough Filter section). Valid segments pass QC and are included in the analysis, whereas invalid segments are discarded. For multi-segment cough recordings, candidate segments are initially selected using the same environmental noise assessment and cough segmentation procedures. Each candidate segment then undergoes the single-cough segment QC process, with only valid segments retained for inference or analysis. Segments deemed invalid at any stage are discarded.
Overall, this workflow ensures that cough audio used in both training and inference is rigorously filtered for environmental noise, accurately segmented by cough phase, and verified for validity. This process enhances model robustness and reliability while minimizing the impact of background noise or invalid cough data on model training and inference.
Cough segmentation
We employed a multi-stage signal processing pipeline to segment cough phases from audio recordings, enabling the extraction of precise temporal markers for the onset, peak, and intermediate phases of each cough event (Algorithm 1).
First, the raw audio signal x(t) was loaded and normalized by dividing the waveform by its maximum absolute amplitude:
| 1 |
which scales the signal to the range [−1, 1]. Candidate cough onset locations were detected using a frequency-domain onset detection procedure. The audio waveform was first transformed into the time-frequency domain via a short-time Fourier transform (STFT). For each frame, the spectral energy ratio between consecutive time steps was computed, and abrupt increases in high-frequency energy were identified. Frames exhibiting energy ratio changes exceeding an adaptive threshold were marked as potential onsets, and temporally adjacent detections were merged to produce the final set of candidate cough onset locations. For each candidate onset, a 0.7 s segment surrounding the onset (0.2 s before and 0.5 s after) was extracted. Within each segment, the precise explosive phase start was determined using root-mean-square (RMS) amplitude analysis and thresholding. The end of the explosive phase was detected via a spectral comparison method, where consecutive frames were analyzed for frequency-domain changes, and the first significant drop in the frequency energy ratio indicated the phase termination. Subsequently, the intermediate phase end was identified by detecting valleys in the low-pass filtered RMS signal.
All detected markers were corrected for segment offsets and filtered to remove closely spaced duplicates, ensuring that consecutive coughs were separated by at least 0.1 s. The resulting set of temporal markers for each cough event is represented as
Finally, after obtaining the temporal markers of each cough event, the corresponding audio signals were segmented for subsequent analysis. Specifically, each segment was centered on the detected explosive start marker, extending 1 s before and 2 s after it, thereby capturing the complete acoustic dynamics surrounding the cough phase transition. This segmentation ensured that both the onset and decay characteristics of the cough were preserved for downstream feature extraction and model training.
This segmentation process enables high-fidelity delineation of cough phases, which is critical for downstream feature extraction and classification.
Algorithm 1
Cough phase segmentation and marker detection
Require: Audio x(t), sampling rate Fs, detection parameters θ
Ensure: Cough phase markers markers = [tstart, texplosive_end, tintermediate_end]
1: Normalize audio:
2: Preliminary cough onsets: candidate_onsets ← OnsetCough(x, Fs, θ)
3: Initialize marker list:
4: for each tcand in candidate_onsets do
5: segment ← x[tcand − 0.2s: tcand + 0.5s] ⊳ Segment around candidate
6: tstart ← RMS_Threshold_Search(segment)
7: [ratio, texplosive_end] ← Phase1EndDetect(segment, tstart)
8: tintermediate_end ← Phase2EndDetect(segment, tstart, texplosive_end)
9: Correct to original signal: tmarkers ← [tstart, texplosive_end, tintermediate_end] + tcand − 0.2s
10: if texplosive_end ≠ None then
11: markers_list ← markers_list ∪ tmarkers
12: end if
13: end for
14: Remove duplicates with minimum 0.1 s separation: markers_list ← RemoveDuplicates(markers_list, 0.1s)
15: return markers_list
Cough filter
After segmentation, each cough-centered audio clip underwent a standardized preprocessing procedure to suppress device-related noise and harmonic interference. To mitigate baseline drift introduced by external microphones, a fourth-order Butterworth high-pass filter with a cutoff frequency of 50 Hz was applied to remove low-frequency noise components. Subsequently, harmonic-percussive source separation (HPSS) was performed to further attenuate stationary harmonic noise while preserving the transient cough components. The resulting signal was then used for downstream cough detection and validity assessment.
Cough detection was implemented to identify the presence of cough events within audio filter results. This stage was trained on a dataset of 3000 manually annotated cough segments, capturing a wide variety of cough patterns and acoustic environments. A recurrent neural network24 was employed to model the temporal dynamics of cough sounds, allowing robust detection across diverse recording conditions. The output of this module serves as the initial filter, selecting segments that likely contain cough events for further analysis.
Following detection, each candidate cough segment was evaluated for its validity and suitability for downstream analysis, such as disease classification. This step was trained on a subset of 1000 carefully curated and high-quality annotated cough segments. A random forest25 classifier was employed to assess the segment quality, incorporating features such as signal-to-noise ratio, temporal consistency, and spectral characteristics. Only segments passing this validity check were retained for subsequent model training and evaluation.
Data preprocessing and feature extraction
To ensure high-quality input for downstream modeling, each recorded audio sample underwent a systematic preprocessing and feature extraction pipeline. This pipeline converts raw discrete-time waveforms into normalized log-mel filterbank representations, which capture both spectral and temporal characteristics of the signal while reducing sensitivity to amplitude variations and environmental noise. Specifically, the processing consists of waveform quantization, frame segmentation and windowing, short-time Fourier analysis, mel-scale filtering, logarithmic compression, and feature normalization. The following formalism describes each step in detail:
For each recorded audio sample, we denote the discrete-time waveform by
| 2 |
where xn is the amplitude of the n-th sample, N is the total number of samples, and Fs is the sampling rate in Hz. The waveform is transformed into a normalized log-mel filterbank (FBank) representation via three steps: waveform quantization, short-time spectral analysis with mel-filtering, and feature normalization.
- Waveform quantization The floating-point waveform is scaled to a 16-bit integer range:
where is the quantized sample. For simplicity, we denote it as xn in the subsequent steps.3 - Frame segmentation and windowing Frames of length Wms = 25ms with overlap Oms = 10ms are used. In samples:
where H is the hop size (frame shift). The k-th frame is extracted as4
where is the total number of frames. A Hamming window is applied:5
where ⊙ denotes element-wise multiplication.6 - Short-time Fourier transform and power spectrum For each windowed frame, the FFT of length NFFT is
and the power spectrum is7 8 - Mel filter bank and energy computation A triangular mel filter bank with M = 512 filters is applied:
where fp−1, fp, fp+1 are the corner frequencies of the p-th filter. The mel energy of frame k is9 10 - Logarithmic compression
where ε is a small constant (e.g., 10−8) for numerical stability. Stacking all frames yields the log-mel spectrogram11 12 - Feature normalization
where μ and σ are the mean and standard deviation over the training set. The factor of 2 in the denominator is empirically chosen for training stability.13 - Summary The entire process from waveform to normalized log-mel features can be compactly written as
14
Structured textual information—including patient demographics (e.g., age, gender, height, weight), smoking history, and symptom profiles—was first consolidated into standardized textual strings. For each patient, individual variables were concatenated with descriptive labels to form a coherent text representation. As an illustrative example, a subject’s information might be formatted as:
"gender: male; age: 64 years; height: 161 cm; weight: 59 kg; smoking history: never" .
In cases of missing values, the corresponding entry was consistently replaced with the placeholder "unknown", ensuring that the resulting text strings remained complete and structurally uniform for all subjects. For instance, if the weight value were unavailable, the corresponding string would become:
"gender: male; age: 64 years; height: 161 cm; weight: unknown; smoking history: never" .
These textual representations were subsequently tokenized using a pretrained tokenizer (RobertaTokenizerFast26), which converts each token into its corresponding vocabulary ID, and generates an attention mask identifying valid tokens. The resulting tensors—comprising both the input IDs and attention masks—served as the encoded input to the downstream model.
Model architecture
In this study, we present a device-invariant multimodal deep learning framework designed to jointly model cough audio signals and structured textual information, enabling robust detection of respiratory diseases in heterogeneous clinical settings (see Fig. 9a, while the hyperparameter and configuration are provided in Supplementary Tables 8 and 9). The proposed architecture comprises five principal modules: the input module, the encoder module, the multimodal fusion module, the disease classifier, and the adversarial learning branch.
The input module processes two complementary data modalities: (i) preprocessed cough audio represented as mel-spectrogram tensors, and (ii) textual information including symptom descriptions and demographic attributes, each converted into unified vector representations.
The encoder module consists of both an audio encoder and a text encoder, enabling the model to capture complementary acoustic and semantic information.
The audio encoder is built upon the EAT, a transformer-based architecture specifically designed for compact and robust modeling of temporal-spectral patterns in acoustic signals. Following preprocessing, the cough spectrogram is segmented into non-overlapping 16 × 16 patches, which are linearly projected into embedding vectors and augmented with learnable positional encodings. These embeddings are then processed through a series of transformer encoder blocks, each comprising multi-head self-attention, feed-forward sublayers, residual connections, and layer normalization. For a batch of input spectrograms, the audio encoder produces frame-level acoustic features of shape , where T denotes the number of frames and d is the embedding dimension. In the base configuration, the audio encoder comprises 12 transformer layers with a hidden dimension of 768 and 8 attention heads, whereas the large configuration employs 24 layers with a hidden dimension of 1024 and 16 attention heads. To obtain a fixed-length representation, the frame-level features are temporally aggregated via mean pooling and subsequently projected into the shared multimodal fusion space for downstream integration with text embeddings.
The text encoder leverages a RoBERTa-based transformer to capture contextual linguistic and semantic information from symptom descriptions and clinical narratives. RoBERTa enhances language understanding by employing dynamic masking, larger training corpora, and longer sequences during pretraining. We adopt the RoBERTa model (12 layers, 12 attention heads, hidden dimension = 768) for the configuration. Input sequences are tokenized with a maximum length of 256 tokens, and the [CLS] token from the final layer is used as the aggregated text embedding. The output textual features are projected into the shared latent space for multimodal integration.
To mitigate device-induced distribution shifts, an adversarial device branch is introduced during training. Audio features are first aggregated using self-attentive pooling27, then passed through a gradient reversal layer (GRL)28 before being fed into a device classifier. This adversarial optimization encourages the encoder to learn device-invariant representations. For sequence-level compression, attention pooling is applied: attention weights are obtained via linear projection, normalized through a softmax function, and used to compute a weighted sum of sequence elements, yielding fixed-length vectors for both intra- and inter-modal aggregation. The adversarial attention and hidden layers are aligned with the corresponding fusion dimensions in each configuration to ensure consistent representational capacity.
Multimodal fusion is implemented through a bidirectional cross-modal attention mechanism29 that enables effective interaction between audio and text modalities. Initially, audio and text embeddings are linearly projected into a common latent space, forming modality-specific subspaces. Two symmetric cross-modal attention streams—audio-to-text and text-to-audio—allow each modality to selectively attend to salient cues from the other, with each stream composed of multi-head attention layers, followed by dropout and layer normalization. The resulting modality-specific outputs are then aggregated via attention pooling to derive compact embeddings. These embeddings are concatenated and further processed through a self-attentive multimodal fusion block, consisting of multi-head attention30 layers with layer normalization, to enhance inter-modality dependency modeling.
Finally, the fused representation is fed into an enhanced classification head to generate the disease classification logits. The disease classifier is implemented as a multi-layer perceptron (MLP) integrating layer normalization31, dropout32, Gaussian error linear unit (GELU) activations33, and linear output layers. It produces a probability vector of shape , where num_labels corresponds to the number of diagnostic categories. In the base configuration, the classifier employs a 768-dimensional fusion space, 8 attention heads, and a dropout rate of 0.3. The large configuration expands these parameters to a 1024-dimensional fusion space with 16 attention heads and a reduced dropout rate of 0.2, thereby enhancing representational capacity.
Loss function design
To achieve both accurate recognition of respiratory diseases and robust modeling against device variability, we developed a dual-branch multi-objective loss framework that simultaneously optimizes two tasks: device-invariance modeling (device loss) and multi-label disease identification (disease joint loss). The overall architecture is illustrated in Fig. 10a1.
- Device loss The device branch is designed to enhance the model’s generalization capability across heterogeneous acquisition devices and comprises the following two components:
- IRM Penalty: Leveraging the IRM framework, a gradient-based penalty is introduced after partitioning data according to device environments, encouraging the model to learn a shared feature space across devices. Specifically, the primary task loss gradients are computed independently within each device-specific sub-environment, and their variability is minimized to enforce invariance. The process can be formulated as follows:
where is the set of training environments. (Xe, Ye) denotes the input and label pairs from environment e. Φ is the feature extractor shared across all environments. w is a linear classifier applied on top of Φ. is the empirical risk (e.g., cross-entropy loss) for environment e. is the gradient of the loss with respect to w. The squared norm ∥ ⋅ ∥2 penalizes deviations from invariance across environments.15 - Device Classification Loss: A GRL is inserted between the audio feature extractor and the device classifier to adversarially train the audio encoder, thereby confounding device-specific discriminative cues. This objective is implemented using the cross-entropy (CE) loss34:
16
Since the IRM penalty is a gradient-based regularization term, it is not suitable to apply it from the early stages of training. To address this, we employ a sigmoid-based dynamic gradient scaling (Fig. 10b1) for the penalty, which gradually increases its influence, thereby avoiding the performance degradation that can occur if IRM is introduced prematurely. The sigmoid-based dynamic gradient scaling is computed as:
| 17 |
where t is the current training step. T is the total number of steps, computed as:
is the maximum weight assigned to the IRM loss. k is the steepness parameter, controlling the growth rate of λIRM. represents the normalized training progress, ranging from 0 to 1. λIRM acts as a stability enhancement mechanism to gradually adjust the adversarial loss coefficient, and forms the device loss term through a weighted combination with and :
| 18 |
-
2.Disease joint loss For the disease joint loss, considering the complementarity between modalities and the inherent data uncertainty, we designed a unified framework comprising three sub-losses:
- Disease classification loss: Serving as the primary task loss, binary cross-entropy is used to independently classify each disease.
- Audio-text contrastive loss: To enhance the consistency of multimodal representations, an InfoNCE-style contrastive loss35 is employed, pulling closer the audio and text embeddings from the same patient while pushing apart those from different patients.
-
Modality alignment loss: The fused audio/text features are aligned in the shared feature space using KL divergence36 as a regularization constraint.Given the differing optimization difficulty and contribution of each sub-task, the overall joint loss is dynamically fused via an uncertainty weighting mechanism37, where each sub-task is associated with a learnable log-variance parameter:
where M is the total number of tasks or loss components. is the loss associated with the i-th task. denotes the task-dependent variance (uncertainty), which is a learnable parameter. adaptively scales the i-th loss according to its uncertainty. acts as a regularization term to avoid trivial solutions.19
To prevent the device-adversarial task from strongly interfering with disease classification in the early stages of training, we adopt a linear warmup strategy30 (Fig. 10 b2), allowing the audio encoder to prioritize learning disease-relevant features first. Once disease classification stabilizes, device-adversarial constraints are gradually introduced:20 The final optimization objective integrates device robustness and disease recognition into a single weighted total loss:21
Experimental setup and evaluation metrics
To evaluate model performance in real-world clinical scenarios, we designed a comprehensive experimental protocol:
Modality ablation We systematically evaluate the contribution of different modalities through ablation studies (Table 8):
-
2.Adversarial strategy validation We examine two key adversarial strategies:
- Domain adversarial training (DAT): Device discriminator + GRL for device-invariant audio encoding.
- IRM-based optimization: Enhances consistency across heterogeneous device distributions.
Table 8.
Comparison of model parameters
| Modalities | Model structure |
|---|---|
| Audio only | Fig. 9c |
| Audio + Demographics | Fig. 9a (w/o Adversarial Device Module) |
| Audio + Symptoms | Fig. 9a (w/o Adversarial Device Module) |
| Demographics+Symptoms (no audio) | Fig. 9d |
| Full multimodala | Fig. 9a (w/o Adversarial Device Module) |
| Full multimodala + Adversarial training | Fig. 9a |
aFull multimodal (Audio + Demographics + Symptoms).
These are evaluated across multi-label, binary, and ablation tasks. To further assess adversarial effectiveness, we visualize [CLS] token embeddings from the audio encoder using t-SNE, comparing cluster separation (device-wise vs. disease-wise) before and after adversarial training. We also compare training time and performance metrics across three regimes: Non-adversarial, Adversarial (w/o IRM), Adversarial (with IRM).
-
3.Model scaling and deployment readiness We benchmark two model sizes:
- base-size: Moderate parameter count, optimized for low-resource deployment.
- large-size: Approximately twice the parameters of the base, to explore performance ceilings.
We report AUC, inference latency, and resource utilization to inform future deployment on mobile or edge devices.
-
4.Evaluation metrics We adopt clinically meaningful metrics, including:
- AUROC: Discrimination across thresholds; robust to class imbalance.
- AUPRC: Emphasizes precision in low-prevalence classes.
- F1 Score: Balances precision and recall.
- Accuracy, Precision, Sensitivity (Recall): Standard binary classification measures.
- Youden Index: Determines optimal threshold by maximizing (Sensitivity + Specificity - 1), aligning with diagnostic utility.
Supplementary information
Acknowledgements
This work was supported by the National Key R&D Program of China (2022YFC2010005).
Author contributions
M.Y. conceived and designed the study. X.L., W.D., Y.L., W.Z., Z.B., and J.M. collected the data. M.Y., W.Z., and Z.B. analyzed the data. M.Y., Q.W. and Y.L. drafted the manuscript. Q.W., S.C., M.Z., and J.Q. revised the draft. Q.W. supervised the study.
Data availability
The datasets generated and/or analyzed during the current study are not publicly available due to the inclusion of sensitive clinical information collected under institutional and regulatory data-use agreements, as well as proprietary components that cannot be openly released, but are available from the corresponding author upon reasonable request.
Code availability
The code developed in this study is proprietary and has substantial commercial potential; accordingly, it cannot be made publicly available. Due to intellectual property protections and ongoing commercialization activities, the source code cannot be shared at this time. All model development, training, and analysis were conducted using Python 3.10 with PyTorch 2.1.0 (which can be accessed at https://pytorch.org/get-started/previous-versions/). Specific training configurations and parameters used to generate and analyze the datasets are detailed in the “Methods” section.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Xuefei Liu, Wei Du.
Contributor Information
Qian Wang, Email: qian.wang173@hotmail.com.
Si Chen, Email: echo.chen@lucahealthcare.com.
Min Zhou, Email: doctor_zhou_99@163.com.
Jie-ming Qu, Email: jmqu0906@163.com.
Supplementary information
The online version contains supplementary material available at 10.1038/s41746-026-02445-4.
References
- 1.Wang, Z. et al. Global, regional, and national burden of chronic obstructive pulmonary disease and its attributable risk factors from 1990 to 2021: an analysis for the global burden of disease study 2021. Respir. Res.26, 2 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Bhakta, N. R., McGowan, A., Ramsey, K. A. et al. European Respiratory Society/american thoracic society technical statement: standardisation of the measurement of lung volumes, 2023 update. Eur. Respir. J.62, 2201519 (2023). [DOI] [PubMed] [Google Scholar]
- 3.Thawanaphong, S. & Nair, P. Contemporary concise review 2024: chronic obstructive pulmonary disease. Respirology30, 574–586 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Agusti, A. & Vogelmeier, C. F. Gold 2024: a brief overview of key changes. J. Bras. Pneumol.49, e20230369 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Kim, S. H. & Han, M. K. Challenges and the future of pulmonary function testing in chronic obstructive pulmonary disease (copd): toward earlier diagnosis of copd. Tuberc. Respir. Dis.88, 413–418 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Chu, Y. et al. Cycleguardian: a framework for automatic respiratory sound classification based on improved deep clustering and contrastive learning. Complex Intell. Syst.11, 200 (2025). [Google Scholar]
- 7.Isangula, K. G. & Haule, R. J. Leveraging ai and machine learning to develop and evaluate a contextualized user-friendly cough audio classifier for detecting respiratory diseases: Protocol for a diagnostic study in rural Tanzania. JMIR Res. Protoc.13, e54388 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Sharan, R. V. & Xiong, H. Wet and dry cough classification using cough sound characteristics and machine learning: a systematic review. Int. J. Med. Inform.199, 105912 (2025). [DOI] [PubMed] [Google Scholar]
- 9.Huddart, S. et al. A dataset of solicited cough sound for tuberculosis triage testing. Sci. Data11, 1149 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Morocutti, T., Schmid, F., Koutini, K. & Widmer, G. Device-robust acoustic scene classification via impulse response augmentation. In Proc. 31st European Signal Processing Conference (EUSIPCO) 176-180 (IEEE, 2023).
- 11.Mezza, A. I., Habets, E. A., Müller, M. & Sarti, A. Unsupervised domain adaptation for acoustic scene classification using band-wise statistics matching. In Proc. 28th European Signal Processing Conference (EUSIPCO) 11−15 (IEEE, 2020).
- 12.Ma, C., Wang, H. & Hoi, S. C. H. Multi-label thoracic disease image classification with cross-attention networks. In Proc. Int. Conf. on Medical Image Computing and Computer-Assisted Intervention 730−738 (Cham: Springer International Publishing, 2019).
- 13.Lei, T., Hu, Q., Hou, Z. & Lu, J. Enhancing real-world far-field speech with supervised adversarial training. Appl. Acoust.229, 110407 (2025). [Google Scholar]
- 14.Arjovsky, M., Bottou, L., Gulrajani, I. & Lopez-Paz, D. Invariant risk minimization. https://arxiv.org/abs/1907.02893 (2020).
- 15.Wang, J. et al. Joint asymmetric loss for learning with noisy labels. In Proc. IEEE/CVF International Conference on Computer Vision 1947−1956 (IEEE, 2025).
- 16.Gong, Y., Chung, Y.-A. & Glass, J. Ast: audio spectrogram transformer. In Proc. Interspeech 571-575 (2021).
- 17.van der Maaten, L. & Hinton, G. Visualizing data using t-SNE. J. Mach. Learn. Res.9, 2579–2605 (2008). [Google Scholar]
- 18.Santosh, K. C., Rasmussen, N., Mamun, M. & Aryal, S. A systematic review on cough sound analysis for covid-19 diagnosis and screening: is my cough sound COVID-19? PeerJ Comput. Sci.8, e958 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Santamaria, M. et al. Longitudinal voice monitoring in a decentralized bring your own device trial for respiratory illness detection. npj Digit. Med.8, 202 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Mersha, M., Lam, K., Wood, J., AlShami, A. K. & Kalita, J. Explainable artificial intelligence: a survey of needs, techniques, applications, and future direction. Neurocomputing599, 128111 (2024). [Google Scholar]
- 21.Baiardi, A. & Naghi, A. A. The value added of machine learning to causal inference: evidence from revisited studies. Econom. J.27, 213–234 (2024). [Google Scholar]
- 22.Poinsot, A. et al. Position: causal machine learning requires rigorous synthetic experiments for broader adoption. In Proc. 42nd International Conference on Machine Learning 81995-82015 (PMLR, 2025).
- 23.Guo, L.-Z., Jia, L.-H., Shao, J.-J. & Li, Y.-F. Robust semi-supervised learning in open environments. Front. Comput. Sci.19, 198345 (2025). [Google Scholar]
- 24.Schmidt, R. M. Recurrent neural networks (rnns): a gentle introduction and overview https://arxiv.org/abs/1912.05911 (2019).
- 25.Breiman, L. Random forests. Mach. Learn.45, 5–32 (2001). [Google Scholar]
- 26.Liu, Y. et al. Roberta: a robustly optimized bert pretraining approach https://arxiv.org/abs/1907.11692 (2019).
- 27.Chen, F., Datta, G., Kundu, S. & Beerel, P. Self-attentive pooling for efficient deep learning. In Proc. IEEE/CVF Winter Conference on Applications of Computer Vision 3974-3983 (IEEE, 2023).
- 28.Ganin, Y. et al. Domain-adversarial training of neural networks. J. Mach. Learn. Res.17, 1–35 (2016). [Google Scholar]
- 29.Luo, J., Phan, H., Wang, L. & Reiss, J. D. Bimodal connection attention fusion for speech emotion recognition https://arxiv.org/abs/2503.05858 (2025).
- 30.Vaswani, A. et al. Attention is all you need. In Proc. Advances in Neural Information Processing Systems30 (2023).
- 31.Aly, H., Al-Ali, A. K. & Suganthan, P. N. Boosted multilayer feedforward neural network with multiple output layers. Pattern Recognit.156, 110740 (2024). [Google Scholar]
- 32.Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I. & Salakhutdinov, R. Dropout: a simple way to prevent neural networks from overfitting. J. Mach. Learn. Res.15, 1929–1958 (2014). [Google Scholar]
- 33.Hendrycks, D. & Gimpel, K. Gaussian error linear units (GELUs) https://arxiv.org/abs/1606.08415 (2016).
- 34.Rumelhart, D. E., Hinton, G. E. & Williams, R. J. Learning representations by back-propagating errors. Nature323, 533–536 (1986). [Google Scholar]
- 35.Wang, Z., Xu, B., Yuan, Y., Shen, H. & Cheng, X. Infonce is a free lunch for semantically guided graph contrastive learning. In Proc. 48th International ACM SIGIR Conference on Research and Development in Information Retrieval 719–728 10.1145/3726302.3730007 (ACM, 2025).
- 36.He, K. et al. Dalr: dual-level alignment learning for multimodal sentence representation learning. In Proc. Findings of the Association for Computational Linguistics: ACL 2025 3586-3601 (2025).
- 37.Kendall, A., Gal, Y. & Cipolla, R. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proc. IEEE Conference on Computer Vision and Pattern Recognition 7482-7491 (IEEE, 2018).
- 38.Chen, W., Liang, Y., Ma, Z., Zheng, Z. & Chen, X. EAT: Self-supervised pre-training with efficient audio transformer. In Proc. Thirty-Third International Joint Conference on Artificial Intelligence 3807-3815 (2024).
- 39.Kim, J.-W., Toikkanen, M., Choi, Y., Moon, S.-E. & Jung, H.-Y. Bts: bridging text and sound modalities for metadata-aided respiratory sound classification. In Interspeech 2024 1690–1694 (GitHub, 2024).
- 40.Shao, N., Li, X. & Li, X. Fine-tune the pretrained atst model for sound event detection. In Proc.ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 911–915 (IEEE, 2024). [DOI] [PMC free article] [PubMed]
- 41.Chen, S. et al. BEATs: audio pre-training with acoustic tokenizers. In Proc. 40th International Conference on Machine Learning 5178-5193 (2023).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets generated and/or analyzed during the current study are not publicly available due to the inclusion of sensitive clinical information collected under institutional and regulatory data-use agreements, as well as proprietary components that cannot be openly released, but are available from the corresponding author upon reasonable request.
The code developed in this study is proprietary and has substantial commercial potential; accordingly, it cannot be made publicly available. Due to intellectual property protections and ongoing commercialization activities, the source code cannot be shared at this time. All model development, training, and analysis were conducted using Python 3.10 with PyTorch 2.1.0 (which can be accessed at https://pytorch.org/get-started/previous-versions/). Specific training configurations and parameters used to generate and analyze the datasets are detailed in the “Methods” section.










