Graphical abstract
Overview of the study. The graphical elements of this figure were created using Canva (Canva Pty Ltd., Sydney, Australia). EHR: electronic health record.
Abstract
Background
COPD remains a leading cause of global morbidity and mortality, with acute exacerbations driving disease progression and healthcare utilisation. Artificial intelligence (AI) offers new opportunities to predict exacerbation risk by integrating multimodal data such as electronic health records (EHRs), spirometry and wearable sensor inputs.
Methods
This systematic review, conducted in accordance with PRISMA 2020 guidelines and registered in PROSPERO (CRD420251165476), evaluated AI-based models developed for COPD exacerbation prediction using combined data modalities.
Results
Comprehensive searches of PubMed, Embase and Google Scholar identified 859 records, of which five studies published between 2021 and 2025 met inclusion criteria. Study designs ranged from prospective monitoring cohorts to EHR-based and hybrid datasets. Models applied diverse approaches including random forests, gradient boosting, convolutional neural networks and ensemble learning frameworks. Reported discriminative performance was moderate to high, with area under the curve values between 0.73 and 0.92 and accuracies up to 0.92. Most of these performance metrics were derived from internal validation, with limited external testing, which restricts assumptions about generalisability. Sensitivity reached 0.94 in wearable-driven models, while only one study reported formal calibration assessment.
Conclusions
Despite encouraging performance, methodological heterogeneity, limited external validation and incomplete reporting of preprocessing and explainability methods restrict clinical translation. Current evidence supports the potential of multimodal AI to enhance early detection of COPD exacerbations, but future research must prioritise transparent reporting, external validation and integration into real-world care pathways.
Shareable abstract
Integrating multimodal data with artificial intelligence can enhance the prediction of COPD exacerbations, but greater methodological rigour and external validation are essential before clinical translation https://bit.ly/45l4jgW
Introduction
COPD is a leading cause of morbidity and mortality worldwide, responsible for over three million deaths annually and projected to become the fourth leading cause of death globally [1–3]. Characterised by persistent airflow limitation and progressive decline in lung function, COPD is marked by acute exacerbations, which accelerate disease progression and contribute disproportionately to healthcare utilisation [1, 2]. Despite therapeutic advances, early identification of patients at risk for exacerbation remains challenging due to the multifactorial nature of triggers spanning physiological, environmental and behavioural factors. Emerging AI and machine learning (ML) approaches have demonstrated promise in capturing these complex interactions by integrating high-dimensional and heterogeneous data [1, 4, 5]. Traditional statistical or unimodal models relying solely on clinical or spirometric parameters often fail to account for temporal variation and nonlinear relationships in disease trajectories [2, 6]. In contrast, multimodal AI frameworks that combine electronic health records (EHRs), physiological signals and wearable or remote monitoring data may offer improved predictive performance by fusing complementary information [2, 5, 7]. However, published models vary substantially in design, validation strategies and transparency, with persistent concerns regarding data bias, interpretability and external generalisability [4, 6, 8–11]. To address these limitations, the present systematic review aimed to identify and critically appraise AI-based multimodal models predicting COPD exacerbations. Following PRISMA guidelines [12] and evaluating methodological rigour using PROBAST [13], this review synthesises evidence on study design, data modalities, model performance and risk of bias. Although established risk factors such as prior exacerbations, gastroesophageal reflux, emphysema and chronic bronchitis remain central to COPD prognostication, traditional models rely mainly on linear relationships and limited variable sets. AI methods may complement these tools by capturing nonlinear interactions and integrating high-dimensional data sources not easily incorporated into conventional approaches [4, 5, 14]. Findings highlight both the emerging potential and current methodological challenges of multimodal AI for COPD prognostication, emphasising the importance of transparent reporting and robust validation frameworks [4, 6, 15].
Materials and methods
Study design
This systematic review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA 2020) guidelines to ensure methodological transparency and reproducibility [12]. The approach followed the prespecified protocol entitled “The Role of Artificial Intelligence in Predicting COPD Exacerbations Using Multimodal Data”, which was developed before the literature search and implemented without deviations. The protocol was prospectively registered on the International Prospective Register of Systematic Reviews (PROSPERO; ID: CRD420251165476), providing public access to predefined objectives, eligibility criteria and methodological details in line with best practice standards [16]. The Prediction Model Risk of Bias Assessment Tool (PROBAST) was employed to evaluate the quality and risk of bias in included studies, with adaptations for AI-specific contexts [13]. The review aimed to identify, evaluate and synthesise studies applying AI or ML algorithms for predicting exacerbations in COPD using multimodal data sources. These included combinations of EHRs, spirometry, wearable sensor data and clinical trial datasets. To enhance clarity, the overall workflow of the review, from literature identification to study inclusion, is illustrated in figure 1. The flowchart presents the number of records retrieved, screened, excluded and ultimately included at each stage of the process.
FIGURE 1.
Flowchart summarising the study selection process. A total of 859 records were retrieved across PubMed, Google Scholar and Embase; 523 records remained after deduplication. Following title/abstract screening (76 full texts assessed), 16 studies met initial eligibility criteria applied in the full text, and five were retained in the final synthesis as they employed multimodal AI models for predicting COPD exacerbations. The review process followed PRISMA 2020 guidelines and corresponds to PROSPERO registration CRD420251165476. COPDgen: Genetic Epidemiology of COPD study; AECOPD: acute exacerbation of COPD; CT: computed tomography. The graphical elements of this figure were created using Canva (Canva Pty Ltd., Sydney, Australia).
Information sources and search strategy
A comprehensive search strategy was executed across PubMed/MEDLINE, Google Scholar, and Embase databases up to 20 August 2025, following the guidance of the Cochrane Handbook for Systematic Reviews [17]. The search strategy combined Medical Subject Headings (MeSH) and free-text terms related to COPD, AI, ML, exacerbation, and multimodal data integration. To support reproducibility, the search was structured around five thematic query groups (AI/ML prediction, wearables, environmental data, imaging/EHR, and deep-learning methods), each combining MeSH terms with relevant free-text keywords. The complete database-ready search strings are provided in Supplementary Table S3.
The search was structured along three thematic dimensions:
AI/ML-driven prediction of COPD exacerbations,
Integration of multimodal data types (EHRs, imaging, wearable or environmental data), and
Advanced deep learning and neural network approaches.
Boolean operators were used to combine terms, and database-specific syntax was adapted as needed. The reference lists of included studies were screened to identify additional relevant publications [16]. In total, 859 records were retrieved. After removing 336 duplicates, 523 unique studies were screened by title and abstract. The detailed flowchart (figure 1) illustrates the selection pathway.
Study selection
All records were independently screened by two reviewers using the Rayyan web platform, following a two-phase process (title/abstract, then full-text review) [18]. Disagreements were resolved through consensus with a senior reviewer. The PICO framework (Population, Intervention, Comparison, Outcome) guided inclusion to ensure clinical and methodological relevance [19]. Of the 859 records initially retrieved, 523 unique articles remained after removing duplicates. After title and abstract screening, 76 full texts were reviewed in detail. 16 studies met preliminary criteria, and five were finally included after assessing methodological completeness and multimodal integration: Wu et al. (2021) [20], Singla et al. (2021) [21], Singh et al. (2022) [22], Hussain et al. (2021) [23] and Jeon et al. (2025) [24].
Eligibility criteria
The eligibility criteria were identical to those established in the review protocol. Studies were included if they:
presented peer-reviewed original research;
applied AI or ML algorithms to predict COPD exacerbations;
used ≥2 data modalities;
reported at least one model performance metric (e.g., area under the curve (AUC), accuracy, sensitivity, specificity or F1 score).
Exclusion criteria comprised:
studies without AI/ML methodology;
single-modality models;
reviews, protocols, editorials or case reports;
non-human or non-English publications.
These criteria ensured inclusion of empirical studies providing transparent performance reporting, consistent with best practices in AI model evaluation [1, 2].
Data extraction
Data extraction was conducted using a standardised Excel template designed for this review, based on Cochrane methodological recommendations [25]. Extracted variables included study characteristics (authors, year, country, sample size, design), data modalities, model type, validation strategy and performance metrics (AUC, sensitivity, specificity, accuracy, F1 score, Brier score). Information on calibration, decision-curve analysis and bias mitigation techniques (e.g., regularisation, oversampling, early stopping) was also collected. No author contact was necessary, as sufficient data were available from published materials and supplementary files. Reporting on feature preprocessing, missing-data strategies and hyperparameter tuning was often incomplete in the included studies, which constrained the level of methodological detail that could be uniformly extracted and synthesised. Detailed summaries of extracted study characteristics and performance metrics are provided in supplementary Table S1 and supplementary Figure S1.
Risk of bias assessment
Each of the five included studies was evaluated using the PROBAST, following the guidance of Wolff et al. (2019) [13]. The assessment covered four principal domains:
Participants – adequacy of population sampling, inclusion criteria and representativeness of real-world COPD populations.
Predictors – clarity of feature definition, measurement consistency and harmonisation across data modalities.
Outcome – appropriateness and consistency of COPD exacerbation definition and ascertainment.
Analysis – robustness of model development, control of overfitting, validation strategy and reporting transparency.
To address the unique challenges of AI-driven prediction studies, an AI-adapted PROBAST framework was applied, consistent with recent guidance from Riley et al. [9] and Collins et al. [8], which recommends incorporating algorithmic transparency, interpretability, data preprocessing and fairness considerations [4]. Each study was rated as low, moderate or high risk of bias per domain. AI-specific aspects, such as data imbalance, feature selection, hyperparameter optimisation and limited model explainability, were considered potential sources of bias [4, 6, 10]. Applicability (external validity) was also evaluated according to study setting, data source and population coverage. Overall, the participant and outcome domains were commonly judged as low risk, whereas the predictor and analysis domains more frequently exhibited moderate risk, mainly due to incomplete reporting of feature preprocessing and internal validation procedures. Only one study (Jeon et al., 2025) [24] implemented external temporal validation, a recommended practice for AI prognostic models [6, 26]. The domain-level PROBAST ratings for all included studies that cover the participants, predictors, outcome and analysis domains are provided in table 1 to support the risk of bias judgements presented in the Results section.
TABLE 1.
Risk of bias assessment using the PROBAST framework across included studies
| Study | Participants | Predictors | Outcome | Analysis | Overall risk |
|---|---|---|---|---|---|
| Wu et al., 2021 [20] | ▪ Low | ▪ Moderate | ▪ Low | ▪ Moderate | ▪ Moderate |
| Singla et al., 2021 [21] | ▪ Low | ▪ Moderate | ▪ Low | ▪ Moderate | ▪ Moderate |
| Singh et al., 2022 [22] | ▪ Low | ▪ Moderate | ▪ Low | ▪ Moderate | ▪ Moderate |
| Hussain et al., 2021 [23] | ▪ Low | ▪ Moderate–High | ▪ Low | ▪ High | ▪ Moderate–High |
| Jeon et al., 2025 [24] | ▪ Low | ▪ Moderate | ▪ Low | ▪ Moderate | ▪ Moderate |
PROBAST: Prediction Model Risk of Bias Assessment Tool; Low: criteria met with minimal concern for bias; Moderate: some methodological or reporting concerns; High: substantial bias risk likely to affect validity. Coloured symbols represent domain-level ratings: ▪ Low, ▪ Moderate, ▪ Moderate–High, ▪ High.
Data synthesis
Given heterogeneity in study design, data sources and performance metrics, a qualitative narrative synthesis was employed, consistent with Popay et al. (2006) [27]. The five included studies, published between 2021 and 2025, examined diverse designs, from prospective home-monitoring cohorts to retrospective EHR analyses and secondary evaluations of international randomised controlled trials. Sample sizes ranged from under 100 participants in wearable-based studies to over 40 000 in EHR-derived cohorts. All studies included adults with spirometry-confirmed COPD according to Global Initiative for Chronic Obstructive Lung Disease (GOLD) criteria. Every study incorporated multimodal data, integrating combinations of EHR variables, spirometry, trial data or wearable sensors. This approach aligns with prior evidence indicating that multimodal fusion enhances predictive accuracy over unimodal baselines [1, 2, 28]. Outcomes were uniformly defined based on GOLD or trial criteria for moderate (pharmacological intervention) and severe (hospitalisation) exacerbations. Prediction horizons varied: 7-day short-term risk (Wu et al., 2021 [20]) versus annual or 12-month risk models (Singh et al., 2022 [22]; Jeon et al., 2025 [24]; Singla et al., 2021 [21]; Hussain et al., 2021 [23]). Most studies applied internal validation (cross-validation or train/test split), with one performing external validation. Reported AUCs ranged from ∼0.70 to 0.89, comparable to other state-of-the-art AI models in respiratory disease [1, 6]. Short-term multimodal approaches integrating wearable or physiological data yielded the highest discrimination. In PROBAST assessments, participant and outcome domains were predominantly low risk, reflecting transparent selection and standardised end-points. Predictor and analysis domains presented moderate concerns due to insufficient detail on data preprocessing and feature engineering. Applicability was high in EHR-based and wearable studies, though trial-based analyses may have limited generalisability. A complete summary of extracted results and methodological characteristics is available in the supplementary material.
Results
Study selection
The literature search returned 859 records across PubMed, Embase and Google Scholar. Following removal of 336 duplicate entries, 523 unique records underwent title and abstract screening. This screening was facilitated using the Rayyan platform to ensure transparent and efficient management [18]. Of those, 76 full texts were assessed for eligibility, with 16 studies initially meeting inclusion criteria. A subsequent manual review focusing on multimodal data integration and methodological completeness resulted in five final studies: Wu et al. (2021) [20], Singla et al. (2021) [21], Singh et al. (2022) [22], Jeon et al. (2025) [24] and Hussain et al. (2021) [23]. These studies satisfied all predefined inclusion criteria, and their screening trajectory is detailed in figure 1, following the PRISMA 2020 flow schematic [12].
Characteristics of included studies
The five studies included, published between 2021 and 2025, represented a broad range of methodological designs and data environments, encompassing prospective remote-monitoring cohorts, retrospective EHR analyses, and hybrid multimodal frameworks conducted across Europe, Asia and North America. Sample sizes varied from 67 participants in the wearable-based study by Wu et al. (2021) [20] to large institutional datasets of several thousand individuals analysed by Hussain et al. (2021) [23]. All studies enrolled adults with spirometry-confirmed COPD, consistent with GOLD diagnostic standards [1, 2]. Wu et al. (2021) developed a deep neural network integrating continuous physiological signals, environmental exposure and self-reported symptoms to forecast short-term (7-day) exacerbation risk. Singla et al. (2021) [21] leveraged COPDGene data to combine spirometry, computed tomography features and clinical records within a gradient-boosting framework. Singh et al. (2022) [22] employed a random forest model incorporating lung function indices, medication profiles and prior exacerbation history for 1-year prediction. Jeon et al. (2025) [24] implemented a convolutional neural network fusing wearable and EHR data to predict individualised 12-month risk trajectories. Hussain et al. (2021) [23] applied an ensemble learning strategy using structured clinical features and spirometry-derived measurements, rather than wearable or sensor data, to differentiate between mild and severe exacerbations. A consolidated overview of the data modalities and modelling techniques used across the studies is presented in table 2, illustrating the growing adoption of multimodal integration and deep-learning architectures in AI research for COPD management.
TABLE 2.
Summary of data modalities and AI model types across included studies
| Study | EHR | Spirometry | Wearable | Trial data | Model type |
|---|---|---|---|---|---|
| Wu et al., 2021 [20] | ✓ | ✓ | LSTM | ||
| Singla et al., 2021 [21] | ✓ | ✓ | Gradient Boosting | ||
| Singh et al., 2022 [22] | ✓ | ✓ | Random Forest | ||
| Jeon et al., 2025 [24] | ✓ | ✓ | ✓ | CNN | |
| Hussain et al., 2021 [23] | ✓ | ✓ | Soft Voting Ensemble Classifier |
EHR: electronic health records; LSTM: Long Short-Term Memory (a type of recurrent neural network); CNN: Convolutional Neural Network; Gradient Boosting: sequential ensemble of weak learners that iteratively corrects previous errors; Random Forest: ensemble of decision trees using random subsets of data to improve stability and reduce overfitting; Soft Voting Ensemble Classifier: merges probabilities from several models to improve prediction reliability. AI: artificial intelligence; ✓: indicates inclusion of the respective data modality; Spirometry: pulmonary function testing measuring airflow limitation and lung capacity (e.g., forced expiratory volume in 1 s, forced vital capacity); Wearable: sensor-based physiological monitoring devices (e.g., activity trackers, pulse oximeters); Trial data: data derived from clinical trials or longitudinal intervention studies.
AI models, validation approaches and performance synthesis
Across the five studies, a range of AI and ML frameworks were implemented, reflecting the methodological diversity captured in the registered PROSPERO protocol (CRD420251165476). Tree-based ensemble algorithms such as random forests and gradient boosting were widely used for structured clinical or EHR datasets [21, 22]. Deep-learning methods, including recurrent and convolutional architectures, predominated in studies that incorporated continuous or multimodal physiological signals [20, 24]. Hussain et al. (2021) [23] applied an ensemble learning strategy to differentiate exacerbation severity, combining structured clinical and spirometry-derived variables.
Validation practices varied considerably. Three studies [20, 21, 23] conducted internal validation using k-fold cross-validation or train/test splits, whereas Jeon et al. (2025) [24] performed temporal validation using an internal–external split with bootstrapped confidence intervals. None of the studies employed independent external datasets, underscoring the limited evidence for model generalisability [6, 9, 26, 29]. Bootstrap resampling (n=1000) was used in one study [22] to estimate internal stability. Only two studies explicitly described measures to reduce overfitting, including early stopping and feature regularisation.
To facilitate transparency and enable comparison across studies, table 3 summarises the validation frameworks, resampling strategies and calibration procedures used. Internal validation predominated, and reporting of external testing, calibration and decision curve analyses was limited.
TABLE 3.
Summary of AI model validation characteristics across included studies
| Study | Validation type | External validation | Cross-validation | Hyperparameter turning reported | Overfitting control | Bootstrap/resampling | Calibration reported | Decision curve analysis |
|---|---|---|---|---|---|---|---|---|
| Wu et al. 2021 [20] | Train/test with three-fold internal CV | No | Yes (three-fold) | Yes (model layers and settings reported) | Batch normalisation; threshold policy (p>0.7) | No | No | No |
| Singla et al., 2021 [21] | Five-fold internal CV (10 300 subjects) | No | Yes | Partly (architecture details; settings in text/supplementary material) | Class-imbalance random oversampling (ROS) | No | Yes (Hosmer–Lemeshow, p=0.079) | No |
| Singh et al., 2022 [22] | Hold-out internal test set | No | No | Not clearly stated | Not reported | No | Not clearly reported | Not reported |
| Hussain et al., 2021 [23] | Five-fold internal CV | No | Yes | Yes | SMOTE oversampling; some early-stopping language in methods | No | No | No |
| Jeon et al., 2025 [24] | Development+external validation cohort | Yes (separate hospital cohort) |
No | Yes (details in supplementary material) | Early stopping (described in workflow) | Yes (bootstrapped 95% CIs for AUC) |
Not explicitly reported | Not reported |
AI: artificial intelligence; Validation type: approach used to evaluate model performance during development or testing; External validation: assessment of model performance on an independent dataset not used in model training; Cross-validation: (CV) internal resampling procedure in which the dataset is divided into multiple folds to estimate stability and generalisability; Hyperparameter tuning: optimisation of configuration parameters (e.g., learning rate, depth, or number of layers) to improve predictive accuracy; Overfitting control: strategies employed to avoid model overfitting and improve robustness, such as dropout, early stopping, or regularisation; Bootstrap/resampling: repeated sampling procedure to estimate model performance variability and generate confidence intervals; Calibration: comparison between predicted probabilities and observed outcomes, often measured with Brier score or Hosmer–Lemeshow test; Decision curve analysis: evaluation of clinical usefulness of prediction models by assessing net benefit across a range of thresholds; AUC: area under the receiver operating characteristic curve; SMOTE: Synthetic Minority Oversampling Technique (a method to balance minority classes).
TABLE 4.
Summary of key results and evidence interpretation from the included studies
| Research area | Key challenge identified in included studies | Recommended action/future direction |
|---|---|---|
| Data Integration | Multimodal data (EHR, spirometry and wearable signals) were inconsistently structured, with heterogeneous sampling rates and missing data across cohorts | Establish standardised COPD data integration pipelines and shared ontologies to ensure harmonisation across modalities and reproducibility of predictive modelling |
| Validation | All models primarily relied on internal validation; no study used a fully external dataset. Temporal validation was rare | Conduct prospective, multicentre and temporally independent validation studies across geographically and demographically diverse COPD populations to ensure generalisability |
| Transparency and Reporting | Feature preprocessing, missing data handling and hyperparameter tuning were often underreported | Adopt AI-specific reporting guidelines (e.g., TRIPOD-AI, STARD-AI and PROBAST-AI) to promote transparent documentation and reproducibility |
| Calibration and Clinical Utility | Calibration metrics (e.g., slope, Brier score) and decision curve analyses were inconsistently reported | Integrate calibration plots and decision curve analysis into standard performance evaluation to improve clinical interpretability |
| Fairness and Bias | None of the studies evaluated subgroup performance (e.g., sex, age, disease severity), risking biased model outputs | Incorporate bias detection and fairness auditing protocols to ensure equitable model performance across patient subgroups |
| Explainability | Limited use of interpretable modelling or post hoc explanation methods hindered clinical trust | Employ explainable AI (XAI) methods such as SHAP, attention maps or feature attribution analyses to improve clinician acceptance and model accountability |
| Implementation and Clinical Translation | Transition from model development to deployment is limited by interoperability, regulatory and workflow barriers | Develop clinician-centred interfaces, ensure interoperability with EHR systems, and align with regulatory frameworks for safe AI deployment in healthcare |
| Open Science and Collaboration | Restricted data sharing and lack of common evaluation benchmarks limit replication and cumulative progress | Promote open-access repositories, transparent code release and collaborative consortia to accelerate reproducibility and real-world translation |
EHR: electronic health record; AI: artificial intelligence; TRIPOD-AI: Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis in Artificial Intelligence; STARD-AI: Standards for Reporting of Diagnostic Accuracy Studies using Artificial Intelligence; PROBAST-AI: Prediction Model Risk of Bias Assessment Tool – Artificial Intelligence; XAI: Explainable Artificial Intelligence; SHAP: Shapley Additive Explanations.
Performance metrics demonstrated moderate to high discrimination across designs. Reported AUC values ranged from 0.73 to 0.92, with corresponding accuracies between 0.86 and 0.92. The wearable-based model of Wu et al. (2021) [20] achieved the highest short-term predictive sensitivity (0.94) and F1 score (0.92). Singh et al. (2022) [22] reported a Brier score of 0.082 and a positive predictive value of 0.71, with decision curve analysis supporting clinical utility. Specificity values exceeded 0.80 in three studies [20, 21, 23], and only one investigation [22] assessed calibration through graphical plots and slope statistics. No study performed subgroup or fairness analyses. Although discriminative performance was generally strong, calibration was largely underreported across studies. Only one study provided metrics such as a Brier score or calibration slope, limiting insight into how well predicted risks corresponded to actual outcomes. This gap is important, as good discrimination does not guarantee clinically reliable probability estimates. The scarcity of calibration reporting aligns with concerns raised in the broader prediction modelling literature, where probability accuracy remains underemphasised despite its relevance for clinical decision-making [15, 30].
Taken together, the validation evidence indicates that current AI models for predicting COPD exacerbations demonstrate promising internal discrimination but remain methodologically heterogeneous. Consistent external validation, standardised calibration reporting and transparent hyperparameter documentation are needed to enhance reproducibility and clinical translatability [8, 11, 13, 15, 26].
Risk of bias and methodological quality
Risk of bias was assessed for all five studies using the Prediction Model Risk of Bias Assessment Tool (PROBAST), adapted for AI-based prediction models. Overall, methodological quality was rated as moderate, reflecting consistent strengths in participant selection and outcome definition, but limitations in predictor handling and analytical transparency. The participants and outcome domains were generally at low risk of bias, since all studies enrolled adults with spirometry-confirmed COPD and used standardised definitions of moderate or severe exacerbations, consistent with GOLD criteria [1, 2]. In contrast, the predictors domain frequently showed moderate risk, mainly due to incomplete descriptions of feature preprocessing, missing data management and model explainability. This consistent pattern of moderate risk in the predictors and analysis domains indicates that methodological limitations were common across studies and may affect the reliability and comparability of their reported performance metrics. These methodological limitations were also evident in the study-level summaries presented in supplementary table S1, where issues such as inadequate preprocessing documentation, missing data handling and overfitting concerns were repeatedly identified across studies. The analysis domain was the most variable: only two studies [22, 24] explicitly described methods to control overfitting, such as early stopping or regularisation, and just one [22] provided calibration analysis using Brier scores and graphical plots. AI-specific issues were common across the dataset. None of the studies conducted formal fairness, subgroup or external generalisability assessments. Reporting of hyperparameter optimisation, model interpretability and data provenance was inconsistent, reducing overall reproducibility. The wearable-based studies [20, 24] demonstrated innovative use of continuous sensor data but lacked detailed documentation of feature stability and drift correction. Conversely, EHR-based models tended to show stronger internal validity yet remained susceptible to data imbalance and variable quality across health systems [4, 10, 11, 13]. To provide a structured overview, table 1 summarises the PROBAST domain ratings for each included study.
Most investigations achieved low risk for participants and outcomes, moderate risk for predictors and moderate-to-high risk for the analysis domain. These findings highlight the need for standardised reporting frameworks such as TRIPOD-AI and PROBAST-AI [8, 10, 13], which could enhance transparency, comparability and clinical interpretability of AI models in respiratory medicine.
Summary of findings
Across the five included studies, multimodal AI frameworks demonstrated promising capability in predicting COPD exacerbations, though methodological variability limited direct comparability. Reported discriminative performance was moderate to high, with AUC values ranging from 0.73 to 0.92 and accuracies between 0.86 and 0.92. The wearable-based model of Wu et al. (2021) [20] achieved the highest short-term predictive sensitivity (0.94) and F1 score (0.92), reflecting the potential of continuous physiological monitoring. However, the wearable-based models showing the highest short-term sensitivity were also derived from the smallest cohorts and lacked external validation, which limits the certainty with which these results can be interpreted. Singh et al. (2022) [22] and Jeon et al. (2025) [24] incorporated multimodal fusion of spirometry, clinical and sensor data, improving long-term risk estimation compared with unimodal approaches. Specificity values exceeded 0.80 in three studies, while calibration performance was seldom reported, with only Singh et al. (2022) including a Brier score (0.082) and decision curve analysis.
The synthesis of results, presented in supplementary Table S1, highlights consistent trends towards enhanced performance when combining heterogeneous data sources and employing deep learning or ensemble methods. However, as illustrated in supplementary Figure S1, methodological rigour and generalisability remain constrained by the limited use of external validation and inconsistent documentation of preprocessing steps.
As shown in supplementary Figure S1, model performance varied considerably among the included studies. Reported AUC values ranged from 0.73 to 0.92, reflecting moderate to strong discriminative ability across designs. The wearable-based model by Wu et al. (2021) [20] demonstrated the highest overall performance, with an AUC of 0.79, accuracy of 0.92, sensitivity of 0.94, specificity of 0.90 and an F1 score of 0.92. Singla et al. (2021) [21] reported an AUC of 0.73, accuracy of 0.86, sensitivity of 0.84 and specificity of 0.82, showing balanced yet slightly lower discrimination in a clinical-imaging framework. The model developed by Singh et al. (2022) [22] achieved an AUC of 0.73, a positive predictive value of 0.71, a negative predictive value ranging from 0.83 to 0.88 and a Brier score of 0.082, with decision curve analysis supporting its potential clinical utility. Hussain et al. (2021) [23] presented the strongest results overall, with an AUC of 0.92, accuracy of 0.91, specificity around 0.90 and an F1 score of 0.91, indicating highly consistent classification performance. Finally, Jeon et al. (2025) [24] achieved an AUC of 0.895 using a temporally split internal–external validation design with bootstrapped confidence intervals, highlighting improved methodological rigour. Collectively, the findings presented in supplementary Figure S1 confirm the moderate-to-high discriminative capacity of multimodal AI models for predicting COPD exacerbations, while also revealing that reporting remains inconsistent and true external validation is still uncommon. Overall, these findings confirm that multimodal AI architectures can effectively capture complex clinical and physiological patterns associated with COPD exacerbation risk, offering opportunities for individualised prediction and proactive management. Yet, the field remains in an exploratory phase, with reproducibility and standardisation still evolving [6, 8, 11, 13, 15, 26].
Discussion
This systematic review synthesised current evidence on the application of AI and ML approaches for predicting exacerbations in COPD using multimodal data. The integration of diverse data sources, such as EHRs, spirometry, imaging and wearable sensor streams, reflects an expanding effort to move beyond conventional risk prediction tools towards continuous, personalised monitoring. Across the five included studies, AI models demonstrated moderate-to-high discriminative performance, with AUC values ranging from 0.73 to 0.92. These findings align with prior reviews emphasising that multimodal data fusion enhances predictive accuracy compared with unimodal clinical or demographic models [1, 2]. While multimodal designs appeared to outperform unimodal approaches, it remains difficult to determine how much of this advantage reflects true data fusion rather than differences in sample size, sensor availability or model complexity. The strongest-performing studies also incorporated richer physiological signals or larger EHR feature sets, suggesting that these study-specific factors may have contributed to the observed performance rather than multimodality alone. Although none of the included studies incorporated pollution metrics, biomarker panels or external environmental signals, work from adjacent fields has shown that these modalities can meaningfully enhance machine-learning performance [31–33], suggesting potential avenues for broadening multimodal COPD prediction in future research.
The predictive performance observed in the reviewed studies suggests that real-time physiological monitoring, when combined with structured clinical data, can improve the early detection of exacerbations. Notably, the wearable-based framework by Wu et al. (2021) [20] achieved high sensitivity and accuracy in short-term prediction, indicating the potential of sensor-derived continuous signals for proactive disease management. However, the wearable-based models showing the highest short-term sensitivity were also derived from the smallest cohorts and lacked external validation, which limits the certainty with which these results can be interpreted. Similarly, the convolutional architecture employed by Jeon et al. (2025) [24] effectively modelled complex temporal dependencies across heterogeneous datasets. Such results support the hypothesis that deep-learning models appear capable of capturing complex nonlinear interactions between physiological and behavioural patterns, which traditional regression-based models often fail to detect. At the same time, these models appear to complement established clinical risk factors by incorporating dynamic physiological and behavioural signals that traditional predictors such as prior exacerbation history or chronic bronchitis cannot capture [4, 5, 14]. However, this interpretation should be considered tentative, as it is derived from indirect comparisons across heterogeneous studies rather than direct head-to-head evaluations of multimodal AI against traditional modelling approaches.
However, this review also highlights several methodological shortcomings that limit generalisability and reproducibility. Only one study performed temporal validation, and none incorporated fully external datasets for testing, an omission that may lead to optimistic bias in model performance. This limited use of external or temporally independent validation further strengthens the view that the current evidence base should be regarded as preliminary. The reliance on internal resampling creates a risk of overestimating performance, and without replication across independent cohorts, it remains uncertain whether these models would generalise to broader COPD populations. Additionally, regional differences in COPD treatment may affect model transportability. Variability in the use of biologics and in LABA/LAMA/ICS prescribing patterns can influence exacerbation rates, meaning that models developed in one therapeutic context may not perform similarly in others. Furthermore, calibration analyses were infrequently reported, with just one study providing slope and Brier score metrics. This lack of calibration reporting impedes clinical interpretation, as even highly discriminative models can perform poorly when probability estimates are miscalibrated [15, 26]. Transparent documentation of hyperparameter tuning, handling of missing data and overfitting control strategies was inconsistent across studies, underscoring the need for standardised reporting frameworks such as TRIPOD-AI and PROBAST-AI [8, 10, 13, 34]. These gaps in reporting have important implications for reproducibility and may bias the overall interpretation of model performance. When key steps such as feature preprocessing, missing data handling or hyperparameter selection are insufficiently described, it becomes difficult to determine whether observed improvements reflect true model capability or methodological artefacts.
The interpretability of AI systems remains another major challenge. Few studies described feature importance or explainability methods, despite their relevance for clinical decision-making. The absence of model explainability can hinder clinician trust and impede clinical adoption [4, 11, 35]. Furthermore, none of the reviewed works evaluated subgroup fairness or bias mitigation, a crucial omission given growing evidence that AI models may propagate health disparities if not assessed across demographic or socioeconomic strata [5, 36]. The lack of subgroup or fairness analyses makes it unclear whether these models perform consistently across different COPD populations, posing a risk of unintended disparities. To ensure equitable implementation, future research should incorporate bias auditing, fairness testing and reporting of demographic performance stratifications.
From a clinical perspective, AI-driven exacerbation prediction models could enable timely interventions, optimise resource allocation and improve patient outcomes through precision-based management. Nevertheless, before integration into practice, external validation across geographically and demographically diverse populations remains essential. The inclusion of standardised COPD outcome definitions, consistent exacerbation labelling and harmonised multimodal data pipelines would improve comparability across future studies. Moreover, pragmatic trials assessing the impact of AI-guided alerts on patient outcomes would be a critical next step towards clinical translation.
A major outstanding challenge is translating risk estimates into clinically actionable tools that integrate seamlessly into real-world care pathways. Moving from a model that generates a probability score to a system that triggers timely, interpretable alerts requires robust workflow design, clinician-facing interfaces and clear escalation protocols. In practice, this demands alignment with existing care structures, including thresholds for action, responsibility for responding to alerts and mechanisms to avoid alarm fatigue. Moreover, successful implementation depends on interoperability with EHR systems and prospective evaluation of how such tools affect patient outcomes, clinician workload and decision-making processes.
Finally, this review emphasises that the advancement of artificial intelligence in respiratory medicine must go hand in hand with rigorous validation, transparent reporting and clear interpretability of models. The variability observed among the included studies highlights the importance of establishing shared guidelines and data-sharing frameworks to promote consistency and openness in future research. As digital health technologies continue to evolve, the integration of wearable monitoring, remote patient tracking and clinical data analytics has the potential to transform COPD management from a reactive to a preventive approach. Building on the methodological gaps and recurring themes identified in this review, several research priorities are proposed to guide the next phase of work on AI-based COPD prediction. These priorities, outlined in table 4, focus on improving data harmonisation, strengthening validation practices, enhancing interpretability, and ensuring fair and reliable application of predictive models in clinical care.
In summary, while current evidence demonstrates the promise of multimodal AI for predicting COPD exacerbations, significant methodological and ethical challenges remain. Addressing these through rigorous validation, transparent reporting and equitable data integration will be essential for translating algorithmic innovation into trustworthy clinical tools.
Conclusions
This systematic review demonstrates that AI and ML models hold considerable promise for predicting exacerbations in patients with COPD using multimodal data sources. Across the five included studies, multimodal models generally demonstrated higher discriminative performance, although direct comparisons with unimodal approaches were not consistently available across datasets. Despite these encouraging findings, the review also reveals persistent methodological limitations, including scarce external validation, inconsistent calibration reporting and limited interpretability.
Future research should focus on standardised data harmonisation, transparent reporting through TRIPOD-AI and PROBAST-AI frameworks, and the inclusion of fairness and explainability analyses. Prospective multicentre validation studies supported by open data sharing and collaboration are essential to ensure that AI models achieve clinical reliability, equity and regulatory readiness. Strengthening these methodological standards will allow AI-based systems to evolve from proof of concept approaches into practical instruments for personalised COPD management and timely clinical decision support in routine care.
Limitations of this review
This review has several limitations. Although the search strategy followed the PRISMA 2020 framework and was registered in PROSPERO (CRD420251165476), the number of eligible studies was limited, reflecting the emerging nature of AI research in COPD prediction. This narrow yield appears to reflect the early stage of multimodal AI model development rather than a limitation of the search strategy, as relatively few published studies met the methodological and multimodal criteria required for inclusion. Considerable heterogeneity was observed in data modalities, outcome definitions and performance metrics, which restricted direct quantitative synthesis. Moreover, potential publication bias cannot be excluded, as studies reporting negative or nonsignificant findings may remain unpublished. Additionally, the moderate and recurrent risk of bias observed across studies may have influenced the overall certainty of the synthesised findings. Finally, the review relied on the methodological transparency of the included articles, and incomplete reporting of model development and validation details may have influenced the assessment of quality and bias.
Acknowledgments
The authors wish to thank the Medical School of the National and Kapodistrian University of Athens for providing access to academic resources and research databases that supported the completion of this systematic review. The authors also acknowledge and express appreciation to all researchers whose studies contributed to the evidence base analysed in this work. The authors acknowledge the use of Canva (Canva Pty Ltd, Sydney, Australia) for creating the graphical design of figure 1.
Footnotes
Provenance: Submitted article, peer reviewed.
Author contributions: Conceptualisation: C. Kapatais and N. Rovina; methodology: C. Kapatais, T. Karaoulani, A. Papanikolaou and N. Rovina; software: C. Kapatais and A. Bakakos; validation: T. Karaoulani, A. Papanikolaou and Evangelia Koukaki; formal analysis: C. Kapatais and E. Koukaki; investigation: C. Kapatais, A. Papanikolaou and T. Karaoulani; resources: A. Bakakos, C. Anagnostopoulou, A.I. Papaioannou and P. Bakakos; data curation: C. Kapatais; writing – original draft preparation: C. Kapatais; writing – review and editing: C. Kapatais, A.I. Papaioannou, P. Bakakos and N. Rovina; visualisation: C. Kapatais and T. Karaoulani; supervision: N. Rovina and P. Bakakos; project administration: C. Kapatais and N. Rovina. All authors have read and agreed to the published version of the manuscript.
Conflicts of interest: The authors declare no conflict of interest.
Support statement: No funding declared.
Supplementary material
Please note: supplementary material is not edited by the Editorial Office, and is uploaded as it has been supplied by the author.
Supplementary material
01420-2025.SUPPLEMENT
Data availability
No new data were created or analysed in this study. Data supporting the findings of this systematic review are available from the corresponding author upon reasonable request. Supplementary material associated with this article, including detailed study-level summaries and performance metrics (table S1 and figure S1) as well as the complete database search strategies used for the systematic review (supplementary Table S3) are provided.
References
- 1.Topalovic D, Das N, Burgel PR, et al. Artificial intelligence and machine learning in respiratory medicine. ERJ Open Res 2021; 7: 00101-2020. doi: 10.1183/23120541.00101-2020 [DOI] [Google Scholar]
- 2.Islam MM, Yang HC, Nguyen PA, et al. Performance of machine learning algorithms to predict chronic obstructive pulmonary disease exacerbations: a systematic review and meta-analysis. PLoS ONE 2021; 16: e0248127. doi: 10.1371/journal.pone.0248127 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.W orld Health Organization (WHO) . Global Health Observatory: Mortality and global health estimates. Geneva, Switzerland, WHO, 2024. Date last accessed: 27 September 2025. www.who.int/data/gho/data/themes/mortality-and-global-health-estimates [Google Scholar]
- 4.Wiens J, Saria S, Sendak M, et al. Do no harm: a roadmap for responsible machine learning for health care. Nat Med 2019; 25: 1337–1340. doi: 10.1038/s41591-019-0548-6 [DOI] [PubMed] [Google Scholar]
- 5.Rajkomar A, Dean J, Kohane I. Machine learning in medicine. N Engl J Med 2019; 380: 1347–1358. doi: 10.1056/NEJMra1814259 [DOI] [PubMed] [Google Scholar]
- 6.Wynants L, Van Calster B, Collins GS, et al. Prediction models for diagnosis and prognosis of COVID-19: systematic review and critical appraisal. BMJ 2020; 369: m1328. doi: 10.1136/bmj.m1328 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Shickel B, Tighe PJ, Bihorac A, et al. Deep EHR: a survey of recent advances in deep learning techniques for electronic health record (EHR) analysis. IEEE J Biomed Health Inform 2018; 22: 1589–1604. doi: 10.1109/JBHI.2017.2767063 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Collins GS, Reitsma JB, Altman DG, et al. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement. Ann Intern Med 2015; 162: 55–63. doi: 10.7326/M14-0697 [DOI] [PubMed] [Google Scholar]
- 9.Riley RD, Ensor J, Snell KIE, et al. Calculating the sample size required for developing a clinical prediction model. BMJ 2020; 368: m441. doi: 10.1136/bmj.m441 [DOI] [PubMed] [Google Scholar]
- 10.Sounderajah V, Ashrafian H, Rose S, et al. Developing a reporting guideline for artificial intelligence-centred diagnostic test accuracy studies: the STARD-AI protocol. BMJ Open 2021; 11: e047709. doi: 10.1136/bmjopen-2020-047709 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Kelly CJ, Karthikesalingam A, Suleyman M, et al. Key challenges for delivering clinical impact with artificial intelligence. BMC Med 2019; 17: 195. doi: 10.1186/s12916-019-1426-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Moher D, Liberati A, Tetzlaff J, et al. PRISMA Group . Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement. PLoS Med 2009; 6: e1000097. doi: 10.1371/journal.pmed.1000097 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Wolff RF, Moons KGM, Riley RD, et al. PROBAST: a tool to assess risk of bias and applicability of prediction model studies. Ann Intern Med 2019; 170: 51–58. doi: 10.7326/M18-1376 [DOI] [PubMed] [Google Scholar]
- 14.Esteva A, Robicquet A, Ramsundar B, et al. A guide to deep learning in healthcare. Nat Med 2019; 25: 24–29. doi: 10.1038/s41591-018-0316-z [DOI] [PubMed] [Google Scholar]
- 15.Van Calster B, McLernon DJ, van Smeden M, et al. Calibration: the Achilles heel of predictive analytics. BMC Med 2019; 17: 230. doi: 10.1186/s12916-019-1466-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Booth A, Clarke M, Dooley G, et al. The nuts and bolts of PROSPERO: an international prospective register of systematic reviews. Syst Rev 2011; 1: 1–9. doi: 10.1186/2046-4053-1-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Higgins JPT, Thomas J, Chandler J, et al. , eds. Cochrane Handbook for Systematic Reviews of Interventions, Version 6.3. London, Cochrane, 2022. [Google Scholar]
- 18.Ouzzani M, Hammady H, Fedorowicz Z, et al. Rayyan: a web and mobile app for systematic reviews. Syst Rev 2016; 5: 210. doi: 10.1186/s13643-016-0384-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Schardt C, Adams MB, Owens T, et al. Utilization of the PICO framework to improve searching PubMed for clinical questions. BMC Med Inform Decis Mak 2007; 7: 16. doi: 10.1186/1472-6947-7-16 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Wu CT, Li GH, Huang CT, et al. Acute exacerbation of a chronic obstructive pulmonary disease prediction system using wearable device data, machine learning, and deep learning: development and cohort study. JMIR Mhealth Uhealth 2021; 9: e22591. doi: 10.2196/22591 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Singla S, Gong M, Riley C, et al. Improving clinical disease subtyping and future events prediction through a chest CT-based deep learning approach. Med Phys 2021; 48: 1168–1181. doi: 10.1002/mp.14673 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Singh D, Hurst JR, Martinez FJ, et al. Predictive modeling of COPD exacerbation rates using baseline risk factors. Ther Adv Respir Dis 2022; 16: 17534666221107314. doi: 10.1177/17534666221107314 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Hussain A, Choi HE, Kim HJ, et al. Forecast the exacerbation in patients of chronic obstructive pulmonary disease with clinical indicators using machine learning techniques. Diagnostics (Basel) 2021; 11: 829. doi: 10.3390/diagnostics11050829 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Jeon ET, Park H, Lee JK, et al. Deep learning-based chronic obstructive pulmonary disease exacerbation prediction using flow-volume and volume-time curve imaging: retrospective cohort study. J Med Internet Res 2025; 27: e69785. doi: 10.2196/69785 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Microsoft Corporation . Microsoft Excel [computer software]. 2022. www.microsoft.com/excel
- 26.Steyerberg EW, Moons KGM, van der Windt DA, et al. Prognosis research strategy (PROGRESS) 3: prognostic model research. PLoS Med 2013; 10: e1001381. doi: 10.1371/journal.pmed.1001381 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Popay J, Roberts H, Sowden A, et al. Guidance on the Conduct of Narrative Synthesis in Systematic Reviews: A Product from the ESRC Methods Programme. Version 1. Lancaster, Lancaster University, 2006. Date last accessed: 27 September 2025. www.researchgate.net/publication/233866356_Guidance_on_the_conduct_of_narrative_synthesis_in_systematic_reviews_A_product_from_the_ESRC_Methods_Programm. doi: 10.13140/2.1.1018.4643 [DOI] [Google Scholar]
- 28.Johnson AEW, Pollard TJ, Shen L, et al. MIMIC-III, a freely accessible critical care database. Sci Data 2016; 3: 160035. doi: 10.1038/sdata.2016.35 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Roberts M, Driggs D, Thorpe M, et al. Common pitfalls and recommendations for machine learning validation in biomedical studies. Nat Commun 2021; 12: 6045. doi: 10.1038/s41467-021-26281-234663792 [DOI] [Google Scholar]
- 30.Steyerberg EW, Vickers AJ, Cook NR, et al. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology 2010; 21: 128–138. doi: 10.1097/EDE.0b013e3181c30fb2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Li L, Li J, Zhong M, et al. Nanozyme-enhanced tyramine signal amplification probe for preamplification-free myocarditis-related miRNAs detection. Chem Eng J 2025; 503: 158093. doi: 10.1016/j.cej.2024.158093 [DOI] [Google Scholar]
- 32.Liu B, Du H, Zhang J, et al. Developing a new sepsis screening tool based on lymphocyte count, international normalized ratio and procalcitonin (LIP score). Sci Rep 2022; 12: 20002. doi: 10.1038/s41598-022-16744-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Li J, Liang L, Lyu B, et al. Double trouble: the interaction of PM2.5 and O3 on respiratory hospital admissions. Environ Pollut 2023; 338: 122665. doi: 10.1016/j.envpol.2023.122665 [DOI] [PubMed] [Google Scholar]
- 34.Nagendran M, Chen Y, Lovejoy CA, et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies. BMJ 2020; 368: m689. doi: 10.1136/bmj.m689 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Ghassemi M, Oakden-Rayner L, Beam AL. The false hope of current approaches to explainable artificial intelligence in health care. Lancet Digit Health 2021; 3: e745–e750. doi: 10.1016/S2589-7500(21)00173-5 [DOI] [PubMed] [Google Scholar]
- 36.Rajkomar A, Hardt M, Howell MD, et al. Ensuring fairness in machine learning to advance health equity. Ann Intern Med 2018; 169: 866–872. doi: 10.7326/M18-1990 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Please note: supplementary material is not edited by the Editorial Office, and is uploaded as it has been supplied by the author.
Supplementary material
01420-2025.SUPPLEMENT
Data Availability Statement
No new data were created or analysed in this study. Data supporting the findings of this systematic review are available from the corresponding author upon reasonable request. Supplementary material associated with this article, including detailed study-level summaries and performance metrics (table S1 and figure S1) as well as the complete database search strategies used for the systematic review (supplementary Table S3) are provided.


