Abstract
Background
Obstructive sleep apnea (OSA) is a highly prevalent sleep disorder. Misdiagnosis might lead to several systemic conditions, including hypertension, vascular damage, and cognitive impairment. The gold-standard diagnostic tool for OSA is polysomnography, which is expensive, time-consuming, and not accessible everywhere. Artificial intelligence (AI) algorithms can facilitate diagnosis by detecting patients’ signs and symptoms. In this systematic review, we evaluated the diagnostic accuracy of AI models in detecting sleep apnea.
Methods
We searched six major databases, PubMed®, Cochrane, Web of Science, Scopus, Embase, and IEEE Xplore, using keywords related to AI and OSA. Eligible studies focused on adult populations, used in-laboratory PSG as the reference standard, and applied AI models trained on multiple clinical features. Reviews, pediatric studies, and articles lacking accuracy metrics were excluded. From the included articles, data were extracted regarding patients and datasets, type of AI model applied, accuracy report, and explainability of the AI model. A risk of bias assessment was done using the QUADAS-2 checklist.
Results
Thirteen studies were included in our final analysis. The AI models consisted of deep learning, machine learning, and hybrid models with various architectures. The reported accuracy of studies ranged from 67.03 to 98.6%, with the highest being related to hybrid and deep learning models. Risk of bias assessment showed that 7 of the studies had a low risk of bias, indicating high reliability.
Conclusions
AI-driven models, particularly deep learning and hybrid architectures, show significant promise in diagnosing obstructive sleep apnea. However, challenges such as transparency, explainability, and variability in performance necessitate diverse training datasets to improve generalizability for clinical adoption.
Registration of systematic reviews
The protocol of this systematic review was registered in PROSPERO (CRD42023453789), available from: https://www.crd.york.ac.uk/prospero/display_record.php?ID=CRD42023453789.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12911-025-03129-x.
Keywords: Artificial intelligence, Obstructive sleep apnea, Diagnostic accuracy, Deep learning, Machine learning
Background
Obstructive sleep apnea (OSA), also called obstructive sleep apnea syndrome (OSAS), is a potentially severe sleep disorder characterized by repeated episodes of partial or complete obstruction of the upper airway during sleep [1]. These episodes manifest as significant reductions in airflow or complete pauses in breathing, known as hypopneas (partial airway obstruction) and apneas (complete airway obstruction) [2]. The etiology of OSA involves a multifactorial interplay of anatomical susceptibilities and neuromuscular control of the airway [3]. Risk factors include obesity, male gender, advancing age, and genetic predispositions, among others [4]. The pathogenesis of OSA includes intermittent hypoxia-induced oxidative stress and systemic inflammation [5], which result from disrupted sleep architecture and increased sympathetic nervous system activity [6]. When OSA remains untreated, the clinical manifestations range from daytime fatigue to cognitive impairment [7], increasing the risk of uncontrolled systemic hypertension, cardiovascular disease, diabetes, chronic kidney disease, and vehicular accidents [8]. These significant consequences underline the critical need for accurate diagnosis and effective management.
The gold standard for diagnosis of OSA is an overnight polysomnography (PSG). PSG involves monitoring several body functions during sleep, including brain activity (as measured by EEG), eye movement, muscle activity, heart rate, respiratory effort, and blood oxygen levels. This comprehensive approach allows for detailed observation and analysis of the sleep stages and identification of any disturbances related to breathing or other vital signs [9]. However, the PSG is resource-intensive and requires specialized facilities and personnel [10]. Thus, its complexity and cost highlight the necessity for more accessible diagnostic modalities [11].
Artificial intelligence (AI) technologies, particularly machine learning (ML) and deep learning (DL), are transforming diagnostic approaches in sleep medicine. ML algorithms can learn from structured clinical data using handcrafted features, while DL models extract patterns directly from raw signals such as airflow, oximetry, or electrocardiogram (ECG) [12]. This ability to autonomously recognize subtle spatial and temporal patterns makes DL especially suited for detecting OSA-related disruptions [13]. By leveraging large datasets, these models can achieve high diagnostic accuracy and offer scalable solutions that augment traditional assessment tools.
Various automated approaches aim to simplify OSA diagnosis and overcome the limitations of PSG. The availability of simplified, automated, and reliable alternative tools to PSG would enable the accurate diagnosis of OSA, which would have several other advantages for patients. For instance, less need for too many sensors would help improve patient comfort. Moreover, the tool could be available for home testing, therefore reducing the long waiting lists within healthcare for patients to undergo a PSG [14]. Additionally, an automated tool would dramatically decrease the effort and time needed for the specialists to analyze overnight physiological signals. Consequently, the advantages mentioned would facilitate patients’ access to treatment while maintaining appropriate accuracy in diagnosis [15].
Many studies have evaluated the accuracy of AI models in diagnosing OSA; however, differences in levels of accuracy, AI models used, and various performance metrics have yielded conflicting results.
Therefore, it is necessary to systematically review and evaluate the performance of current AI models for diagnosing OSA.
This systematic review evaluated the diagnostic accuracy and performance of deep learning-based AI models in identifying and detecting OSA in adult patients.
Methods
The protocol of this systematic review was registered in PROSPERO (CRD42023453789), available from: https://www.crd.york.ac.uk/prospero/display_record.php?ID=CRD42023453789. The study was conducted following the Preferred Reporting Items for Systematic Review and Meta-Analyses (PRISMA) [16].
Eligibility Criteria
Studies were selected based on the eligibility criteria and consensus from all the authors.
To be included, the studies needed to (A) be original investigations assessing the diagnostic performance of AI algorithms in identifying OSA among adult populations; (B) original research designs, including randomized clinical trials, cross-sectional, and case-control studies; (C) utilize in-laboratory PSG as the reference standard; (D) incorporate multiple clinical factors to train the AI models; (E) be published in English language; and (F) be published in the last 5 years. This time window of the last 5 years was chosen to ensure a comprehensive review of the most recent advancements in AI algorithms. Conversely, studies were excluded if they (A) did not report accuracy metrics; (B) were review articles, meta-analyses, or conference abstracts; (C) were not conducted in human subjects; (D) involved pediatric patients (< 18 years); (E) lacked full-text availability; (F) were published in languages other than English; (G) relied solely on one-organ symptoms and signs, such as cardiovascular or respiratory data, to identify OSA; or (H) evaluated aspects of OSA other than diagnosis (such as management, adherence and treatment outcomes).
Information Sources and Bibliography Search
A comprehensive literature search was conducted on October 15, 2023, across six databases (PubMed®, Cochrane, Web of Science, Scopus, Embase, and IEEE Xplore) to identify studies evaluating the diagnostic accuracy of artificial intelligence (AI) systems (Intervention) compared to polysomnography (PSG) (Comparator) in diagnosing obstructive sleep apnea (OSA) in adults (Population). The research question guiding this review was: “What is the diagnostic accuracy of AI-based systems compared to the gold standard PSG in detecting OSA in adults?”
Search Strategy
We included keywords for three main domains, namely AI, diagnosis, and obstructive sleep apnea, with AND Boolean operators between the keywords. Supplementary Table 1 in Additional File 1 presents the search query for each database.
Selection Process
To ensure a comprehensive search, we conducted a two-phase screening process. Rayyan Web Platform [17] was used to screen articles and remove duplicates. Initially, a team of four collaborators screened the titles and abstracts of potential articles using predefined eligibility criteria. In the second phase, two additional reviewers independently reassessed the articles labeled as ‘included’ or ‘maybe’ to confirm their relevance and improve screening accuracy. Final inclusion decisions were made by consensus among the review authors. This collaborative approach helped to minimize bias and ensure the inclusion of relevant studies.
Data Collection Process and Data Items
We categorized models as machine learning, deep learning, or hybrid based on their architectural structure. ML models included traditional classifiers that relied on engineered features [18], while DL models referred to multilayer neural networks capable of learning directly from raw data. Hybrid models were defined as those integrating both ML and DL components within a single diagnostic pipeline [19]. Models composed entirely of DL or ML components were classified under their respective category to maintain consistency.
The diagnostic accuracy of AI systems was evaluated by extracting the overall accuracy (correct prediction rate) from all included studies. Accuracy was selected as the primary performance metric due to the substantial heterogeneity in model types, input features, and study designs, as well as its consistent reporting across all studies. Additional performance metrics, such as the F1 score (harmonic mean of precision and recall/sensitivity), sensitivity (true positive rate, TP), specificity (true negative rate, TN), false negatives (FN), false positives (FP), and the area under the curve (AUC), were also extracted when reported.
To evaluate model performance, we extracted an aggregated accuracy metric from each study, calculated as (TP + TN) / (TP + TN + FP + FN). These values were summarized in tables.
![]() |
In addition to primary performance metrics, we extracted detailed characteristics including dataset type, participant demographics, input signals, and model architectures.
Risk of Bias
The QUADAS-2 (Quality Assessment Tool for Diagnostic Accuracy Studies - Version 2) checklist was used to evaluate the quality of diagnostic accuracy studies [20]. This tool evaluates risk of bias across four domains: (1) patient selection, which assesses whether participants were enrolled in a way that avoids bias; (2) index test, which evaluates whether the test under investigation was conducted and interpreted without knowledge of the reference standard; (3) reference standard, which examines the reliability and appropriateness of the diagnostic benchmark used; and (4) flow and timing, which assesses whether all participants received the same reference standard and whether there were delays that could affect results. In addition, the first three domains are examined for concerns about applicability (e.g., how well the study matches the review question). Each domain was rated as having a low, high, or unclear risk of bias.
Results
Study Selection
The search yielded 2907 articles. After removing duplicates (N = 1507), 1400 studies were initially screened. Following full-text evaluation, 13 studies were included for the qualitative analysis. The reasons for exclusion and the number of studies in each category are summarized in the PRISMA flowchart (Fig. 1).
Fig. 1.
PRISMA flowchart of the study
Study Characteristics
The qualitative analysis was performed on 13 included studies, including a total of 12,631 participants. Three studies (23.08%) did not report the ratio of male/female [21–23], and three did not report mean age [21, 22, 24]. Among the remaining 10 studies (N = 4708), there were a total of 1485 females (31.54%) and 3223 males (68.45%), and the overall mean age was 45.82 ± 14.21.
We divided the models into three categories: DL, ML, and hybrid. The number of studies in the DL, ML, and hybrid models was 4, 6, and 3, respectively. There were a total of 21 types of algorithms developed or used in the studies, consisting of 10 ML algorithms (Logistic Regression, Random Forest, Support Vector Machine, AdaBoost Learning, Gradient Boosting, K Nearest Neighbor, Neural Network, CatBoost Classifier, Extra Trees Classifier and Decision Tree), 6 DL algorithms (Convolutional Neural Network, Long Short Term Memory, Bi-directional Long Short Term Memory, ResNet101, DeepsleepNet, CMS-2-Net), and five hybrid algorithms (FNN, Complex Tree, RuBoosted Trees, GRU and ML Meta-learners). All studies reported the accuracy of the model, while other performance metrics, including sensitivity, specificity, AUC, and F1 score, were reported in 12, 9, 7, and 3 of the studies, respectively. The reported accuracy of AI systems in diagnosing OSA ranged from 67.03 to 98.60%. Table 1 presents the AI model type, total number of subjects, and reported midpoint accuracy for each study. (The list of abbreviations is provided in Additional File 1, Table 2).
Table 1.
Model type, sample size and reported accuracy per study
| 1st Author | Algorithms used | Accuracy (Midpoint, %) | Sample size | Country |
|---|---|---|---|---|
| J et al. [22] | Bi-LSTM, ResNet 101, DeepSleepNet | 83.79% | 7745 a | South Korea |
| Leong [25] | LR, RF, SVM, AdaB, GB, NN, kNN | 87.9% | 2996 | Singapore |
| Zhang [23] | CMS2-Net, CNN | 73.5% | 128 | China |
| Arslan [21] | DNN, RNN, LSTM, GRU, ML Meta-Learner | 88.7% | 50 | Turkey |
| Strumpf [26] | CNN | 88% | 84 | USA |
| Zhuang [24] | RF | 95.14% | 10 | China |
| Moussa [27] | RF, SVM, kNN, LR | 82.70% | 150 | UAE |
| Hafezi [28] | CNN, LSTM | 83% | 69 | Canada |
| Chen [29] | Top 5: RF, LR, CATBOOST, ET, GB b | 69% | 653 | China |
| Elwali & Moussavi [30] | RF | 80.15% | 145 | Canada |
| He [31] | CNN, GB | 81.05% | 393 | China |
| Z. Zhang [24] | DT, RF, kNN, SVM | 82.90% | 27 | USA |
| Li [32] | FNN, SVM, Complex Tree, RUSBoosted Trees, LR | 93.10% | 181 | China |
a: Final image-based dataset
b: A total of 18 models were used, but 13 others were in the supplementary table and not accessible
Table 2 summarizes the datasets and physiological inputs used to train or evaluate each AI model.
Table 2.
Dataset and input signal types across included studies
| 1st Author | Dataset | Input features |
|---|---|---|
| J et al. [22] | Korean National Information Society Agency | Full PSG |
| Leong [25] | Local Sleep Medicine Database | Full PSG, Demographic & Anthropometric Data |
| Zhang [23] | Local Dataset (Peking University Sixth Hospital) + Public (Dream Open Dataset) | EEG, EMG, EOG, ECG |
| Arslan [21] | Local Dataset (Yozgat Bozok University, Department of Chest Diseases Sleep Laboratory) | Full PSG |
| Strumpf [26] | University Hospitals Cleveland Medical Center Bolwell and Beachwood Sleep Labs | Full PSG |
| Zhuang [24] | Sleep Center of Huai’an First People’s Hospital in Jiangsu Province | PSG, Physiological Data (radar signals, vital signs) |
| Moussa [27] | American Center for Psychiatry and Neurology, Stanford Technology Analytics and Genomics in Sleep (STAGES) | EEG, ECG, Breathing Signals |
| Hafezi [28] | Local Dataset (Sleep Laboratory at the Toronto Rehabilitation Institute) |
Chest & abdominal movements, respiratory inductance plethysmography, airflow by nasal pressure cannula, SpO2,Tracheal Movements /Apnea Hypopnea index (AHI) |
| Chen [29] | Local Dataset (Not Specified) | Clinical Data (including PSG and PM) + Craniofacial Images |
| Elwali & Moussavi [30] | Local Dataset (Sleep Disorders Center in Misericordia Health Centre (Winnipeg, Canada)) | Full PSG |
| He [31] | Sleep Laboratories at the Department of Otolaryngology Head and Neck Surgery, Beijing Tongren Hospital (Beijing, China) | Craniofacial Images + PSG |
| Z. Zhang [24] | Weill Cornell Center for Sleep Medicine (New York, USA) | PSG, SpO2, airflow/bed-integrated radio-frequency sensor by near-field coherent sensing |
| Li [32] | Sleep Medicine Center of Beijing Tongren Hospital (Bei Jing Shi, China) | PSG, BMI, ECG, SpO2 |
Table 3 summarizes other findings (such as sensitivity and specificity) for the included studies.
Table 3.
Summary of other findings for the included studies
| 1st Author | AUC | Specificity(%) | Sensitivity(%) | F1 Score | Other |
|---|---|---|---|---|---|
| J et al. [22] | N/A | N/A | N/A |
weighted F1 score: 80.66–86.54 macro F1 score: 80.68–83.60 |
- |
| Leong [25] | 0.848–0.928 | 40.4%-60.6% | 90.9%-97.4% | N/A | - |
| Zhang [23] | N/A | N/A | 52%-69% | 53–69 |
precision: 64–76% Kappa: 56–73% |
| Arslan [21] | 0.99 | N/A | 95.76%-91.70% | N/A | - |
| Strumpf [26] | 0.92–0.95 | 88%-96% | 83%-95% | N/A | - |
| Zhuang [24] | N/A | 97.32% | 72.60% | N/A | - |
| Moussa [27] | 0.79–0.85 | 83.69%-73.17% | 69.35%-96.44% | N/A | - |
| Hafezi [28] | N/A | 36%-94% | 67%-98% | N/A | - |
| Chen [29] | 0.72–0.76 | N/A | 68%-75% | N/A | - |
| Elwali & Moussavi [30] | N/A | 63.9%-100% | 60.6%-94.7% | 0.66–0.92 | - |
| He [31] | 0.747–0.963 | 55.8%-96.7% | 81.8%-98.8% | N/A | - |
| Z. Zhang [24] | N/A | 72.9%-89.1% | 56.7%-74.3% | N/A | - |
| Li [32] | 0.92–0.98 | 76.0% − 93.9% | 89.0% − 98.6% | N/A | - |
Additional File 2 provides full study-level details including dataset, input type, sample size, model architecture, accuracy range, sex distribution, and data split.
Risk of Bias Assessment
Seven of the included studies were judged to have a low risk of bias across all domains, indicating overall methodological reliability. The remaining studies raised concerns primarily in the patient selection domain: Arslan [21], Moussa et al. [27], and Z. Zhang et al. [24] showed unclear risk, while Zhuang et al. [33] had a high risk due to unclear inclusion criteria.
All studies showed low applicability concerns in the index test and reference standard domains. However, Leong et al. [25], Chen et al. [29], and Z. Zhang et al. [24] raised unclear concerns in the patient selection domain, typically due to limited information about sampling methods or study population characteristics. Full assessments are presented in Table 4; Figs. 2 and 3.
Table 4.
QUADAS-2 results of the included studies
Fig. 2.
Diagram of risk of bias assessment
Fig. 3.
Diagram of applicability concern
Discussion
Sleep Disorders Epidemiology and AI in Diagnosis
Obstructive sleep apnea affects an estimated one billion adults worldwide, with 85–95% of cases remaining undiagnosed. Left untreated, OSA is associated with a wide range of complications, including cardiovascular and metabolic disorders, neurocognitive dysfunction, and a diminished quality of life. The burden of undiagnosed cases underscores the critical need for accurate and accessible diagnostic tools [8, 10].
Although PSG remains the gold standard for diagnosing OSA, its widespread use is hindered by high costs, the need for specialized equipment and trained personnel, and the requirement for overnight monitoring in dedicated sleep laboratories. These barriers are especially pronounced in low-resource settings, where access to sleep medicine infrastructure is limited. Additionally, PSG can be inconvenient and uncomfortable for patients, potentially impacting sleep quality and diagnostic yield [34].
As a result, many individuals with OSA never receive a diagnosis, often due to the condition’s nonspecific symptoms, with only around 20% of affected individuals reporting them to healthcare providers [35]. Against this backdrop, artificial intelligence offers promising solutions by enabling automated models that can interpret complex physiological data. These tools have the potential to augment or even replace conventional diagnostic methods, improving early detection and expanding access to care, particularly in underserved populations [36].
AI Models Diversity
The 13 studies included in this review utilized a wide range of AI methodologies, including six traditional ML models, four DL architectures, and three hybrid approaches. ML models, such as Support Vector Machines (SVMs) and Decision trees, rely on handcrafted features and require substantial domain expertise for optimal performance. They tend to be more interpretable and computationally efficient but may struggle with complex pattern recognition.
In contrast, DL models, including CNNs and RNNs, automatically learn multi-level features directly from raw physiological data. This makes them well-suited for modeling the complex temporal and spatial patterns characteristic of OSA diagnostics [37]. While DL approaches typically yield higher diagnostic accuracy, they also require large datasets and considerable computational resources to minimize overfitting [38].
Hybrid models that combine ML and DL components attempt to leverage the strengths of both strategies, offering a balance between adaptability, performance, and interpretability.
Accuracy of AI Models
The diagnostic accuracy of AI models for detecting OSA in the included studies ranged from 67.03 to 98.60%, with a median value of 89.66%. Notably, hybrid models, particularly those combining Convolutional and Recurrent Neural Networks (CNN-RNN), demonstrated the highest diagnostic performance. For example, the model developed by Arslan et al. reported the highest accuracy at 98.60% [21], followed closely by Moussa et al. (98.36%) and Li et al. (97.80%) [27, 32].
These multi-layered hybrid architectures appear particularly effective in capturing both spatial and temporal characteristics of physiological signals relevant to OSA. By integrating feature extraction with sequential pattern analysis, they leverage the complementary strengths of CNNs and RNNs, resulting in enhanced predictive performance [39].
Code Availability
Five of the thirteen studies shared links to their open-access code. The remaining eight articles did not specify whether their code was available, and no further attempts were made to contact the authors regarding code sharing. This gap underscores the ongoing challenges in AI research regarding reproducibility and open science practices.
Comparison with other Similar Articles
A recent systematic review focusing on automated sleep apnea detection using physiological signal data emphasized the strength of deep learning models in capturing complex spatial and temporal patterns [40]. That review also reported improved accuracy in hybrid architectures, consistent with our findings.
Similarly, a systematic review of pediatric OSA detection showed that traditional ML models, while slightly less accurate than DL models, offered advantages such as lower computational demands and easier implementation in low-resource settings [15].
Overall, our review supports the emerging consensus that hybrid and deep learning models outperform traditional ML in terms of diagnostic accuracy. However, these gains come at the cost of increased data requirements and computational complexity. Notably, the diagnostic performance varied widely across studies, reinforcing the importance of external validation in diverse clinical populations before these models can be deployed at scale.
Clinical significance
Integrating AI into clinical practice presents a number of practical and ethical challenges. While AI systems offer objective analyses of physiological data, their outputs may be perceived as opaque or rigid, especially when the underlying algorithms are not transparent. This can lead to skepticism among clinicians and patients, potentially hindering adoption. For AI to be trusted in real-world settings, its decisions must be explainable. Clinicians need to understand not just the output, but the reasoning behind it, making explainable AI (XAI) frameworks critical to clinical integration [41].
Transparency remains a persistent issue. In our review, only 5 out of 13 studies provided access to source code, limiting reproducibility and independent validation. Moreover, AI models often require large training datasets and high computational power, constraints that may limit their implementation in smaller clinics or low-resource environments. Biases in training data can also affect diagnostic accuracy across demographic groups, raising important concerns about fairness and equity.
A recent meta-analysis evaluating wearable AI technologies reported a pooled diagnostic accuracy of 86.9% in detecting apneic events. While these devices offer real-time monitoring and accessibility benefits, their performance was deemed insufficient for standalone clinical use, reinforcing the need for AI tools to function as adjuncts to standard diagnostic workflows [42].
Limitations
Future research should focus on large-scale, multicenter studies, the development of publicly accessible datasets, and standardizing methodologies to validate AI models across diverse populations. Including diverse patient demographics and implementing standardized methodologies will improve AI model generalizability, ensuring that diagnostic tools meet the needs of various patient groups.
Conclusion
AI-driven models, especially DL and hybrid architectures, demonstrate significant potential for diagnosing obstructive sleep apnea, with reported accuracies ranging from 67.03 to 98.60%. Hybrid models appear particularly effective at capturing OSA-specific physiological patterns due to their ability to integrate spatial and temporal analysis.
While these models outperform traditional machine learning approaches in terms of diagnostic accuracy, the latter still hold value in settings with limited computational resources, given their lower complexity and interpretability.
However, before AI tools can be widely adopted in clinical sleep medicine, several barriers must be addressed. These include the lack of model transparency, limited explainability, and variability in diagnostic performance across patient populations. Future research should focus on developing XAI systems, standardizing validation methods, and ensuring consistent performance across diverse demographic and clinical settings. Addressing these challenges will be essential for translating AI’s technical capabilities into meaningful clinical impact.
Supplementary Information
Below is the link to the electronic supplementary material.
Supplementary Material 1: Additional File 1: Data: Supplementary Table 1: Search query for each database, Supplementary Table 2: List of Abbreviations.
Supplementary Material 2: Additional File 2: Data: 1st Author, Source of Data (EEG, ECG, airflow, SpO2, etc.), Midpoint (percentage), Accuracy range (percentage), Number of patients, Train/Validation/Test, Age (Mean±SD), Male/Female, Note (Model Architecture (e.g., DL, ML or Hybrid and Code Availability), Country, Dataset.
Acknowledgements
Not applicable.
Abbreviations
- AdaB
Adaptive Boosting
- AI
Artificial Intelligence
- AUC
Area Under the Curve
- AUROC
Area Under the Receiver Operating characteristic Curve
- Bi-LTSM
Bidirectional Long Short-Term Memory
- CATBOOST
CatBoost Classifier
- CMS2-Net
Co-attention Meta Sleep Staging Network
- CNN
Convolutional Neural Network
- DL
Deep Learning
- DNN
Deep Neural Network
- DT
Decision Tree
- ECG
Electrocardiography
- EEG
Electroencephalography
- EMG
Electromyography
- EOG
Electrooculography
- ET
Extra Trees Classifier
- FN
False Negative
- FP
False Positive
- GB
Gradient Boosting
- GRU
Gated Recurrent Unit
- kNN
K Nearest Neighbor
- LR
Logistic Regression
- LSTM
Long Short-Term Memory
- ML
Machine Learning
- N/A
Not Available
- NN
Neural Network
- OSA
Obstructive Sleep Apnea
- PSG
Polysomnography
- RF
Random Forest
- RNN
Recurrent Neural Network
- SVM
Support Vector Machine
- TN
True Negative
- TP
True Positive
- UAE
United Arab Emirates
- USA
The United States of America
- XAI
Explainable Artificial Intelligence
Author Contributions
SH: Project administration, Conceptualization, Data Curation, Investigation, Writing - Original Draft, Writing - Review & Editing. LS: Project administration, Conceptualization, Supervision, Writing - Review & Editing, Validation. MJ: Investigation, Data Curation, Writing - Original Draft, Writing - Review & Editing, Formal analysis. SA: Data Curation, Investigation, Resources. JI: Investigation, Data Curation, Resources. AS: Investigation, Data Curation, Resources. JB: Investigation, Data Curation, Resources. ZH: Investigation, Data Curation, Resources. AK: Investigation, Supervision, Writing - Review & Editing, Validation. All authors reviewed the manuscript.
Funding
No funding was received for the preparation or publication of this article.
Data Availability
Data is provided within the manuscript or supplementary information files.
Declarations
Ethics approval and consent to participate
The study was conducted following the Preferred Reporting Items for Systematic Review and Meta-Analyses (PRISMA) [16] and the ethical principles of the Declaration of Helsinki.
Consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Arnold J, Sunilkumar M, Krishna V, Yoganand SP, Kumar MS, Shanmugapriyan D. Obstructive sleep apnea. J Pharm Bioallied Sci. 2017 Nov;9(Suppl 1):S26–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Wenner J, Cheema R, Ayas N. Clinical manifestations and consequences of obstructive sleep apnea. Lippincott Williams & Wilkins. 2009 Mar;29(2):76–83. [DOI] [PubMed]
- 3.Demirgüneş DD, Eroğul O, Akçam T, Telatar Z. Analysis of respiration, oxygen saturation and acoustic signals of snoring patients. 2009 May.
- 4.Al Lawati NM, Patel SR, Ayas NT, Epidemiology. Risk factors, and consequences of obstructive sleep apnea and short sleep duration. Prog Cardiovasc Dis. 2009 Jan;51(4):285–93. [DOI] [PubMed] [Google Scholar]
- 5.Pham LV, Schwartz AR. The pathogenesis of obstructive sleep apnea. J Thorac Dis. 2015 Aug;7(8):1358–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Khayat R, Patt B, Hayes D. Obstructive sleep apnea: the new cardiovascular disease. Part I: obstructive sleep apnea and the pathogenesis of vascular disease. Heart Fail Rev. 2009 Sept 1;14(3):143–53. [DOI] [PMC free article] [PubMed]
- 7.Eckert DJ, Malhotra A. Pathophysiology of adult obstructive sleep apnea. Proceedings of the American Thoracic Society. 2012 Dec 20. [DOI] [PMC free article] [PubMed]
- 8.Melamed KH, Goldhaber SZ. Obstr Sleep Apnea Circulation. 2015;132(6):e114–6. [DOI] [PubMed] [Google Scholar]
- 9.Polysomnography. and Other sleep studies. Springer Publishing Company; 2023.
- 10.Park JG, Ramar K, Olson EJ. Updates on Definition, Consequences, and Management of Obstructive Sleep Apnea. Mayo Clinic Proceedings. 2011 June 1;86(6):549–55. [DOI] [PMC free article] [PubMed]
- 11.Keenan SA. Chapter 3 an overview of polysomnography. In: Guilleminault C, editor. Handbook of clinical neurophysiology. Handbook of clinical neurophysiology. Elsevier. 2005;6:33–50.
- 12.Belk RW, Belanche D, Flavián C. Key concepts in artificial intelligence and technologies 4.0 in services. Springer Sci + Bus Media. 2023;17(1):1–9. [Google Scholar]
- 13.Pesapane F, Codari M, Sardanelli F. Artificial intelligence in medical imaging: threat or opportunity? Radiologists again at the forefront of innovation in medicine. Eur Radiol Experimental. 2018 Oct 24;2(1):35. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Lachenmeier W, Lachenmeier DW. Home monitoring of oxygen saturation using a Low-Cost wearable device with haptic feedback to improve sleep quality in a lung cancer patient: A case report. Multidisciplinary Digit Publishing Inst. 2022 Mar;7(2):43–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Gutiérrez-Tobal GC, Álvarez D, Kheirandish-Gozal L, Del Campo F, Gozal D, Hornero R. Reliability of machine learning to diagnose pediatric obstructive sleep apnea: systematic review and meta-analysis. Pediatr Pulmonol. 2022 Aug;57(8):1931–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Moher D, Liberati A, Tetzlaff J, Altman DG, The PRISMA group. Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement. PLoS Med. 2009;6(7):e1000097. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Ouzzani M, Hammady H, Fedorowicz Z, Elmagarmid A. Rayyan—a web and mobile app for systematic reviews. Syst Reviews. 2016;5(1):210. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Jordan MI, Mitchell TM. Machine learning: trends, perspectives, and prospects. Science. 2015;349(6245):255–60. [DOI] [PubMed] [Google Scholar]
- 19.Goodfellow I, Bengio Y, Courville A, Bengio Y. Deep learning. MIT press Cambridge. 2016;1.
- 20.Whiting PF, Rutjes AW, Westwood ME, Mallett S, Deeks JJ, Reitsma JB, et al. QUADAS-2: A revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529–36. [DOI] [PubMed] [Google Scholar]
- 21.Arslan RS. Sleep disorder and apnea events detection framework with high performance using two-tier learning model design. PeerJ Comput Sci. 2023;9:e1554. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.W JJ, Jg Y, Dk LDKYW. K, Standardized image-based polysomnography database and deep learning algorithm for sleep-stage classification. Sleep. 2023 Dec 11;46(12). [DOI] [PubMed]
- 23.Zhang C, Yu W, Li Y, Sun H, Zhang Y, De Vos M. CMS2-Net: Semi-supervised sleep staging for diverse obstructive sleep apnea severity. IEEE J Biomed Health Inf. 2022 July;26(7):3447–57. [DOI] [PubMed]
- 24.Zhang Z, Conroy T, Krieger A, Kan E. Detection and prediction of sleep disorders by Covert Bed-Integrated RF sensors. IEEE Trans Biomed Eng. 2022 Jan 1;1–11. [DOI] [PubMed] [Google Scholar]
- 25.Leong ZH, Loh SRH, Leow LC, Ong TH, Toh ST. A machine learning approach for the diagnosis of obstructive sleep Apnoea using oximetry, demographic and anthropometric data. Singap Med J. 2023 May 2. [DOI] [PMC free article] [PubMed]
- 26.Strumpf Z, Gu W, Tsai CW, Chen PL, Yeh E, Leung L, et al. Belun ring (Belun sleep system BLS-100): deep learning-facilitated wearable enables obstructive sleep apnea detection, apnea severity categorization, and sleep stage classification in patients suspected of obstructive sleep apnea. Sleep Health. 2023;9(4):430–40. [DOI] [PubMed] [Google Scholar]
- 27.Moussa M, Alzaabi Y, Khandoker A. Explainable Computer-Aided detection of obstructive sleep apnea and depression. IEEE Access. 2022 Oct 19;10.
- 28.Hafezi M, Montazeri N, Saha S, Zhu K, Gavrilovic B, Yadollahi A, et al. Sleep apnea severity Estimation using a deep learning model from tracheal movements. IEEE Access. 2020 Jan 24;8:1–1. [Google Scholar]
- 29.Chen Q, Liang Z, Wang Q, Ma C, Lei Y, Sanderson JE, et al. Self-helped detection of obstructive sleep apnea based on automated facial recognition and machine learning. Sleep Breath. 2023 Dec;27(6):2379–88. [DOI] [PubMed] [Google Scholar]
- 30.Elwali A, Moussavi Z. Predicting polysomnography parameters from anthropometric features and breathing sounds recorded during wakefulness. Diagnostics (Basel). 2021 May 19;11(5):905. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.He S, Su H, Li Y, Xu W, Wang X, Han D. Detecting obstructive sleep apnea by craniofacial image-based deep learning. Sleep Breath. 2022;26(4):1885–95. [DOI] [PubMed] [Google Scholar]
- 32.Li Z, Li Y, Zhao G, Zhang X, Xu W, Han D. A model for obstructive sleep apnea detection using a multi-layer feed-forward neural network based on electrocardiogram, pulse oxygen saturation, and body mass index. Sleep Breath. 2021 Dec;25(4):2065–72. [DOI] [PubMed] [Google Scholar]
- 33.Zhuang Z, Wang F, Yang X, Zhang L, Fu CH, Xu J, et al. Accurate contactless sleep apnea detection framework with signal processing and machine learning methods. Methods. 2022 Sept;205:167–78. [DOI] [PubMed]
- 34.US Preventive Services Task Force, Mangione CM, Barry MJ, Nicholson WK, et al. Screening for obstructive sleep apnea in adults: US preventive services task force recommendation statement. JAMA. 2022;328(19):1945–50. [DOI] [PubMed] [Google Scholar]
- 35.Sangalli L, Yanez-Regonesi F, Fernandez-Vial D, Moreno-Hay I. Self-reported improvement in obstructive sleep apnea symptoms compared to treatment response with mandibular advancement device therapy: a retrospective study. Sleep Breath. 2022. [DOI] [PubMed]
- 36.Goldstein C, Berry R, Kent D, Kristo D, Seixas A, Redline S et al. Artificial intelligence in sleep medicine: background and implications for clinicians. J Clin Sleep Med. 2020 Feb 17;16. [DOI] [PMC free article] [PubMed]
- 37.Bahrami M, Forouzanfar M. Sleep apnea detection from Single-Lead ECG: A comprehensive analysis of machine learning and deep learning algorithms. IEEE Trans Instrum Meas. 2022;71:1–11. [Google Scholar]
- 38.Salam SS, Rafi R. Deep Learning Approach for Sleep Apnea Detection Using Single Lead ECG: Comparative Analysis Between CNN and SNN. 2023 26th International Conference on Computer and Information Technology (ICCIT). 2023;1–6.
- 39.Monowar MM, Nobel S, Afroj M, Hamid MA, Uddin M, Kabir M et al. Advanced sleep disorder detection using multi-layered ensemble learning and advanced data balancing techniques. Front Artif Intell. 2025 Jan 28;7. [DOI] [PMC free article] [PubMed]
- 40.Tyagi PK, Agarwal D. Systematic review of automated sleep apnea detection based on physiological signal data using deep learning algorithm: a meta-analysis approach. Biomed Eng Lett. 2023 Aug 1;13(3):293–312. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.La Fisca L, Jennebauffe C, Bruyneel M, Ris L, Lefebvre L, Siebert X, et al. Enhancing OSA assessment with explainable AI. Annu Int Conf IEEE Eng Med Biol Soc. 2023 July;2023:1–6. [DOI] [PubMed]
- 42.Abd-Alrazaq A, Aslam H, AlSaad R, Alsahli M, Ahmed A, Damseh R et al. Detection of sleep apnea using wearable AI: systematic review and Meta-Analysis. J Med Internet Res 2024 Sept 10;26:e58187. [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Material 1: Additional File 1: Data: Supplementary Table 1: Search query for each database, Supplementary Table 2: List of Abbreviations.
Supplementary Material 2: Additional File 2: Data: 1st Author, Source of Data (EEG, ECG, airflow, SpO2, etc.), Midpoint (percentage), Accuracy range (percentage), Number of patients, Train/Validation/Test, Age (Mean±SD), Male/Female, Note (Model Architecture (e.g., DL, ML or Hybrid and Code Availability), Country, Dataset.
Data Availability Statement
Data is provided within the manuscript or supplementary information files.





