Skip to main content
BMC Medical Informatics and Decision Making logoLink to BMC Medical Informatics and Decision Making
. 2025 Jul 28;25:278. doi: 10.1186/s12911-025-03129-x

Diagnostic accuracy of artificial intelligence for obstructive sleep apnea detection: a systematic review

Sara Haghighat 1,2, Muhammed Joghatayi 3,, Julien Issa 2,4, Sarina Azimian 2,5, Janet Brinz 2,6, Ali Ashkan 7, Akhilanand Chaurasia 8, Zahra Rahimian 9, Linda Sangalli 10
PMCID: PMC12306116  PMID: 40722158

Abstract

Background

Obstructive sleep apnea (OSA) is a highly prevalent sleep disorder. Misdiagnosis might lead to several systemic conditions, including hypertension, vascular damage, and cognitive impairment. The gold-standard diagnostic tool for OSA is polysomnography, which is expensive, time-consuming, and not accessible everywhere. Artificial intelligence (AI) algorithms can facilitate diagnosis by detecting patients’ signs and symptoms. In this systematic review, we evaluated the diagnostic accuracy of AI models in detecting sleep apnea.

Methods

We searched six major databases, PubMed®, Cochrane, Web of Science, Scopus, Embase, and IEEE Xplore, using keywords related to AI and OSA. Eligible studies focused on adult populations, used in-laboratory PSG as the reference standard, and applied AI models trained on multiple clinical features. Reviews, pediatric studies, and articles lacking accuracy metrics were excluded. From the included articles, data were extracted regarding patients and datasets, type of AI model applied, accuracy report, and explainability of the AI model. A risk of bias assessment was done using the QUADAS-2 checklist.

Results

Thirteen studies were included in our final analysis. The AI models consisted of deep learning, machine learning, and hybrid models with various architectures. The reported accuracy of studies ranged from 67.03 to 98.6%, with the highest being related to hybrid and deep learning models. Risk of bias assessment showed that 7 of the studies had a low risk of bias, indicating high reliability.

Conclusions

AI-driven models, particularly deep learning and hybrid architectures, show significant promise in diagnosing obstructive sleep apnea. However, challenges such as transparency, explainability, and variability in performance necessitate diverse training datasets to improve generalizability for clinical adoption.

Registration of systematic reviews

The protocol of this systematic review was registered in PROSPERO (CRD42023453789), available from: https://www.crd.york.ac.uk/prospero/display_record.php?ID=CRD42023453789.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12911-025-03129-x.

Keywords: Artificial intelligence, Obstructive sleep apnea, Diagnostic accuracy, Deep learning, Machine learning

Background

Obstructive sleep apnea (OSA), also called obstructive sleep apnea syndrome (OSAS), is a potentially severe sleep disorder characterized by repeated episodes of partial or complete obstruction of the upper airway during sleep [1]. These episodes manifest as significant reductions in airflow or complete pauses in breathing, known as hypopneas (partial airway obstruction) and apneas (complete airway obstruction) [2]. The etiology of OSA involves a multifactorial interplay of anatomical susceptibilities and neuromuscular control of the airway [3]. Risk factors include obesity, male gender, advancing age, and genetic predispositions, among others [4]. The pathogenesis of OSA includes intermittent hypoxia-induced oxidative stress and systemic inflammation [5], which result from disrupted sleep architecture and increased sympathetic nervous system activity [6]. When OSA remains untreated, the clinical manifestations range from daytime fatigue to cognitive impairment [7], increasing the risk of uncontrolled systemic hypertension, cardiovascular disease, diabetes, chronic kidney disease, and vehicular accidents [8]. These significant consequences underline the critical need for accurate diagnosis and effective management.

The gold standard for diagnosis of OSA is an overnight polysomnography (PSG). PSG involves monitoring several body functions during sleep, including brain activity (as measured by EEG), eye movement, muscle activity, heart rate, respiratory effort, and blood oxygen levels. This comprehensive approach allows for detailed observation and analysis of the sleep stages and identification of any disturbances related to breathing or other vital signs [9]. However, the PSG is resource-intensive and requires specialized facilities and personnel [10]. Thus, its complexity and cost highlight the necessity for more accessible diagnostic modalities [11].

Artificial intelligence (AI) technologies, particularly machine learning (ML) and deep learning (DL), are transforming diagnostic approaches in sleep medicine. ML algorithms can learn from structured clinical data using handcrafted features, while DL models extract patterns directly from raw signals such as airflow, oximetry, or electrocardiogram (ECG) [12]. This ability to autonomously recognize subtle spatial and temporal patterns makes DL especially suited for detecting OSA-related disruptions [13]. By leveraging large datasets, these models can achieve high diagnostic accuracy and offer scalable solutions that augment traditional assessment tools.

Various automated approaches aim to simplify OSA diagnosis and overcome the limitations of PSG. The availability of simplified, automated, and reliable alternative tools to PSG would enable the accurate diagnosis of OSA, which would have several other advantages for patients. For instance, less need for too many sensors would help improve patient comfort. Moreover, the tool could be available for home testing, therefore reducing the long waiting lists within healthcare for patients to undergo a PSG [14]. Additionally, an automated tool would dramatically decrease the effort and time needed for the specialists to analyze overnight physiological signals. Consequently, the advantages mentioned would facilitate patients’ access to treatment while maintaining appropriate accuracy in diagnosis [15].

Many studies have evaluated the accuracy of AI models in diagnosing OSA; however, differences in levels of accuracy, AI models used, and various performance metrics have yielded conflicting results.

Therefore, it is necessary to systematically review and evaluate the performance of current AI models for diagnosing OSA.

This systematic review evaluated the diagnostic accuracy and performance of deep learning-based AI models in identifying and detecting OSA in adult patients.

Methods

The protocol of this systematic review was registered in PROSPERO (CRD42023453789), available from: https://www.crd.york.ac.uk/prospero/display_record.php?ID=CRD42023453789. The study was conducted following the Preferred Reporting Items for Systematic Review and Meta-Analyses (PRISMA) [16].

Eligibility Criteria

Studies were selected based on the eligibility criteria and consensus from all the authors.

To be included, the studies needed to (A) be original investigations assessing the diagnostic performance of AI algorithms in identifying OSA among adult populations; (B) original research designs, including randomized clinical trials, cross-sectional, and case-control studies; (C) utilize in-laboratory PSG as the reference standard; (D) incorporate multiple clinical factors to train the AI models; (E) be published in English language; and (F) be published in the last 5 years. This time window of the last 5 years was chosen to ensure a comprehensive review of the most recent advancements in AI algorithms. Conversely, studies were excluded if they (A) did not report accuracy metrics; (B) were review articles, meta-analyses, or conference abstracts; (C) were not conducted in human subjects; (D) involved pediatric patients (< 18 years); (E) lacked full-text availability; (F) were published in languages other than English; (G) relied solely on one-organ symptoms and signs, such as cardiovascular or respiratory data, to identify OSA; or (H) evaluated aspects of OSA other than diagnosis (such as management, adherence and treatment outcomes).

Information Sources and Bibliography Search

A comprehensive literature search was conducted on October 15, 2023, across six databases (PubMed®, Cochrane, Web of Science, Scopus, Embase, and IEEE Xplore) to identify studies evaluating the diagnostic accuracy of artificial intelligence (AI) systems (Intervention) compared to polysomnography (PSG) (Comparator) in diagnosing obstructive sleep apnea (OSA) in adults (Population). The research question guiding this review was: “What is the diagnostic accuracy of AI-based systems compared to the gold standard PSG in detecting OSA in adults?”

Search Strategy

We included keywords for three main domains, namely AI, diagnosis, and obstructive sleep apnea, with AND Boolean operators between the keywords. Supplementary Table 1 in Additional File 1 presents the search query for each database.

Selection Process

To ensure a comprehensive search, we conducted a two-phase screening process. Rayyan Web Platform [17] was used to screen articles and remove duplicates. Initially, a team of four collaborators screened the titles and abstracts of potential articles using predefined eligibility criteria. In the second phase, two additional reviewers independently reassessed the articles labeled as ‘included’ or ‘maybe’ to confirm their relevance and improve screening accuracy. Final inclusion decisions were made by consensus among the review authors. This collaborative approach helped to minimize bias and ensure the inclusion of relevant studies.

Data Collection Process and Data Items

We categorized models as machine learning, deep learning, or hybrid based on their architectural structure. ML models included traditional classifiers that relied on engineered features [18], while DL models referred to multilayer neural networks capable of learning directly from raw data. Hybrid models were defined as those integrating both ML and DL components within a single diagnostic pipeline [19]. Models composed entirely of DL or ML components were classified under their respective category to maintain consistency.

The diagnostic accuracy of AI systems was evaluated by extracting the overall accuracy (correct prediction rate) from all included studies. Accuracy was selected as the primary performance metric due to the substantial heterogeneity in model types, input features, and study designs, as well as its consistent reporting across all studies. Additional performance metrics, such as the F1 score (harmonic mean of precision and recall/sensitivity), sensitivity (true positive rate, TP), specificity (true negative rate, TN), false negatives (FN), false positives (FP), and the area under the curve (AUC), were also extracted when reported.

To evaluate model performance, we extracted an aggregated accuracy metric from each study, calculated as (TP + TN) / (TP + TN + FP + FN). These values were summarized in tables.

graphic file with name d33e464.gif

In addition to primary performance metrics, we extracted detailed characteristics including dataset type, participant demographics, input signals, and model architectures.

Risk of Bias

The QUADAS-2 (Quality Assessment Tool for Diagnostic Accuracy Studies - Version 2) checklist was used to evaluate the quality of diagnostic accuracy studies [20]. This tool evaluates risk of bias across four domains: (1) patient selection, which assesses whether participants were enrolled in a way that avoids bias; (2) index test, which evaluates whether the test under investigation was conducted and interpreted without knowledge of the reference standard; (3) reference standard, which examines the reliability and appropriateness of the diagnostic benchmark used; and (4) flow and timing, which assesses whether all participants received the same reference standard and whether there were delays that could affect results. In addition, the first three domains are examined for concerns about applicability (e.g., how well the study matches the review question). Each domain was rated as having a low, high, or unclear risk of bias.

Results

Study Selection

The search yielded 2907 articles. After removing duplicates (N = 1507), 1400 studies were initially screened. Following full-text evaluation, 13 studies were included for the qualitative analysis. The reasons for exclusion and the number of studies in each category are summarized in the PRISMA flowchart (Fig. 1).

Fig. 1.

Fig. 1

PRISMA flowchart of the study

Study Characteristics

The qualitative analysis was performed on 13 included studies, including a total of 12,631 participants. Three studies (23.08%) did not report the ratio of male/female [2123], and three did not report mean age [21, 22, 24]. Among the remaining 10 studies (N = 4708), there were a total of 1485 females (31.54%) and 3223 males (68.45%), and the overall mean age was 45.82 ± 14.21.

We divided the models into three categories: DL, ML, and hybrid. The number of studies in the DL, ML, and hybrid models was 4, 6, and 3, respectively. There were a total of 21 types of algorithms developed or used in the studies, consisting of 10 ML algorithms (Logistic Regression, Random Forest, Support Vector Machine, AdaBoost Learning, Gradient Boosting, K Nearest Neighbor, Neural Network, CatBoost Classifier, Extra Trees Classifier and Decision Tree), 6 DL algorithms (Convolutional Neural Network, Long Short Term Memory, Bi-directional Long Short Term Memory, ResNet101, DeepsleepNet, CMS-2-Net), and five hybrid algorithms (FNN, Complex Tree, RuBoosted Trees, GRU and ML Meta-learners). All studies reported the accuracy of the model, while other performance metrics, including sensitivity, specificity, AUC, and F1 score, were reported in 12, 9, 7, and 3 of the studies, respectively. The reported accuracy of AI systems in diagnosing OSA ranged from 67.03 to 98.60%. Table 1 presents the AI model type, total number of subjects, and reported midpoint accuracy for each study. (The list of abbreviations is provided in Additional File 1, Table 2).

Table 1.

Model type, sample size and reported accuracy per study

1st Author Algorithms used Accuracy (Midpoint, %) Sample size Country
J et al. [22] Bi-LSTM, ResNet 101, DeepSleepNet 83.79% 7745 a South Korea
Leong [25] LR, RF, SVM, AdaB, GB, NN, kNN 87.9% 2996 Singapore
Zhang [23] CMS2-Net, CNN 73.5% 128 China
Arslan [21] DNN, RNN, LSTM, GRU, ML Meta-Learner 88.7% 50 Turkey
Strumpf [26] CNN 88% 84 USA
Zhuang [24] RF 95.14% 10 China
Moussa [27] RF, SVM, kNN, LR 82.70% 150 UAE
Hafezi [28] CNN, LSTM 83% 69 Canada
Chen [29] Top 5: RF, LR, CATBOOST, ET, GB b 69% 653 China
Elwali & Moussavi [30] RF 80.15% 145 Canada
He [31] CNN, GB 81.05% 393 China
Z. Zhang [24] DT, RF, kNN, SVM 82.90% 27 USA
Li [32] FNN, SVM, Complex Tree, RUSBoosted Trees, LR 93.10% 181 China

a: Final image-based dataset

b: A total of 18 models were used, but 13 others were in the supplementary table and not accessible

Table 2 summarizes the datasets and physiological inputs used to train or evaluate each AI model.

Table 2.

Dataset and input signal types across included studies

1st Author Dataset Input features
J et al. [22] Korean National Information Society Agency Full PSG
Leong [25] Local Sleep Medicine Database Full PSG, Demographic & Anthropometric Data
Zhang [23] Local Dataset (Peking University Sixth Hospital) + Public (Dream Open Dataset) EEG, EMG, EOG, ECG
Arslan [21] Local Dataset (Yozgat Bozok University, Department of Chest Diseases Sleep Laboratory) Full PSG
Strumpf [26] University Hospitals Cleveland Medical Center Bolwell and Beachwood Sleep Labs Full PSG
Zhuang [24] Sleep Center of Huai’an First People’s Hospital in Jiangsu Province PSG, Physiological Data (radar signals, vital signs)
Moussa [27] American Center for Psychiatry and Neurology, Stanford Technology Analytics and Genomics in Sleep (STAGES) EEG, ECG, Breathing Signals
Hafezi [28] Local Dataset (Sleep Laboratory at the Toronto Rehabilitation Institute)

Chest & abdominal movements, respiratory inductance plethysmography, airflow

by nasal pressure cannula, SpO2,Tracheal

Movements /Apnea Hypopnea index (AHI)

Chen [29] Local Dataset (Not Specified) Clinical Data (including PSG and PM) + Craniofacial Images
Elwali & Moussavi [30] Local Dataset (Sleep Disorders Center in Misericordia Health Centre (Winnipeg, Canada)) Full PSG
He [31] Sleep Laboratories at the Department of Otolaryngology Head and Neck Surgery, Beijing Tongren Hospital (Beijing, China) Craniofacial Images + PSG
Z. Zhang [24] Weill Cornell Center for Sleep Medicine (New York, USA) PSG, SpO2, airflow/bed-integrated radio-frequency sensor by near-field coherent sensing
Li [32] Sleep Medicine Center of Beijing Tongren Hospital (Bei Jing Shi, China) PSG, BMI, ECG, SpO2

Table 3 summarizes other findings (such as sensitivity and specificity) for the included studies.

Table 3.

Summary of other findings for the included studies

1st Author AUC Specificity(%) Sensitivity(%) F1 Score Other
J et al. [22] N/A N/A N/A

weighted F1 score: 80.66–86.54

macro F1 score: 80.68–83.60

-
Leong [25] 0.848–0.928 40.4%-60.6% 90.9%-97.4% N/A -
Zhang [23] N/A N/A 52%-69% 53–69

precision: 64–76%

Kappa: 56–73%

Arslan [21] 0.99 N/A 95.76%-91.70% N/A -
Strumpf [26] 0.92–0.95 88%-96% 83%-95% N/A -
Zhuang [24] N/A 97.32% 72.60% N/A -
Moussa [27] 0.79–0.85 83.69%-73.17% 69.35%-96.44% N/A -
Hafezi [28] N/A 36%-94% 67%-98% N/A -
Chen [29] 0.72–0.76 N/A 68%-75% N/A -
Elwali & Moussavi [30] N/A 63.9%-100% 60.6%-94.7% 0.66–0.92 -
He [31] 0.747–0.963 55.8%-96.7% 81.8%-98.8% N/A -
Z. Zhang [24] N/A 72.9%-89.1% 56.7%-74.3% N/A -
Li [32] 0.92–0.98 76.0% − 93.9% 89.0% − 98.6% N/A -

Additional File 2 provides full study-level details including dataset, input type, sample size, model architecture, accuracy range, sex distribution, and data split.

Risk of Bias Assessment

Seven of the included studies were judged to have a low risk of bias across all domains, indicating overall methodological reliability. The remaining studies raised concerns primarily in the patient selection domain: Arslan [21], Moussa et al. [27], and Z. Zhang et al. [24] showed unclear risk, while Zhuang et al. [33] had a high risk due to unclear inclusion criteria.

All studies showed low applicability concerns in the index test and reference standard domains. However, Leong et al. [25], Chen et al. [29], and Z. Zhang et al. [24] raised unclear concerns in the patient selection domain, typically due to limited information about sampling methods or study population characteristics. Full assessments are presented in Table 4; Figs. 2 and 3.

Table 4.

QUADAS-2 results of the included studies

graphic file with name 12911_2025_3129_Tab4_HTML.jpg

Fig. 2.

Fig. 2

Diagram of risk of bias assessment

Fig. 3.

Fig. 3

Diagram of applicability concern

Discussion

Sleep Disorders Epidemiology and AI in Diagnosis

Obstructive sleep apnea affects an estimated one billion adults worldwide, with 85–95% of cases remaining undiagnosed. Left untreated, OSA is associated with a wide range of complications, including cardiovascular and metabolic disorders, neurocognitive dysfunction, and a diminished quality of life. The burden of undiagnosed cases underscores the critical need for accurate and accessible diagnostic tools [8, 10].

Although PSG remains the gold standard for diagnosing OSA, its widespread use is hindered by high costs, the need for specialized equipment and trained personnel, and the requirement for overnight monitoring in dedicated sleep laboratories. These barriers are especially pronounced in low-resource settings, where access to sleep medicine infrastructure is limited. Additionally, PSG can be inconvenient and uncomfortable for patients, potentially impacting sleep quality and diagnostic yield [34].

As a result, many individuals with OSA never receive a diagnosis, often due to the condition’s nonspecific symptoms, with only around 20% of affected individuals reporting them to healthcare providers [35]. Against this backdrop, artificial intelligence offers promising solutions by enabling automated models that can interpret complex physiological data. These tools have the potential to augment or even replace conventional diagnostic methods, improving early detection and expanding access to care, particularly in underserved populations [36].

AI Models Diversity

The 13 studies included in this review utilized a wide range of AI methodologies, including six traditional ML models, four DL architectures, and three hybrid approaches. ML models, such as Support Vector Machines (SVMs) and Decision trees, rely on handcrafted features and require substantial domain expertise for optimal performance. They tend to be more interpretable and computationally efficient but may struggle with complex pattern recognition.

In contrast, DL models, including CNNs and RNNs, automatically learn multi-level features directly from raw physiological data. This makes them well-suited for modeling the complex temporal and spatial patterns characteristic of OSA diagnostics [37]. While DL approaches typically yield higher diagnostic accuracy, they also require large datasets and considerable computational resources to minimize overfitting [38].

Hybrid models that combine ML and DL components attempt to leverage the strengths of both strategies, offering a balance between adaptability, performance, and interpretability.

Accuracy of AI Models

The diagnostic accuracy of AI models for detecting OSA in the included studies ranged from 67.03 to 98.60%, with a median value of 89.66%. Notably, hybrid models, particularly those combining Convolutional and Recurrent Neural Networks (CNN-RNN), demonstrated the highest diagnostic performance. For example, the model developed by Arslan et al. reported the highest accuracy at 98.60% [21], followed closely by Moussa et al. (98.36%) and Li et al. (97.80%) [27, 32].

These multi-layered hybrid architectures appear particularly effective in capturing both spatial and temporal characteristics of physiological signals relevant to OSA. By integrating feature extraction with sequential pattern analysis, they leverage the complementary strengths of CNNs and RNNs, resulting in enhanced predictive performance [39].

Code Availability

Five of the thirteen studies shared links to their open-access code. The remaining eight articles did not specify whether their code was available, and no further attempts were made to contact the authors regarding code sharing. This gap underscores the ongoing challenges in AI research regarding reproducibility and open science practices.

Comparison with other Similar Articles

A recent systematic review focusing on automated sleep apnea detection using physiological signal data emphasized the strength of deep learning models in capturing complex spatial and temporal patterns [40]. That review also reported improved accuracy in hybrid architectures, consistent with our findings.

Similarly, a systematic review of pediatric OSA detection showed that traditional ML models, while slightly less accurate than DL models, offered advantages such as lower computational demands and easier implementation in low-resource settings [15].

Overall, our review supports the emerging consensus that hybrid and deep learning models outperform traditional ML in terms of diagnostic accuracy. However, these gains come at the cost of increased data requirements and computational complexity. Notably, the diagnostic performance varied widely across studies, reinforcing the importance of external validation in diverse clinical populations before these models can be deployed at scale.

Clinical significance

Integrating AI into clinical practice presents a number of practical and ethical challenges. While AI systems offer objective analyses of physiological data, their outputs may be perceived as opaque or rigid, especially when the underlying algorithms are not transparent. This can lead to skepticism among clinicians and patients, potentially hindering adoption. For AI to be trusted in real-world settings, its decisions must be explainable. Clinicians need to understand not just the output, but the reasoning behind it, making explainable AI (XAI) frameworks critical to clinical integration [41].

Transparency remains a persistent issue. In our review, only 5 out of 13 studies provided access to source code, limiting reproducibility and independent validation. Moreover, AI models often require large training datasets and high computational power, constraints that may limit their implementation in smaller clinics or low-resource environments. Biases in training data can also affect diagnostic accuracy across demographic groups, raising important concerns about fairness and equity.

A recent meta-analysis evaluating wearable AI technologies reported a pooled diagnostic accuracy of 86.9% in detecting apneic events. While these devices offer real-time monitoring and accessibility benefits, their performance was deemed insufficient for standalone clinical use, reinforcing the need for AI tools to function as adjuncts to standard diagnostic workflows [42].

Limitations

Future research should focus on large-scale, multicenter studies, the development of publicly accessible datasets, and standardizing methodologies to validate AI models across diverse populations. Including diverse patient demographics and implementing standardized methodologies will improve AI model generalizability, ensuring that diagnostic tools meet the needs of various patient groups.

Conclusion

AI-driven models, especially DL and hybrid architectures, demonstrate significant potential for diagnosing obstructive sleep apnea, with reported accuracies ranging from 67.03 to 98.60%. Hybrid models appear particularly effective at capturing OSA-specific physiological patterns due to their ability to integrate spatial and temporal analysis.

While these models outperform traditional machine learning approaches in terms of diagnostic accuracy, the latter still hold value in settings with limited computational resources, given their lower complexity and interpretability.

However, before AI tools can be widely adopted in clinical sleep medicine, several barriers must be addressed. These include the lack of model transparency, limited explainability, and variability in diagnostic performance across patient populations. Future research should focus on developing XAI systems, standardizing validation methods, and ensuring consistent performance across diverse demographic and clinical settings. Addressing these challenges will be essential for translating AI’s technical capabilities into meaningful clinical impact.

Supplementary Information

Below is the link to the electronic supplementary material.

12911_2025_3129_MOESM1_ESM.docx (17.4KB, docx)

Supplementary Material 1: Additional File 1: Data: Supplementary Table 1: Search query for each database, Supplementary Table 2: List of Abbreviations.

12911_2025_3129_MOESM2_ESM.xlsx (25KB, xlsx)

Supplementary Material 2: Additional File 2: Data: 1st Author, Source of Data (EEG, ECG, airflow, SpO2, etc.), Midpoint (percentage), Accuracy range (percentage), Number of patients, Train/Validation/Test, Age (Mean±SD), Male/Female, Note (Model Architecture (e.g., DL, ML or Hybrid and Code Availability), Country, Dataset.

Acknowledgements

Not applicable.

Abbreviations

AdaB

Adaptive Boosting

AI

Artificial Intelligence

AUC

Area Under the Curve

AUROC

Area Under the Receiver Operating characteristic Curve

Bi-LTSM

Bidirectional Long Short-Term Memory

CATBOOST

CatBoost Classifier

CMS2-Net

Co-attention Meta Sleep Staging Network

CNN

Convolutional Neural Network

DL

Deep Learning

DNN

Deep Neural Network

DT

Decision Tree

ECG

Electrocardiography

EEG

Electroencephalography

EMG

Electromyography

EOG

Electrooculography

ET

Extra Trees Classifier

FN

False Negative

FP

False Positive

GB

Gradient Boosting

GRU

Gated Recurrent Unit

kNN

K Nearest Neighbor

LR

Logistic Regression

LSTM

Long Short-Term Memory

ML

Machine Learning

N/A

Not Available

NN

Neural Network

OSA

Obstructive Sleep Apnea

PSG

Polysomnography

RF

Random Forest

RNN

Recurrent Neural Network

SVM

Support Vector Machine

TN

True Negative

TP

True Positive

UAE

United Arab Emirates

USA

The United States of America

XAI

Explainable Artificial Intelligence

Author Contributions

SH: Project administration, Conceptualization, Data Curation, Investigation, Writing - Original Draft, Writing - Review & Editing. LS: Project administration, Conceptualization, Supervision, Writing - Review & Editing, Validation. MJ: Investigation, Data Curation, Writing - Original Draft, Writing - Review & Editing, Formal analysis. SA: Data Curation, Investigation, Resources. JI: Investigation, Data Curation, Resources. AS: Investigation, Data Curation, Resources. JB: Investigation, Data Curation, Resources. ZH: Investigation, Data Curation, Resources. AK: Investigation, Supervision, Writing - Review & Editing, Validation. All authors reviewed the manuscript.

Funding

No funding was received for the preparation or publication of this article.

Data Availability

Data is provided within the manuscript or supplementary information files.

Declarations

Ethics approval and consent to participate

The study was conducted following the Preferred Reporting Items for Systematic Review and Meta-Analyses (PRISMA) [16] and the ethical principles of the Declaration of Helsinki.

Consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Arnold J, Sunilkumar M, Krishna V, Yoganand SP, Kumar MS, Shanmugapriyan D. Obstructive sleep apnea. J Pharm Bioallied Sci. 2017 Nov;9(Suppl 1):S26–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Wenner J, Cheema R, Ayas N. Clinical manifestations and consequences of obstructive sleep apnea. Lippincott Williams & Wilkins. 2009 Mar;29(2):76–83. [DOI] [PubMed]
  • 3.Demirgüneş DD, Eroğul O, Akçam T, Telatar Z. Analysis of respiration, oxygen saturation and acoustic signals of snoring patients. 2009 May.
  • 4.Al Lawati NM, Patel SR, Ayas NT, Epidemiology. Risk factors, and consequences of obstructive sleep apnea and short sleep duration. Prog Cardiovasc Dis. 2009 Jan;51(4):285–93. [DOI] [PubMed] [Google Scholar]
  • 5.Pham LV, Schwartz AR. The pathogenesis of obstructive sleep apnea. J Thorac Dis. 2015 Aug;7(8):1358–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Khayat R, Patt B, Hayes D. Obstructive sleep apnea: the new cardiovascular disease. Part I: obstructive sleep apnea and the pathogenesis of vascular disease. Heart Fail Rev. 2009 Sept 1;14(3):143–53. [DOI] [PMC free article] [PubMed]
  • 7.Eckert DJ, Malhotra A. Pathophysiology of adult obstructive sleep apnea. Proceedings of the American Thoracic Society. 2012 Dec 20. [DOI] [PMC free article] [PubMed]
  • 8.Melamed KH, Goldhaber SZ. Obstr Sleep Apnea Circulation. 2015;132(6):e114–6. [DOI] [PubMed] [Google Scholar]
  • 9.Polysomnography. and Other sleep studies. Springer Publishing Company; 2023.
  • 10.Park JG, Ramar K, Olson EJ. Updates on Definition, Consequences, and Management of Obstructive Sleep Apnea. Mayo Clinic Proceedings. 2011 June 1;86(6):549–55. [DOI] [PMC free article] [PubMed]
  • 11.Keenan SA. Chapter 3 an overview of polysomnography. In: Guilleminault C, editor. Handbook of clinical neurophysiology. Handbook of clinical neurophysiology. Elsevier. 2005;6:33–50.
  • 12.Belk RW, Belanche D, Flavián C. Key concepts in artificial intelligence and technologies 4.0 in services. Springer Sci + Bus Media. 2023;17(1):1–9. [Google Scholar]
  • 13.Pesapane F, Codari M, Sardanelli F. Artificial intelligence in medical imaging: threat or opportunity? Radiologists again at the forefront of innovation in medicine. Eur Radiol Experimental. 2018 Oct 24;2(1):35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Lachenmeier W, Lachenmeier DW. Home monitoring of oxygen saturation using a Low-Cost wearable device with haptic feedback to improve sleep quality in a lung cancer patient: A case report. Multidisciplinary Digit Publishing Inst. 2022 Mar;7(2):43–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Gutiérrez-Tobal GC, Álvarez D, Kheirandish-Gozal L, Del Campo F, Gozal D, Hornero R. Reliability of machine learning to diagnose pediatric obstructive sleep apnea: systematic review and meta-analysis. Pediatr Pulmonol. 2022 Aug;57(8):1931–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Moher D, Liberati A, Tetzlaff J, Altman DG, The PRISMA group. Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement. PLoS Med. 2009;6(7):e1000097. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Ouzzani M, Hammady H, Fedorowicz Z, Elmagarmid A. Rayyan—a web and mobile app for systematic reviews. Syst Reviews. 2016;5(1):210. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Jordan MI, Mitchell TM. Machine learning: trends, perspectives, and prospects. Science. 2015;349(6245):255–60. [DOI] [PubMed] [Google Scholar]
  • 19.Goodfellow I, Bengio Y, Courville A, Bengio Y. Deep learning. MIT press Cambridge. 2016;1.
  • 20.Whiting PF, Rutjes AW, Westwood ME, Mallett S, Deeks JJ, Reitsma JB, et al. QUADAS-2: A revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529–36. [DOI] [PubMed] [Google Scholar]
  • 21.Arslan RS. Sleep disorder and apnea events detection framework with high performance using two-tier learning model design. PeerJ Comput Sci. 2023;9:e1554. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.W JJ, Jg Y, Dk LDKYW. K, Standardized image-based polysomnography database and deep learning algorithm for sleep-stage classification. Sleep. 2023 Dec 11;46(12). [DOI] [PubMed]
  • 23.Zhang C, Yu W, Li Y, Sun H, Zhang Y, De Vos M. CMS2-Net: Semi-supervised sleep staging for diverse obstructive sleep apnea severity. IEEE J Biomed Health Inf. 2022 July;26(7):3447–57. [DOI] [PubMed]
  • 24.Zhang Z, Conroy T, Krieger A, Kan E. Detection and prediction of sleep disorders by Covert Bed-Integrated RF sensors. IEEE Trans Biomed Eng. 2022 Jan 1;1–11. [DOI] [PubMed] [Google Scholar]
  • 25.Leong ZH, Loh SRH, Leow LC, Ong TH, Toh ST. A machine learning approach for the diagnosis of obstructive sleep Apnoea using oximetry, demographic and anthropometric data. Singap Med J. 2023 May 2. [DOI] [PMC free article] [PubMed]
  • 26.Strumpf Z, Gu W, Tsai CW, Chen PL, Yeh E, Leung L, et al. Belun ring (Belun sleep system BLS-100): deep learning-facilitated wearable enables obstructive sleep apnea detection, apnea severity categorization, and sleep stage classification in patients suspected of obstructive sleep apnea. Sleep Health. 2023;9(4):430–40. [DOI] [PubMed] [Google Scholar]
  • 27.Moussa M, Alzaabi Y, Khandoker A. Explainable Computer-Aided detection of obstructive sleep apnea and depression. IEEE Access. 2022 Oct 19;10.
  • 28.Hafezi M, Montazeri N, Saha S, Zhu K, Gavrilovic B, Yadollahi A, et al. Sleep apnea severity Estimation using a deep learning model from tracheal movements. IEEE Access. 2020 Jan 24;8:1–1. [Google Scholar]
  • 29.Chen Q, Liang Z, Wang Q, Ma C, Lei Y, Sanderson JE, et al. Self-helped detection of obstructive sleep apnea based on automated facial recognition and machine learning. Sleep Breath. 2023 Dec;27(6):2379–88. [DOI] [PubMed] [Google Scholar]
  • 30.Elwali A, Moussavi Z. Predicting polysomnography parameters from anthropometric features and breathing sounds recorded during wakefulness. Diagnostics (Basel). 2021 May 19;11(5):905. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.He S, Su H, Li Y, Xu W, Wang X, Han D. Detecting obstructive sleep apnea by craniofacial image-based deep learning. Sleep Breath. 2022;26(4):1885–95. [DOI] [PubMed] [Google Scholar]
  • 32.Li Z, Li Y, Zhao G, Zhang X, Xu W, Han D. A model for obstructive sleep apnea detection using a multi-layer feed-forward neural network based on electrocardiogram, pulse oxygen saturation, and body mass index. Sleep Breath. 2021 Dec;25(4):2065–72. [DOI] [PubMed] [Google Scholar]
  • 33.Zhuang Z, Wang F, Yang X, Zhang L, Fu CH, Xu J, et al. Accurate contactless sleep apnea detection framework with signal processing and machine learning methods. Methods. 2022 Sept;205:167–78. [DOI] [PubMed]
  • 34.US Preventive Services Task Force, Mangione CM, Barry MJ, Nicholson WK, et al. Screening for obstructive sleep apnea in adults: US preventive services task force recommendation statement. JAMA. 2022;328(19):1945–50. [DOI] [PubMed] [Google Scholar]
  • 35.Sangalli L, Yanez-Regonesi F, Fernandez-Vial D, Moreno-Hay I. Self-reported improvement in obstructive sleep apnea symptoms compared to treatment response with mandibular advancement device therapy: a retrospective study. Sleep Breath. 2022. [DOI] [PubMed]
  • 36.Goldstein C, Berry R, Kent D, Kristo D, Seixas A, Redline S et al. Artificial intelligence in sleep medicine: background and implications for clinicians. J Clin Sleep Med. 2020 Feb 17;16. [DOI] [PMC free article] [PubMed]
  • 37.Bahrami M, Forouzanfar M. Sleep apnea detection from Single-Lead ECG: A comprehensive analysis of machine learning and deep learning algorithms. IEEE Trans Instrum Meas. 2022;71:1–11. [Google Scholar]
  • 38.Salam SS, Rafi R. Deep Learning Approach for Sleep Apnea Detection Using Single Lead ECG: Comparative Analysis Between CNN and SNN. 2023 26th International Conference on Computer and Information Technology (ICCIT). 2023;1–6.
  • 39.Monowar MM, Nobel S, Afroj M, Hamid MA, Uddin M, Kabir M et al. Advanced sleep disorder detection using multi-layered ensemble learning and advanced data balancing techniques. Front Artif Intell. 2025 Jan 28;7. [DOI] [PMC free article] [PubMed]
  • 40.Tyagi PK, Agarwal D. Systematic review of automated sleep apnea detection based on physiological signal data using deep learning algorithm: a meta-analysis approach. Biomed Eng Lett. 2023 Aug 1;13(3):293–312. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.La Fisca L, Jennebauffe C, Bruyneel M, Ris L, Lefebvre L, Siebert X, et al. Enhancing OSA assessment with explainable AI. Annu Int Conf IEEE Eng Med Biol Soc. 2023 July;2023:1–6. [DOI] [PubMed]
  • 42.Abd-Alrazaq A, Aslam H, AlSaad R, Alsahli M, Ahmed A, Damseh R et al. Detection of sleep apnea using wearable AI: systematic review and Meta-Analysis. J Med Internet Res 2024 Sept 10;26:e58187. [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

12911_2025_3129_MOESM1_ESM.docx (17.4KB, docx)

Supplementary Material 1: Additional File 1: Data: Supplementary Table 1: Search query for each database, Supplementary Table 2: List of Abbreviations.

12911_2025_3129_MOESM2_ESM.xlsx (25KB, xlsx)

Supplementary Material 2: Additional File 2: Data: 1st Author, Source of Data (EEG, ECG, airflow, SpO2, etc.), Midpoint (percentage), Accuracy range (percentage), Number of patients, Train/Validation/Test, Age (Mean±SD), Male/Female, Note (Model Architecture (e.g., DL, ML or Hybrid and Code Availability), Country, Dataset.

Data Availability Statement

Data is provided within the manuscript or supplementary information files.


Articles from BMC Medical Informatics and Decision Making are provided here courtesy of BMC

RESOURCES