Abstract
Purpose
obstructive sleep apnea is underdiagnosed due to limited access to polysomnography (PSG). We aimed to assess the performances of Apneal®, an application recording sound and movements thanks to a smartphone’s microphone, accelerometer and gyroscope, to estimate patients’ apnea-hypopnea index (AHI).
Methods
monocentric proof-of-concept study with a first manual scoring step, then automatic detection of respiratory events from recorded signals using a sequential deep-learning model (version 0.1 of Apneal® automatic scoring of respiratory events, end 2022), in adult patients.
Results
46 patients (women 34%, BMI 28.7 kg/m²) were included. Sensitivity of manual scoring was 0.91 (95% CI [0.8-1]) for IAH > 15 and 0.85 [0.67-1] for AHI > 30, and positive predictive values (PPV) 0.89 [0.76–0.97] and 0.94 [0.8-1]. We obtained an AUC-ROC of 0.85 (95% CI [0.69–0.96]) and AUC-PR of 0.94 (95% CI [0.84–0.99]) for the identification of AHI > 15, and AUC-ROC of 0.95 [0.860.99] and AUC-PR of 0.93 [0.81–0.99] for AHI > 30. The ICC between the AHI estimated manually, and from the PSG is 0.89 (p = 6.7 × 10− 17), Pearson correlation 0.90 (p = 1.25 × 10− 17). Automatic scoring found sensitivity of 1 [0.95-1], PPV of 0.9 [0.8–0.9] for AHI > 15, and sensitivity 0.95 [0.84-1], PPV 0.69 [0.52–0.85] for AHI > 30. The ICC between the estimated AHI, and PSG scorings is 0.84 (p = 5.4 × 10− 11) and Pearson correlation is 0.87 (p = 1.7 × 10− 12).
Conclusion
Manual scoring of smartphone-based signals is possible and accurate compared to PSG-based scorings. Automatic scoring method based on a deep learning model provides promising results.
Trial registration
Keywords: Sleep apnea, Artificial intelligence, Deep learning, Smartphone, Diagnosis
Introduction
Sleep apnea-hypopnea syndrome (SAHS) affects 15–25% of the population in developed countries [1, 2]. It is characterized by repeated interruptions of ventilation during sleep, resulting in sleep disturbance, chronic intermittent hypoxia, and hypercapnia. SAHS triggers cardiovascular diseases and excessive daytime sleepiness, which can result in occupational and road traffic accidents. Proper treatment of SAHS, mostly by continuous positive airway pressure (CPAP), reduces the risk of accidents and injuries [3].
Diagnosis of SAHS is based on a polysomnographic recording (PSG) of respiratory- (chest belt, abdominal belt, nasal flow, nasal pressure, snoring, oxygen saturation), cardiac activity (heart rate, photoplethysmogram, electrocardiogram), brain- (EEG), and motor activities (EMG, actimeter, position sensor). While PSG is considered as the reference to precisely evaluate the SAHS severity of a patient, in clinical practice in adults, a simplified ambulatory polygraph recording (PG) is generally sufficient for the diagnosis of SAHS and records at least the following sensors: abdominal and thoracic belts, nasal scope, microphone, pulse oximeter, position, and oxygen saturation [4]. The metric used to evaluate the severity of SAHS is the Apnea Hypopnea Index (AHI), which is the number of apneas and hypopneas per hour (of sleep). Patients will be considered normal (without SAHS) if their AHI is under 5, as having mild SAHS if their AHI is above 5 and below 15, as having moderate SAHS if their AHI is between 15 and 30, and severe SAHS if their AHI is above 30 [5].
However, SAHS screening and diagnosis are too infrequent, and a large number of apneic patients are unaware of their condition. For example, 80% of apnea sufferers in the USA remain undiagnosed [6]. The main obstacles to diagnosis are patients’ awareness, general practitioners’ lack of training in this pathology, the cost of diagnosis and treatment, and access to diagnostic tests [7].
For screening purposes in the general population, some questionnaires can be used like the STOP-BANG [8] or the Berlin questionnaire [9]. While these questionnaires are easy to use, their screening results are either too sensitive or do not show enough sensitivity [8, 9].
Recently, to obtain more accurate results than questionnaires, alternative signals have been proposed, with devices that offer a simpler installation than a PSG, and are used at patients’ homes. For example, some devices use tracheal sounds [10], mandibular movement [11], or movement detected through sleep mattresses [12]. Some of these are validated of SAHS diagnosis or screening, and offer multinight-testing, limiting the risk of misdiagnosis [13]. However, all these devices require dedicated hardware.
To overcome this need of a specific device, inducing economic and environmental questions, and to take advantage of the powerful sensors that exist in smartphones, researchers have considered methods to screen SAHS either using smartphones’ accelerometers [14] or the sound recorded by a smartphone’s microphone [15], and obtained promising but insufficient results. The use of smartphone sensors seems also interesting in reconstructing other vital signs that can be useful to doctors when screening for SAHS, like heart rate [16].
In this paper, we propose a new method that uses no hardware but smartphone’s available sensors to estimate patients’ AHI, called Apneal®. This application records the sound using a smartphone’s microphone and movements thanks to a smartphone’s accelerometer and gyroscope. This proof-of-concept study was conducted with a first manual scoring step, and then automatic detection of respiratory events from the recorded signals using a sequential deep-learning model which was released internally at Apneal® at the end of 2022 (version 0.1 of Apneal® automatic scoring of respiratory events). The principal objective of the study was to assess the metrology validity of Apneal® in comparison to polysomnography, the gold standard.The present study therefore proceeds in two steps: (i) demonstrate that respiratory events can be identified manually from smartphone signals, establishing interpretability; (ii) evaluate a fully automated algorithm designed for population-scale screening.
Methods
Study design
We conducted a transversal study at a university hospital sleep center (Centre du sommeil, hôpital Bichat Claude Bernard, Paris, France). Consecutive adult patients referred for videopolysomnography (VPSG) as part of routine care were offered to participate. Other inclusion criteria were ability to understand French, health insurance coverage. Exclusion criteria were known cardiac rhythm disorder, pacemaker, diabetes mellitus, CPAP treatment during the night of the study, or refusal to participate. In this proof-of-concept study, no power calculation was made. A total of 50 patients were planned for inclusion.
Data acquisition
After obtaining informed consent, patients were equipped with both devices: the polysomnography (PSG) sensors were installed and a smartphone in airplane mode (iPhone 12 mini) was attached to their chest thanks to an adhesive band. The Apneal® application was then switched on. The smartphone was placed with a microphone directed towards the mouth of the patient, and the screen facing up (see Fig. 1).
Fig. 1.
full installation of a smartphone and the PSG on a patient
Full VPSG (Alice 6, Philips Respironics) was performed and scored following American Association of Sleep Medicine (AASM) 2012 guidelines by trained technologists and doctors [17]. We named this first scoring of the PSG PSG-first. After this first scoring, the scoring of the respiratory events was verified and modified if needed by other trained technologists. T his second version of the scoring of the PSG is referred to as PSG-second.
The installation of both devices on the patients are visible on Fig. 1.
The Apneal® application recorded the sound (8000 Hz) produced by the patients during the night, along with the 3D acceleration (100 Hz) and the 3D angular velocity (100 Hz) of the smartphone.
Proprietary APNEAL algorithm.
Proprietary AI algorithms were applied to the raw signals recorded by the smartphones, to extract the position of the patients during the night, the probability of breathing, and the probability of the presence of snores. The sound (hearable by the scorers), the sound volume, the filtered acceleration on the three axes, and the activity were also extracted from the signals recorded by the smartphone.
The activity was extracted with the following method: the algorithm took the accelerometer L2-norm of the x and y axes, then the first derivative, then the absolute values of the first derivative. Value was then multiplied by ten, and clipped to [0, 10] g/s.
AI model, training and validation
While Table 1 lists the complete set of smartphone-derived signals made available to manual scorers, the first automatic-scoring prototype (Apneal® v0.1) deliberately works with a paired-down input: the heart-rate time-series and the acceleration envelope, both resampled at 10 Hz and segmented into consecutive, non-overlapping 3-minute windows (1 800 samples). Each channel passes through six one-dimensional convolutional layers (ReLU activation), the resulting feature maps are concatenated and processed by a bidirectional gated recurrent unit, and a final 1-D convolution with softmax activation produces, every second, the probability of a respiratory event versus normal breathing.
Table 1.
Signals used by scorers
| Variable | Sample rate | Resolution |
|---|---|---|
| 3D acceleration, filtered using a bandpass filter between 0.1 Hz and 1 Hz | 100 Hz | 4.10− 5 g |
| Position (left, right, supine, prone or upright) | 0.033 Hz | - |
| Probability of breathing | 3.2 Hz | 3.10− 5 |
| Snoring probability | 3.2 Hz | 3.10− 5 |
| Audio | 8000 Hz | 16-bit |
| Audio power | 40 Hz | 4.10− 3dB |
| Activity | 1 Hz | 2.10− 3 |
| ECG Heart rate | 1 Hz | 2.10− 3bpm |
Model optimisation employs weighted binary cross-entropy (event to non-event weight 5 : 1) and stochastic gradient descent initialised at 1 × 10⁻⁵. The learning rate is halved whenever validation loss fails to improve for five epochs, and training stops if no improvement is seen for ten consecutive epochs.
Generalisability was assessed with a patient-level four-fold cross-validation: the 44 recordings were split once into four folds balanced for sex, BMI and age (see Table 2) as well as for respiratory-event sub-types. In each iteration two folds served for training, one for validation and the remaining fold for testing, ensuring that every patient appeared exactly once as unseen data.
Table 2.
Patients’ characteristics
| Age | Sex | BMI | N3 (min) | REM% | Sleep time (min) | AHI (PSG-first) | AHI (PSG-second) |
|---|---|---|---|---|---|---|---|
| med : 57.5, IQR: 18 | F 34% M 66% | med: 28.7, IQR: 7 | med: 78.8, IQR: 48.6 |
med: 18.5% IQR: 8.7% |
med: 371 IQR:82 |
med 26.0 IQR: 31.1 | med 27.6 IQR: 32.1 |
Pre-processing (per-patient z-scaling) and training were carried out with TensorFlow 2.8 on a single NVIDIA T4 GPU.
Manual scoring of Apneal® recordings.
The signals in Table 1 (all extracted from Apneal® recording apart from the ECG heart rate) were used by seven non-expert scorers (employees of Mitral) to blindly score sleep stages (30 s epochs, wake and sleep periods only) and respiratory events (central or obstructive apneas and hypopneas, and RERAs), following common guidelines. An illustration of the kind of signals they used can be seen during various respiratory events in Fig. 2 (top) along with some of the signals recorded by the PSG during the same events (bottom). All scorers went through the following steps for each patient: First, look for when the patient falls asleep for the first time and wakes up for the last time to set up the limits of the studied period. Scorers were then asked to identify intra-sleep awakening periods (labeled as wake epochs if they are longer than 30 s) and arousals (less than 30 s long, and at least 3 s long). Once this was done for the whole studied period, scorers were asked to go back to the beginning and look for respiratory events (wake epochs could also be added during this step) and finally to score all the epochs not scored as wake periods as sleep periods. Scorers were encouraged to listen to the audio when in doubt.
Fig. 2.
Illustration of various respiratory events: central apnea, obstructive hypopnea, and obstructive apnea (from left to right), with, in parallel, some of the Apneal® signals used during manual scoring (top) and signals extracted from the PSG (bottom)
All seven candidate scorers first attended a 1-hour standardized tutorial. Immediately afterwards each candidate took a certification exam on one pre-selected Apneal recording (“certification record”, patient ID fixed before any analysis). A scorer was certified if the Apneal®-derived AHI deviated by ≤ ± 10% from the PSG-first AHI for that record. Four of seven candidates met this criterion and are hereafter referred to as certified scorers. For the remaining 43 patients, each recording was scored exactly once by a single, randomly assigned scorer (certified or uncertified). The certification record was the only one scored by all seven scorers; for all performance analyses its manual Apneal® AHI was randomly selected from the four certified scorers.
The scorings produced by this process are considered as Apneal® manual scorings.
Automatic scoring of respiratory events.
In parallel, a sequential deep learning model was also used to automatically detect respiratory events during sleep periods, using input features extracted from Apneal® recordings. The sequential model was trained using cross-validation, with four folds of training, validation, and test sets, containing different patients’ recordings. The scorings produced by this process are considered as Apneal® automatic scorings.
Data processing
After manual scoring, raw VPSG data on Alice 6 were extracted. After manual and automatic scorings of Apneal® recordings, the scorings of sleep stages and respiratory events of the different sources were extracted. See Fig. 3 for a summary of our data processing pipeline.
Fig. 3.
Explanation of our data processing pipeline
For each patient and each setting, we extracted the Apnea and Hypopnea Index (AHI) found, which is the number of hypopneas and apneas during sleep periods divided by the number of hours of sleep. We compared the PSG-first AHI with the AHI obtained using the Apneal® manual, as they were done in a similar setting (one scorer per recording, one reading of the signal). Thus, every patient contributes one manual Apneal AHI value; for the certification patient this value is the random certified score described above. We then compared PSG-second with Apneal® automatic scorings.
We also compared the respiratory events segmentation produced by these different methods.
Statistics
Patient characteristics were described as a median and interquartile range for quantitative variables and percentages for categorical variables.
We studied the ability of the Apneal® device to detect moderate-to-severe SAHS, i.e. classify patients under and above the AHI threshold of 15/h, the current threshold for SAHS treatment (SPLF RPC 2009). For that, we computed positive predictive value and sensitivity. These values were obtained using the raw values of the predicted AHI and the threshold of value 15/h. To obtain 95% confidence intervals for these values, we used bootstrapping over patients, with N = 10,000.
We also computed the Area Under the ROC Curve (AUC-ROC) and the Area Under the Precision-Recall Curve (AUC-PR) for the identification of patients with an AHI above 15/h, and above 30/h. For the automatic scoring, we estimated the minimal threshold to set to obtain a precision of at least 0.9 and of at least 0.95 for the identification of patients above 15/h and above 30/h of AHI.
We computed the intraclass correlation coefficient (ICC) and Pearson correlation between PSG-scored AHI values and those obtained by the manual (with PSG-first) and automatic (with PSG-second) Apneal® scorings and provided the p-value for these values to assess the performances of the App.
We also compared the segmentation of respiratory events obtained from the PSG and the one obtained by the manual (with PSG-first) and automatic (with PSG-second) Apneal® scorings. We did that using positive predictive value and sensitivity metrics, considering a predicted event as a true positive when it overlaps with a PSG-based respiratory event for at least three seconds, with a margin for PSG-based respiratory events of 20 s. To obtain 95% confidence intervals for these values, we used bootstrapping over patients, with N = 10,000.
Ethical aspects
This study is part of the Evaluation of the Metrological Reliability of Connected Objects in the Measurement of Medical Physiological Parameters (EvalExplo) study [NCT03803098]. Ethics approval was obtained from Comité de Protection des Personnes Sud Est VI (approval number AU 1443), and written non-opposition was obtained according to the Jardé decree in France.
Results
Patient characteristics
Due to a change in French legislation on research during the study, the final number of 50 patients could not be reached at the end of the inclusion period, and this period could not be extended. Thus, a total of 46 patients were included. One patient was excluded because of a recording issue. Another patient was excluded because the phone placed on his chest was removed during the night, and put back in the wrong direction. The final sample included 44 patients with available PSG and Apneal® recordings. Patients’ characteristics are described in Table 2.
Manual scoring
For classifying patients below and above AHI 15, the manual Apneal® scorings compared to the PSG scorings (PSG-first) lead to a sensitivity of 0.91 (95% CI [0.8, 1]), and a positive predictive value of 0.89 (CI 95% [0.76, 0.97]). For the AHI threshold of 30, we obtained a sensitivity of 0.85 (95% CI [0.67, 1]), and a positive predictive value of 0.94 (CI 95% [0.8, 1]).
Using varying thresholds on the predicted AHI we obtained an AUC-ROC of 0.85 (95% CI [0.69, 0.96]) and an AUC-PR of 0.94 (95% CI [0.84, 0.99]) for the identification of patients with an AHI above 15, and AUC-ROC of 0.95 (95% CI [0.86, 0.99]) and AUC-PR of 0.93 (95% CI [0.81, 0.99]) for the identification of patient with an AHI above 30.
The ICC between the AHI estimated manually, and the one obtained from the PSG scorings is 0.89 (p-value = 6.7 × 10− 17), Pearson correlation is 0.90 (p-value = 1.25 × 10− 17). Bland-Altman analysis showed a mean bias of + 2.70 events·h⁻¹ with 95% limits of agreement from − 19.65 to + 25.05 (Fig. 4C).
Fig. 4.
Results of the manual scoring. (A) Confusion matrix between the SAHS severities obtained from the PSG AHI and the Manual Apneal® AHI. (B) Regression plot between the PSG AHI and theManual Apneal® AHI. (C) Bland Altman plot between the PSG AHI and the Manual Apneal® AHI
Confusion matrix, regression plot, and Bland Altman plot can be seen in Fig. 4. Results obtained at the event segmentation level are visible in Table 3.
Table 3.
Event-per-event segmentation of respiratory events results for all scoring types produced from Apneal® recordings compared to PSG-based scorings (PSG-first for manual and PSG-second for automatic). The values here represent the ability of Apneal® to identify each individual respiratory event during a patient’s night. The values displayed are median and 95% CI
| Scoring type | PPV for segmentation | Sensitivity for segmentation |
|---|---|---|
| Manual scorers | 0.68 [0.6, 0.74] | 0.7 [0.61, 0.77] |
| Manual certified scorers | 0.7 [0.62, 0.76] | 0.74 [0.66, 0.8] |
| Automatic algorithm | 0.69 [0.6, 0.76] | 0.77 [0.71, 0.82] |
Manual scoring by certified scorers.
Among the seven candidates, four passed the predefined certification exam (see Methods) and are hereafter analysed as certified scorers.
We studied the results of Apneal® manual scorings for the subset of recordings that were scored by these certified scorers. Inter-rater agreement among certified scorers on the certification record was excellent (maximum absolute AHI difference < 5 events·h⁻¹). Among certified scorers, we observed a sensitivity of 0.96 (95% CI [0.88, 1]) and a positive predictive value of 0.93 (95% CI [0.81, 1]) for the AHI threshold of 15. For the AHI threshold of 30, we obtained a sensitivity of 0.87 (95% CI [0.64, 1]), and a positive predictive value of 0.99 (CI 95% [0.95, 1]).
Using varying thresholds on the predicted AHI we obtained an AUC-ROC of 0.95 (95% CI [0.84, 1]) and an AUC-PR of 0.99 (95% CI [0.95, 1]) for the identification of patients with an AHI above 15, and AUC-ROC of 0.96 (95% CI [0.87, 1]) and AUC-PR of 0.96 (95% CI [0.86, 1]) for the identification of patients with an AHI above 30.
The ICC between the AHI estimated manually by certified scorers and the one obtained from the PSG (PSG-first) is 0.93 (p-value = 2.6 × 10− 16), and the Pearson correlation is 0.95 (p-value = 1.8 × 10− 17). The corresponding limits of agreement (Bland Altman analysis) for certified scorers were − 12.73 to + 20.44 events·h⁻¹ around a bias of + 3.85 events·h⁻¹ (Fig. 5C).
Fig. 5.
Results of the manual scoring from certified scorers. (A) Confusion matrix between the SAHS severities obtained from the PSG AHI and the Manual certified Apneal® AHI. (B) Regression plot between the PSG AHI and the Manual certified Apneal® AHI. (C) Bland Altman plot between the PSG AHI and the Manual certified Apneal® AHI
Confusion matrix, regression plot, and Bland Altman plot can be seen in Fig. 5. Results obtained at the event segmentation level are visible in Table 3.
Automatic scoring
For classifying patients below and above AHI 15, the automatic Apneal® scorings compared to the PSG scorings (PSG-second) lead to a sensitivity of 1 (95% CI [0.95, 1]), and a positive predictive value of 0.9 (95% CI [0.8, 0.9]) for the AHI threshold of 15. For the AHI threshold of 30, we obtained a sensitivity of 0.95 (95% CI [0.84, 1]), and a positive predictive value of 0.69 (CI 95% [0.52, 0.85]). The confusion matrix, regression plot, and Bland Altman plot can be seen in Fig. 6.
Fig. 6.
Results of the automatic scoring. (A) Confusion matrix between the SAHS severities obtained from the PSG-second AHI and the automatic Apneal® AHI, using the severity threshold of IAH (< 5, 5–15 and > 15/h). (B) Regression plot between the PSG-second AHI and the automatic Apneal® AHI. (C) Bland Altman plot between the PSG-second AHI and the automatic Apneal® AHI
Using varying thresholds on the predicted AHI we obtained an AUC-ROC of 0.85 (95% CI [0.64, 0.99]) and an AUC-PR of 0.97 (95% CI [0.9, 1]) for the identification of patients with an AHI above 15, and AUC-ROC of 0.87 (95% CI [0.74, 0.96]) and AUC-PR of 0.88 (95% CI [0.74, 0.96]) for the identification of patients with an AHI above 30.
We modified the thresholds used on the predicted AHI to identify patients’ severity in order to obtain optimal predictive positive values. We obtained the updated following metrics. Using the smallest threshold to obtain a positive predictive value above 0.9: for 15 sensitivity of 0.97 (95% CI [0.91, 1]) and positive predictive value of 0.9 (95% CI [0.79, 0.98]); for 30 sensitivity of 0.57 (95% CI [0.35, 0.78]) and positive predictive value of 0.93 (95% CI [0.75, 1]).
The ICC between the AHI estimated, and the one obtained from the PSG scorings is 0.84 (p-value = 5.4 × 10− 11) and the Pearson correlation found is 0.87 (p-value = 1.7 × 10− 12). For automatic scoring, the mean bias was − 6.72 events·h⁻¹ with limits − 32.39 to + 18.95 events·h⁻¹ (Fig. 6C).
Results obtained at the event segmentation level are visible in Table 3.
Discussion
This study presents the first steps for the validation of a new, minimally invasive tool to diagnose sleep apnea. Using only the smartphone’s sensors (microphone, accelerometer, and gyroscope), respiratory events were detected properly. The signal recorded by smartphone sensors is enough for non-expert scorers to score respiratory events, and to detect apneic patients, with a sensitivity of 0.91, and a positive predictive value of 0.89 for the AHI threshold of 15. Moreover, this scoring could be automated, using Apneal® version 0.1 automatic scoring of respiratory events obtaining a sensitivity of 1 and a PPV of 0.9. The predicted AHI values seem to be relevant to identifying the SAHS severity of patients (AHI threshold of 15 and 30, the only ones that can be tested considering the AHI distribution of patients that were included in this study), as we obtained an AUC-ROC of 0.95 (threshold 15) and 0.96 (threshold 30) for the manual scoring (certified scorers) and an AUC-ROC of 0.85 (threshold 15) and 0.87 (threshold 30) for the automatic scoring. These results are encouraging and provide us with a proof of concept of this method.
Furthermore, interestingly, this method does not just provide with a global AHI but enables us to segment individual respiratory events accurately (see Table 3), as we show that manual scoring can reach up to 0.7 of PPV and 0.74 of sensitivity for individual respiratory events segmentation, and the automatic scoring 0.69 of PPV and 0.77 of sensitivity. Further work on the algorithm is necessary to increase the global performance of the solution.
Compared to existing solutions, our AI-driven solution shows good performance in screening and diagnosis. Indeed, screening relies on questionnaires such as the STOP-BANG, Berlin Questionnaire. Their sensitivity and specificity are lower than 80% in the general population (Abrishami et al., 2010; Chung & Vairavanathan, 2008), and they are not adapted to specific populations (e.g. pregnant women, children, psychiatry…). Epworth Sleepiness Scale (ESS) is intended to detect excessive daytime sleepiness and is often used although not recommended as a screening tool. Our device, although it would need specific validation in children and pregnant women, represents an easily accessible tool for SAHS objective detection. Other currently available AI-driven solutions include pulse tonometry and mandibular movements. Their diagnostic performances are equivalent to the ones we find in this preliminary study [11, 18].
As far as the cost of diagnosis is concerned, this has been considerably reduced by the systematic introduction of home ventilatory polygraphy, which can be performed in private practice by doctors of various specialties (pulmonologist, ENT specialist, cardiologist, psychiatrist) or by the general practitioner specializing in sleep [19]. Polygraphy involves a limited number of sensors, enabling patients to equip themselves independently and carry out the recording at home. Although one night’s hospitalization is avoided and the cost is, therefore, lower [20], the investment in equipment and the logistics of the examination (loan of equipment, recovery, disinfection, reading of tracings) remain a barrier to scaling up healthcare systems diagnostics capacity. New devices using derivative signals and AI, such as jaw movements (Sunrise®), ballistocardiography embedded in devices placed under the mattress (Withings®) or arterial pulse tonometry (Watchpat®) have shown good performance for sleep apnea detection. In addition, it has been demonstrated with such devices that AHI may vary considerably from one night to another, leading to consider the need to record for several nights to set a proper diagnosis [13]. Such a consideration further pinpoints the need for simplified devices for SAHS diagnosis. However, the later methods still require a dedicated device, inducing the need for an in-person visit (at least to a pharmacy or a healthcare professional) to set up the exam and/or retrieve results, but also addresses the question of device recycling or elimination. Our device, using only the patient’s smartphone, ensures access to sleep apnea diagnosis in remote areas and has a lower carbon footprint.
Strengths and limits of the study
Strengths of our study include the blinded manual scoring to assess the quality of the recorded signals, the use of polysomnography and not polygraphy as a gold standard, and the high number of respiratory events to be detected, thanks to the inclusion of patients with a high probability of sleep apnea.
Manual scoring by briefly trained laypersons confirms that the smartphone signals carry sufficient physiological information for human interpretation; large-scale deployment will rely on the automated algorithm that replicates these human readings.
One of the limitations of our work is the use of the heart rate provided by the ECG of the PSG during manual scorings. The algorithm to extract the heart rate from the accelerometer and the gyroscope was developed subsequently, and thus could not be used for this first step. However, articles on the extraction of heart rate from the seismocardiogram show that an equivalent heart rate can be measured using that type of signal [21]. This would be enough to reproduce these results using only signals extracted from the smartphone. This error reduces to 3.8 bpm during movement-free and artifact-free periods.
Outliers in the results figures for the automatic scoring (Fig. 6) are visible: the automatic scoring tends to overestimate the number of respiratory events for some patients. These outliers may be due to multiple factors: first, the confusion of periodic leg movements with movements due to respiratory recovery, and thus respiratory events like apneas or hypopneas. The distinction between these two types of events will be taken care of by future versions of Apneal® automatic scoring to solve this overestimation issue. Second, the inclusion of patients under betablockers. Indeed, Apneal® uses heart rate variability (HRV) to score respiratory events. In patients with betablockers, HRV is limited, which could lead to an underestimation of the AHI. This underlines the need to record symptoms, to assess the eligibility of patients for Apneal®, and to repeat the exam if symptoms and AHI do not match, as recommended for PSG [17]. Finally, although the phone was placed by a trained technician, we can not exclude the possibility of phone misplacement, which could have led to poorer diagnostic capacity in those patients. Again, this highlights the necessity to repeat recording nights when needed, a possibility easily offered by Apneal ®.
Another limitation is that manual and automatic Apneal® scorings are benchmarked against two slightly different PSG references. Manual Apneal® scores are compared with PSG-first, a single-pass scoring performed by the night technologist, because both involve a human reader who may overlook subtle events. By contrast, the automatic algorithm is compared with PSG-second, the same tracing after senior-technologist quality control, providing the strictest available standard for a fatigue-free machine. We want to underline that these two scorings could be used as a reference. However, as PSG-first and PSG-second are closely aligned (though not identical), this choice cannot inflate the automated results but aligns each Apneal® output with the most appropriate comparator.
This study was performed in a controlled environment (in-lab PSG). This is of importance, since the phone was placed by a trained technician, and not by the patient, but also since patients were alone in the room and no ambient noises were expected. Preliminary results from Apneal show a good recording capacity even at home, although ambient noises are recorded by the phone. The phone is placed upside down to ensure that the microphone is close as possible to the mouth, and filtering of ambient noises is made by the algorithm. The SESAME study enrolling 500 patients who will perform both PSG and Apneal at home is ongoing in France (NCT065778390) and even allows patients to sleep with another person in the bed to assess the real performance of Apneal in an in-home setting, since this is the dedicated purpose of the application.
Last, our sample is relatively small. The monocentric set-up in a reference sleep center allows for a proper gold standard with highly trained technicians and sleep doctors for scoring but induces a selection bias by including patients with severe sleep complaints. This selection bias produces a marked class imbalance—only seven of 44 subjects had an AHI ≤ 15 h⁻¹—which in turn inflates the confidence intervals of specificity and NPV. We therefore focused on prevalence-independent indices such as sensitivity, PPV, ICC and Bland-Altman bias. Bland-Altman analysis showed acceptable mean biases with wide 95% limits ; these seemingly broad limits are driven by a few outliers. Including subjects with a low pretest probability is mandatory in the next steps to ensure proper device validation in a less specific population. As this paper showed the feasibility of using only smartphone sensors to help diagnose sleep apnea, the next steps are to improve the automatic scoring of respiratory events using these signals andto generalize this method on more and diverse patients. Larger multicentric studies that dispense with the “sleep-apnoea suspicion” inclusion criterion are now in progress—SESAME (primary-care recruitment, NCT065778390) and EASY (hospitalisation setting, NCT03803098)—and will provide balanced data for precise estimates of specificity and NPV as well as tighter limits of agreement.These studies will compare PSG (centralised scoring by one expert scorer to assess the low inter agreement between 2 PSG scorers) to Apneal (automatic scoring), with a primary objective focusing on AHI and secondary objectives focusing on the comparison to different screening (ESS, STOP-BANG, Berlin, nocturnal oxymetry) and diagnostic tools (HSAT mostly).
Conclusion
This work introduces a proof of concept demonstrating the potential of smartphone-based recordings of various signals for detecting respiratory events associated with SAHS diagnosis. Our findings indicate that manual scoring of these signals is possible and accurate compared to PSG-based scorings, demonstrating the interpretability of the recorded signals. Additionally, we also present the first version of an automatic scoring method based on a deep learning model, which provides promising results. A larger multicentric validation study, involving subjects with different SAHS severity is required to confirm these results. Further work will also be done to improve the performances of SAHS severity identification in future versions of the automatic scoring algorithm, and the planned multicentre studies will provide the evidence base needed for CE-marking.When CE-marking is achieved, medico-economic studies will allow to insert our solution in the patients’ journey towards a fast and reliable diagnosis of OSAS.
Acknowledgements
The authors are thankful to the patients who agreed to participate.
The authors are thankful to the sleep technicians (Rémi Cellot, Béatrice Guy, Marie-Cécile Flottes, Carine Ecourtemer, Axelle Lefebvre-Roque, Rosine Zana, Claire Petrovic) and to Fedja Kerzabi for managing the Apneal® application and smartphones.
The authors are thankful for manual scoring, to Danica Despotović, Rima El Kosseifi, Séverin Benizri, Anton Prodanet, employees of Mitral SAS.
Funding
This study was funded by Mitral, and MSD France (unrestricted grant to the Assistance Publique Hôpitaux de Paris Foundation). This research is partially supported by the Agence Nationale de la Recherche as part of the “Investissements d’avenir” program (reference ANR-19-P3IA-0001; PRAIRIE 3IA Institute).
Data availability
Data will be made available upon reasonable request and in accordance to GDPR and French Loi Jardé. (french regulation on human research).
Declaration
Ethical approval
All procedures performed in studies involving human participants were in accordance with the ethical standards of the Jardé law in France and with the 1964 Helsinki declaration and its later amendments or comparable ethical standards.
Informed consent was obtained from all individual participants included in the study.
Ethical Approval was obtained from Comité de Protection des Personnes Sud Est VI (approval number AU 1443).
Conflict of interest
Justine Frija is the PI of the EASY study, funded by Mitral SA and EIT Health. Juliette Millet is an employee of Mitral. Emilie Béquignon has no COI. Ala Covali has no COI. Guillaume Cathelain is an employee of Mitral. Josselin Houenou has no COI. Hélène Benzaquen has no COI. Pierre A. Geoffroy has no COI. Emmanuel Bacry has no COI. Mathieu Grajoszex has no COI. Marie-Pia d’Ortho has no COI.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Benjafield AV, Ayas NT, Eastwood PR et al (2019) Estimation of the global prevalence and burden of obstructive sleep apnoea: a literature-based analysis. Lancet Respiratory Med 7:687–698. 10.1016/S2213-2600(19)30198-5 [Google Scholar]
- 2.Balagny P, Vidal-Petiot E, Renuy A et al (2023) Prevalence, treatment and determinants of obstructive sleep Apnoea and its symptoms in a population-based French cohort. ERJ Open Res 9:00053–02023. 10.1183/23120541.00053-2023 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Tregear S, Reston J, Schoelles K, Phillips B (2010) Continuous positive airway pressure reduces risk of motor vehicle crash among drivers with obstructive sleep apnea: systematic review and Meta-analysis. Sleep 33:1373–1380. 10.1093/sleep/33.10.1373 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Berry RB et al The AASM manual for the scoring of sleep and associated events: rules, terminology and technical specifications, 2012th ed. American Academy of Sleep Medicine, Darien, IL
- 5.Meurice J-C, Gagnadoux F (2010) Préambule. Rev Mal Respir 27:S113–S114. 10.1016/S0761-8425(10)70016-4 [DOI] [PubMed] [Google Scholar]
- 6.Watson NF (2016) Health care savings: the economic value of diagnostic and therapeutic care for obstructive sleep apnea. J Clin Sleep Med 12:1075–1077. 10.5664/jcsm.6034 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Ye L, Li W, Willis DG (2022) Facilitators and barriers to getting obstructive sleep apnea diagnosed: perspectives from patients and their partners. J Clin Sleep Med 18:835–841. 10.5664/jcsm.9738 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Pivetta B, Chen L, Nagappa M et al (2021) Use and performance of the STOP-Bang questionnaire for obstructive sleep apnea screening across geographic regions: A systematic review and Meta-Analysis. JAMA Netw Open 4:e211009. 10.1001/jamanetworkopen.2021.1009 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Tan A, Yin JDC, Tan LWL et al (2017) Using the Berlin questionnaire to predict obstructive sleep apnea in the general population. J Clin Sleep Med 13:427–432. 10.5664/jcsm.6496 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Devani N, Pramono RXA, Imtiaz SA et al (2021) Accuracy and usability of acupebble SA100 for automated diagnosis of obstructive sleep Apnoea in the home environment setting: an evaluation study. BMJ Open 11:e046803. 10.1136/bmjopen-2020-046803 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Kelly JL, Ben Messaoud R, Joyeux-Faure M et al (2022) Diagnosis of sleep Apnoea using a mandibular monitor and machine learning analysis: One-Night agreement compared to in-Home polysomnography. Front NeuroSci 16
- 12.Edouard P, Campo D, Bartet P et al (2021) Validation of the withings sleep analyzer, an under-the-mattress device for the detection of moderate-severe sleep apnea syndrome. J Clin Sleep Med 17:1217–1227. 10.5664/jcsm.9168 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Lechat B, Naik G, Reynolds A et al (2022) Multinight prevalence, variability, and diagnostic misclassification of obstructive sleep apnea. Am J Respir Crit Care Med 205:563–569. 10.1164/rccm.202107-1761OC [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Ferrer-Lluis I, Castillo-Escario Y, Montserrat JM, Jane R (2019) Automatic Event Detector from Smartphone Accelerometry: Pilot mHealth Study for Obstructive Sleep Apnea Monitoring at Home. In: 2019 41st Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). IEEE, Berlin, Germany, pp 4990–4993
- 15.Le VL, Kim D, Cho E et al (2023) Real-Time detection of sleep apnea based on breathing sounds and prediction reinforcement using home noises: algorithm development and validation. J Med Internet Res 25:e44818. 10.2196/44818 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Lahdenoja O, Hurnanen T, Iftikhar Z et al (2018) Atrial fibrillation detection via accelerometer and gyroscope of a smartphone. IEEE J Biomed Health Inf 22:108–118. 10.1109/JBHI.2017.2688473 [Google Scholar]
- 17.Berry RB, Budhiraja R, Gottlieb DJ et al (2012) Rules for scoring respiratory events in sleep: update of the 2007 AASM manual for the scoring of sleep and associated events. Deliberations of the sleep apnea definitions task force of the American academy of sleep medicine. J Clin Sleep Med 8:597–619. 10.5664/jcsm.2172 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Van Pee B, Massie F, Vits S et al (2022) A multicentric validation study of a novel home sleep apnea test based on peripheral arterial tonometry. Sleep 45:zsac028. 10.1093/sleep/zsac028 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Safadi A, Etzioni T, Fliss D et al (2014) The effect of the transition to home monitoring for the diagnosis of OSAS on test availability, waiting time, patients’ satisfaction, and outcome in a large health provider system. Sleep Disord 2014:418246. 10.1155/2014/418246 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Stewart SA, Penz E, Fenton M, Skomro R (2017) Investigating cost implications of incorporating level III At-Home testing into a polysomnography based sleep medicine program using administrative data. Can Respir J 2017:8939461. 10.1155/2017/8939461 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Morra S, Hossein A, Gorlier D et al (2019) Modification of the mechanical cardiac performance during end-expiratory voluntary apnea recorded with ballistocardiography and seismocardiography. Physiol Meas 40:105005. 10.1088/1361-6579/ab4a6a [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Data will be made available upon reasonable request and in accordance to GDPR and French Loi Jardé. (french regulation on human research).






