Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Feb 12, 2026
Date Accepted: Jun 12, 2026
Accuracy of Machine Learning Algorithms Based on Electroencephalogram in Sleep Apnea Detection: A Systematic Review and Meta-Analysis
ABSTRACT
Background:
Sleep apnea is a serious sleep disorder, and its diagnostic gold standard, polysomnography (PSG), is costly and time-consuming. Electroencephalogram (EEG) signals, due to their direct correlation with neural activity and ease of extraction, represent a promising tool. Despite increasing research on traditional machine learning (TML) and deep learning (DL) for EEG-based sleep apnea detection yielding significant results, the performance of these models has not been consistently evaluated.
Objective:
The objective of this systematic review was to evaluate the accuracy of machine learning (ML) in detecting sleep apnea from EEG data and provide an evidence base for further clinical application and future research.
Methods:
Following the PRISMA-DTA guidelines, we systematically searched PubMed, Embase, Web of Science, Cochrane Library (CENTRAL), Scopus, IEEE Xplore, and ClinicalTrials.gov databases. Studies evaluating the value of machine learning algorithms for detecting sleep apnea based only on electroencephalogram (EEG) data were included, while those failing to meet inclusion criteria or representing duplicate publications were excluded. The QUADAS-2 and PROBAST+AI tools were used to assess the risk of bias in each study. Data synthesis and statistical analysis were performed using the MIDAS and METANDI modules in Stata 18.0 software and Meta-DiSc 1.4 software.
Results:
A total of 24 retrospective studies were included. The results of the meta-analysis at the event-level indicate that machine learning demonstrated high diagnostic accuracy, with a pooled sensitivity of 0.89 (95% CI: 0.83–0.93; I² = 99.87%) and a pooled specificity of 0.91 (95% CI: 0.87–0.94; I²=99.97%). The diagnostic odds ratio was 79.65 (95% CI: 34.58–183.48), the positive likelihood ratio (LR+) was 9.99(95% CI: 6.39–15.60), and the negative likelihood ratio (LR-) was 0.13 (95% CI: 0.08–0.19). The pooled receiver operating characteristic curve showed an area under the curve (AUC) of 0.96 (95% CI: 0.93–0.99) for the machine learning model in diagnosing sleep apnea, indicating high diagnostic value. Only two patient-level studies met the inclusion criteria. Due to the small sample size, the model could not reliably estimate the pooled effect size, so results from these two studies are described qualitatively.
Conclusions:
Machine learning models based solely on EEG signals demonstrate excellent performance in sleep apnea detection and hold promise as scalable screening or decision support tools. However, current evidence primarily stems from retrospective event-level analyses, which may overestimate diagnostic performance in real-world applications. Further prospective studies with patient-level validation are required before EEG-based machine learning models can be reliably integrated into clinical diagnostic pathways. Clinical Trial: PROSPERO CRD420251244156
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.