Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Feb 2, 2026
Date Accepted: Jul 20, 2026
Performance of Artificial Intelligence–Based Screening Tools for Obstructive Sleep Apnea Across Apnea–Hypopnea Index Thresholds: A Systematic Review and Meta-Analysis
ABSTRACT
Background:
Obstructive sleep apnea (OSA) is a prevalent, underdiagnosed condition associated with substantial morbidity and health burden. Given the limited capacity of definitive diagnostic testing in real-world settings, AI-based screening models leveraging diverse data sources have been explored as scalable front-end triage tools. However, the accuracy of these approaches may vary by the apnea–hypopnea index (AHI) threshold used to define disease severity, and their threshold-specific screening performance has not been comprehensively quantified.
Objective:
To systematically review and meta-analyze the diagnostic accuracy of AI-based approaches for adult OSA screening across different AHI cutoffs.
Methods:
A comprehensive literature search of PubMed, Embase, Scopus, and Web of Science was conducted to identify studies published between January 2015 and June 15, 2025 that evaluated artificial intelligence–based approaches for OSA screening. Study selection and data extraction were independently performed by two reviewers, and the QUADAS-2 tool was used to assess study quality. For different AHI thresholds, pooled estimates of sensitivity and specificity were calculated using a bivariate random-effects model, with subgroup analyses undertaken to investigate potential sources of heterogeneity.
Results:
Across 30 studies comprising 41 independent datasets, AI-based screening models showed high diagnostic accuracy for adult OSA across different AHI thresholds. For any OSA (AHI >5), pooled sensitivity and specificity were 0.92 and 0.72, respectively (AUC 0.91), while performance remained robust for moderate-to-severe OSA (AHI >15; sensitivity 0.87, specificity 0.75). For severe OSA (AHI > 30), the pooled results indicated balanced performance, with a pooled sensitivity of 0.83, specificity of 0.85, and an AUC of 0.92. Subgroup analyses indicated clinically meaningful heterogeneity primarily driven by data modality and modeling strategy.
Conclusions:
AI-based screening tools demonstrate robust discrimination for adult OSA across clinically relevant AHI thresholds, with clinically meaningful variation driven by data modality and modeling paradigm. These approaches may serve as scalable front-end triage tools to optimize PSG/HSAT utilization; however, broader prospective multicenter validation and more rigorous reporting remain essential for clinical translation.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.