Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Sep 10, 2025
Open Peer Review Period: Sep 16, 2025 - Nov 11, 2025
Date Accepted: Aug 4, 2026
(closed for review but you can still tweet)
GPT-4o Powered Pre-Anesthetic AI: Development and Validation
ABSTRACT
Background:
Accurate pre-anesthetic assessment is essential for perioperative risk stratification, but conventional tools such as the American Society of Anesthesiologists (ASA) physical status classification and postoperative nausea and vomiting (PONV) risk scores may be affected by subjective judgment, incomplete documentation, and fragmented clinical data. Large language models may support pre-anesthetic assessment by integrating structured and unstructured clinical information.
Objective:
This study aimed to develop and retrospectively validate a GPT-4o–powered artificial intelligence (AI) system for pre-anesthetic assessment. The primary validation focus was agreement between AI-generated and clinician-assigned ASA physical status classifications. PONV risk stratification was evaluated as an additional clinically relevant performance outcome. A secondary objective was to assess the incremental contribution of National Health Insurance (NHI) cloud data to model performance.
Methods:
This single-center retrospective validation study included 600 surgical patients randomly sampled from 4404 eligible inpatient surgical patients between January and May 2025. The system processed structured and unstructured electronic health record data, with and without NHI cloud data, to generate pre-anesthetic assessment outputs. ASA agreement was assessed using Cohen’s kappa with 95% confidence intervals (CIs). PONV risk stratification was evaluated using sensitivity, specificity, positive predictive value, negative predictive value, confusion matrix analysis, chi-square testing, and Cramér’s V.
Results:
Among the 600 patients, clinician-assigned ASA classifications were ASA I in 39 patients, ASA II in 498 patients, ASA III in 57 patients, and ASA IV in 6 patients. Agreement between AI-generated and clinician-assigned ASA classifications was high when NHI data were incorporated (κ=0.883, 95% CI 0.841–0.946), whereas agreement was lower without NHI data (κ=0.518, 95% CI 0.412–0.691). A total of 49 patients (8.2%) experienced documented PONV within 24 hours after surgery. Using the High risk category as the primary test-positive threshold, the model achieved sensitivity of 34.7% (95% CI 22.9–48.7), specificity of 99.1% (95% CI 97.9–99.6), positive predictive value of 77.3%, and negative predictive value of 94.5%. The association between predicted PONV risk group and observed PONV outcome was statistically significant (χ²₂=169.25, P<.001; Cramér’s V=0.531).
Conclusions:
The GPT-4o–powered pre-anesthetic assessment system demonstrated feasibility and high agreement with clinician-assigned ASA classifications when NHI cloud data were incorporated. For PONV risk stratification, the system showed high specificity and high negative predictive value but modest sensitivity, suggesting greater usefulness for identifying low-risk patients than for detecting all patients who may develop PONV. These findings support the use of large language model–based decision-support tools as adjuncts to, rather than replacements for, anesthesiologist judgment.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.