Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Sep 18, 2025
Date Accepted: Jul 7, 2026
Prediction of Postoperative Vomiting Within 24 Hours Using Machine Learning With Large Language Model–Enhanced Interpretability: Development and Validation Study
ABSTRACT
Background:
Postoperative nausea and vomiting (PONV) is a common complication after anesthesia. However, vomiting represents a clinically distinct and objectively measurable endpoint.
Objective:
This study aimed to develop and internally validate predictive models for postoperative vomiting within 24 hours (POV 24h) using structured perioperative data and unstructured clinical text, while introducing a structured framework that separates feature construction from interpretability using large language models (LLMs).
Methods:
We analyzed 33,460 anesthesia records from a single center (2019–2022). Two temporally defined prediction tasks were constructed to reflect real-world clinical decision-making and prevent information leakage: a preoperative model using variables available before anesthesia induction, and a perioperative model using variables available up to the end of surgery. Structured data were modeled using machine learning algorithms (Logistic Regression, Random Forest, XGBoost, LightGBM). Unstructured clinical text was incorporated through a deterministic, concept-driven preprocessing pipeline, where LLMs were used solely for normalization (temperature = 0) without feature generation, followed by rule-based concept mapping and feature encoding. Post-hoc interpretability was further supported using an LLM-based QAChain module. Model performance was evaluated using ROC-AUC, PR-AUC, calibration metrics (Brier score), and threshold-based measures.
Results:
A total of 33,460 surgeries were included, with 3,607 POV events (10.8%). The preoperative model achieved ROC-AUC of 0.7173 and PR-AUC of 0.2049, while the perioperative model achieved ROC-AUC of 0.7200 and PR-AUC of 0.2095. Calibration analysis showed Brier scores of 0.0897 and 0.0833, respectively. Incorporating text-derived features provided modest improvements, while LLM-based explanation modules enhanced interpretability without substantially improving predictive performance.
Conclusions:
Machine learning models can effectively predict postoperative vomiting within 24 hours using perioperative data. The proposed framework demonstrates that LLMs can be integrated in a controlled and reproducible manner—restricted to deterministic normalization and post-hoc reasoning—thereby enhancing interpretability without introducing information leakage or altering predictive modeling. This design supports clinically grounded, transparent decision-making in perioperative risk assessment.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.