Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Jan 17, 2026
Date Accepted: Aug 14, 2026
Prediction models for in-hospital delirium using routinely collected electronic health record data: a systematic review
ABSTRACT
Background:
Acute in-hospital mental status deterioration is a common and clinically significant manifestation of acute brain dysfunction, associated with increased morbidity, mortality, and long-term cognitive impairment. Prediction models based on routinely collected electronic health record (EHR) data have been proposed to support early identification and targeted prevention. However, the methodological quality, validation rigor, and clinical readiness of these models remain unclear.
Objective:
This systematic review aimed to synthesise and critically evaluate prediction models developed to identify acute in-hospital mental status deterioration using routinely collected EHR data, with a focus on model characteristics, validation strategies, performance, risk of bias, and clinical applicability.
Methods:
A systematic literature search was conducted in PubMed/MEDLINE, Embase, PsycINFO, and Web of Science from inception to 2025/11/11. Studies were eligible if they developed, validated, or evaluated multivariable prediction models using routinely collected EHR or administrative data to predict acute mental status deterioration during adult hospital admissions. Although the eligibility criteria were intentionally broad, all included studies operationalised deterioration as delirium. Data extraction was informed by CHARMS and TRIPOD/TRIPOD-AI guidance. Model performance, validation characteristics, calibration, and implementation features were synthesised narratively. Risk of bias and applicability were assessed using the PROBAST tool.
Results:
Twenty-nine studies met inclusion criteria. Most models were developed using retrospective cohort designs and applied across heterogeneous clinical settings, including general wards, intensive care units, perioperative pathways, and emergency departments. Machine learning approaches—particularly tree-based ensemble models—predominated, while deep learning methods were largely confined to ICU cohorts or large-scale datasets. Internal discrimination was commonly reported, with AUROC values ranging from 0.77 to 0.97; however, external validation was reported in only 12 studies, and calibration assessment in 15 studies. Decision curve analysis was infrequently performed (3 studies), and prospective evaluation or workflow integration was limited. Using PROBAST, overall risk of bias was judged low in 8 studies, unclear in 10, and high in 11, with high risk most frequently driven by limitations in the analysis domain, including apparent-only performance reporting, inadequate handling of missing data, and absence of calibration assessment.
Conclusions:
Prediction models for delirium based on routinely collected EHR data demonstrate technical feasibility and generally favourable discrimination across hospital settings. However, inconsistent validation, limited calibration assessment, and high risk of bias constrain confidence in their generalisability and clinical utility. Future research should prioritise methodological rigor, transparent reporting, robust external and temporal validation, and prospective evaluation within real-world clinical workflows to enable safe and effective clinical deployment.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.