Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Jun 10, 2026
Date Accepted: Jul 27, 2026
Artificial Intelligence Models for Predicting Acute Kidney Injury and Post-AKI Mortality: Systematic Review and Meta-Analysis
ABSTRACT
Background:
Machine learning (ML) models are increasingly used to predict acute kidney injury (AKI), but validation quality and clinical readiness remain uncertain.
Objective:
This systematic review and meta-analysis aimed to summarize discrimination performance and implementation-relevant gaps for AKI occurrence and post-AKI mortality prediction.
Methods:
We searched Cochrane Library, Embase, PubMed, and Web of Science through January 23, 2025. Eligible studies developed or validated machine learning-based prediction models and reported the area under the receiver operating characteristic curve (AUC). Two reviewers screened studies, extracted data, and assessed risk of bias using the Prediction Model Risk Of Bias Assessment Tool + Artificial Intelligence (PROBAST+AI). Logit-transformed AUCs were pooled using restricted maximum likelihood (REML) random-effects meta-analysis with Hartung-Knapp-Sidik-Jonkman (HKSJ) adjusted inference.
Results:
We included 219 studies with 7,343,170 participants and 101 modeling approaches. Primary analyses included 188 AUC estimates for AKI occurrence and 31 for post-AKI mortality. Pooled AUCs were 0.834 (95% CI 0.821-0.846) for AKI occurrence prediction and 0.830 (95% CI 0.807-0.851) for post-AKI mortality prediction, respectively. For AKI occurrence, nonlinear approaches, especially deep learning and tree-based or ensemble methods, generally showed higher pooled AUC point estimates than linear or generalized linear models in exploratory subgroup analyses. By study level, PROBAST+AI rated 129 studies (58.9%) as low risk, 84 studies (38.4%) as high risk, and 6 studies (2.7%) as unclear risk. External validation was uncommon: it was reported in 30 (16.0%) AKI occurrence records and 9 (29.0%) post-AKI mortality records.
Conclusions:
Machine learning models have shown high average discrimination for AKI occurrence and post-AKI mortality, supporting their potential value for AKI risk stratification and early warning. However, high levels of heterogeneity, limited external or prospective validation, and inconsistent reporting of model calibration and clinical utility mean that substantial barriers remain before routine clinical deployment. Pooled AUC estimates reveal that nonlinear models have considerable clinical translational potential. Further refinements to modeling frameworks are warranted to explore feasible strategies for real-world clinical implementation. Future studies should prioritize standardized definitions, robust validation, clinically meaningful thresholds, assessment of alert burden, and evidence that model-guided care improves kidney-protective management or patient outcomes. Clinical Trial: PROSPERO CRD420261333545
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.