Currently submitted to: JMIR AI
Date Submitted: Aug 20, 2026
Open Peer Review Period: Aug 28, 2026 - Oct 23, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Automating the Comprehensive Complication Index Using Large Language Models: A Methodological Validation Study
ABSTRACT
Background:
The Comprehensive Complication Index (CCI) is a validated measure of postoperative morbidity that quantifies total postoperative complication burden. It remains underutilised because manual Clavien-Dindo Classification (CDC) grading and score aggregation is tedious and susceptible to subjectivity and inconsistency. Although large language models (LLMs) can classify postoperative complications from free-text clinical documentation, their ability to compute the more complex CCI remains unclear.
Objective:
To evaluate the feasibility, accuracy and reproducibility of contemporary LLMs for automated CCI derivation from real-world surgical discharge summaries.
Methods:
Six commercially available LLMs were assessed using a tiered validation framework. After conceptual testing, ChatGPT-5.2 was selected for the main analysis as the model with the highest accuracy in complication extraction and CCI computation. 40 synthetic discharge summaries were analysed under naïve and structured prompting conditions. The prompted pipeline was then applied to 200 retrospectively anonymised real-world surgical discharge summaries. Outputs were compared with human reference values using Bland-Altman analysis, intraclass correlation coefficients, and kappa statistics.
Results:
Compared with naïve analysis, structured prompting increased agreement in LLM-derived vs human reference CCI scores of synthetic cases, reducing mean bias and narrowing the limits of agreement (0.36; LoA -8.96 to 8.24 vs 0.13; LoA -0.90 to 0.65, respectively). Prompted analysis of real-world discharge summaries equally showed a high level of agreement, and limits of agreement remained narrow (mean bias 0.18; LoA -4.71 to 5.07). Exact CCI agreement occurred in 193/200 cases (97%), and highest CDC grade agreement was near-perfect (κw = 0.99).
Conclusions:
Structured prompting results in accurate and reproducible CCI derivation from surgical discharge summaries. LLM-based CCI computation is therefore a promising tool for scalable, automated postoperative morbidity assessment. Multicenter validation is warranted before wider implementation.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.