Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR AI

Date Submitted: Aug 20, 2026
Open Peer Review Period: Aug 28, 2026 - Oct 23, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Automating the Comprehensive Complication Index Using Large Language Models: A Methodological Validation Study

  • Pia Borgas; 
  • Christian Nebiker; 
  • Urs Pfefferkorn; 
  • Sebastian Manuel Staubli

ABSTRACT

Background:

The Comprehensive Complication Index (CCI) is a validated measure of postoperative morbidity that quantifies total postoperative complication burden. It remains underutilised because manual Clavien-Dindo Classification (CDC) grading and score aggregation is tedious and susceptible to subjectivity and inconsistency. Although large language models (LLMs) can classify postoperative complications from free-text clinical documentation, their ability to compute the more complex CCI remains unclear.

Objective:

To evaluate the feasibility, accuracy and reproducibility of contemporary LLMs for automated CCI derivation from real-world surgical discharge summaries.

Methods:

Six commercially available LLMs were assessed using a tiered validation framework. After conceptual testing, ChatGPT-5.2 was selected for the main analysis as the model with the highest accuracy in complication extraction and CCI computation. 40 synthetic discharge summaries were analysed under naïve and structured prompting conditions. The prompted pipeline was then applied to 200 retrospectively anonymised real-world surgical discharge summaries. Outputs were compared with human reference values using Bland-Altman analysis, intraclass correlation coefficients, and kappa statistics.

Results:

Compared with naïve analysis, structured prompting increased agreement in LLM-derived vs human reference CCI scores of synthetic cases, reducing mean bias and narrowing the limits of agreement (0.36; LoA -8.96 to 8.24 vs 0.13; LoA -0.90 to 0.65, respectively). Prompted analysis of real-world discharge summaries equally showed a high level of agreement, and limits of agreement remained narrow (mean bias 0.18; LoA -4.71 to 5.07). Exact CCI agreement occurred in 193/200 cases (97%), and highest CDC grade agreement was near-perfect (κw = 0.99).

Conclusions:

Structured prompting results in accurate and reproducible CCI derivation from surgical discharge summaries. LLM-based CCI computation is therefore a promising tool for scalable, automated postoperative morbidity assessment. Multicenter validation is warranted before wider implementation.


 Citation

Please cite as:

Borgas P, Nebiker C, Pfefferkorn U, Staubli SM

Automating the Comprehensive Complication Index Using Large Language Models: A Methodological Validation Study

JMIR Preprints. 20/08/2026:109961

DOI: 10.2196/preprints.109961

URL: https://preprints.jmir.org/preprint/109961

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.