Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Jan 29, 2026
Date Accepted: Jul 16, 2026
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Understanding Surgical Complications in Clinical Text: Automated Clavien–Dindo Grading using a Zero-Shot Large Language Model Approach in a Collective of Liver Surgery Patients
ABSTRACT
Background:
The standardized extraction of postoperative complications from unstructured routine clinical documentation remains a major unresolved challenge in digital surgery and health informatics. Although the Clavien–Dindo classification is the established standard for grading postoperative complications, its application in routine clinical documentation is largely implicit and unstructured.
Objective:
To assess the capability of open-weight and proprietary large language models (LLMs) to classify postoperative complications according to the Clavien–Dindo system using discharge letters, benchmarked against expert consensus annotations.
Methods:
We analyzed discharge letters from 650 surgical cases of patients who underwent hepatobiliary surgery between 2010 and 2024. The cohort included Grade I–II complications in 24%, Grade III–IV in 19%, and Grade V (death) in 6% of patients. Representative open-weight (Qwen 3, Llama 3.3, Ministral 3, GPT-OSS) and proprietary (GPT 5.1, Gemini 3 Pro) LLMs were prompted to infer complication grades directly from the discharge letters. Model performance was evaluated against expert assessment using accuracy and a detailed deviation analysis.
Results:
All models were capable of identifying and classifying complications from the unstructured documentation in the discharge letters. On the full 650-case dataset, open-weight models achieved accuracies up to 0.775 for fine-grading prediction and 0.94 for binary classification. On a balanced 50-case subset, proprietary models achieved the highest performance, with accuracies of 0.78 for fine-grading and 0.98 for binary classification. In contrast, open-weight models reached accuracies up to 0.70 in fine-grading and 0.94 for binary grading, with lower computational requirements. An ensemble approach yielded additional gains in classification performance.
Conclusions:
LLMs can accurately classify postoperative complications from discharge letters, enabling scalable and objective monitoring of surgical outcomes. Their use may reduce manual abstraction workload and promote consistent, data-driven quality assessment in surgical care.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.