Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Jan 12, 2026
Date Accepted: Jul 23, 2026
A Multidisciplinary Team-Based Large Language Model Framework for Predicting Postoperative Neurological Complications in Acute Type A Aortic Dissection: Model Development and Validation Study
ABSTRACT
Background:
Postoperative neurological complications (PNC) after surgery for acute type A aortic dissection (ATAAD) lead to poor outcomes. While multidisciplinary team (MDT) consultations are essential for risk assessment, they are resource-intensive and may not be feasible under time constraints.
Objective:
This study aimed to develop an MDT-based large language model (LLM) framework, enhanced by population-level context engineering, for predicting PNC, and to systematically compare its performance with traditional machine learning (ML) models and single-agent LLMs.
Methods:
Patients diagnosed with ATAAD at West China Hospital, Sichuan University from January 2020 to June 2024 (retrospective cohort, n=763) and from July 2024 to June 2025 (prospective cohort, n=120) were used for model development and prospective validation, respectively. The retrospective cohort was randomly allocated to a training set and an internal validation set at a 7:3 ratio. Four traditional ML models were developed. Concurrently, two LLMs (DeepSeek-V3 and GPT-5) were evaluated under single-agent and multi-agent configurations, with and without population-level statistical context engineering. The multi-agent configuration simulated the MDT framework with four role-specialized agents (cardiovascular surgeon, anesthesiologist, neurologist and supervisor) that jointly reason over real-world clinical cases.
Results:
The incidence of PNC was 13.0%. Among ML models, random forest showed the highest discriminative ability (AUC=0.8308). The optimal performance was achieved by GPT-5 utilizing both the MDT framework and context engineering (AUC=0.8419). Context engineering focused the LLMs' reasoning on clinically established risk factors. Similar results and trends were confirmed in the prospective validation. Both the SHAP analysis for ML models and the word frequency analysis for LLMs identified intraoperative peak lactate as the most important predictor.
Conclusions:
An MDT-based LLM framework augmented with population-level context engineering effectively predicts PNC risk, outperforming traditional ML models and providing a transparent, collaborative reasoning process that mimics clinical decision-making.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.