Currently submitted to: Journal of Medical Internet Research
Date Submitted: Aug 13, 2026
Open Peer Review Period: Aug 14, 2026 - Oct 9, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Evaluating human–large language model interaction in clinical prediction rule calculation: a study of four pediatric trauma rules
ABSTRACT
Background:
Clinical prediction rules (CPRs) are widely used in medicine, including for pediatric injury care. For retrospective validation, predictor variable extraction must be accurate and reproducible to avoid bias. However, extracting predictor variables from electronic health record (EHR) clinical notes is labor–intensive, difficult to scale, and prone to reviewer error.
Objective:
We aimed to evaluate a large language model (LLM) for its ability to calculate risk categories and extract predictor variables from emergency department (ED) physician notes in the EHR for four pediatric trauma CPRs. We also sought to characterize human–LLM interaction in clinical data extraction.
Methods:
We conducted a cross–sectional study of 750 de–identified pediatric emergency department visits in order to assess four published trauma CPRs in children: traumatic brain injury (TBI) in children younger than 2 years (TBI in younger), TBI in children older than 2 years (TBI in older), cervical spine injury (CSI), and intra–abdominal injury (IAI). The LLM performed zero–shot extraction using prespecified prompts (development cohort; n=300 visits), and performance was compared with expert reviewer extraction methods (validation cohort; n=450 visits). The primary outcome was LLM performance for CPR calculation and predictor variable extraction, including non–inferiority to expert reviewers using a prespecified 7.5% margin. Secondary outcomes examined the role of human–LLM interaction in reference standard creation. These included the number of expert reviewer extraction errors identified by the LLM, the resulting changes in children's risk categories, and the percentage decrease in extractions requiring full physician review.
Results:
The LLM’s risk category calculation accuracy was 91% (95% CI 85-96) for TBI in younger children, 97% (95% CI 94-100) for TBI in older, 95% (95% CI 91-98) for CSI, and 98% (95% CI 97-100) for IAI. The LLM’s risk category calculation accuracy and sensitivity were non–inferior to expert reviewers for TBI in older children, CSI, and IAI. LLM data extraction accuracy and sensitivity were also non–inferior to expert reviewers for 25 of 27 (93%) predictor variables. During reference standard creation, the LLM identified 34 human extraction errors (20% of corrections made), updated risk categories for 11 children (2.4%), and reduced physician review burden by 72%.
Conclusions:
An interactive human–LLM pipeline achieved expert–level performance for pediatric trauma CPR calculation and predictor variable extraction. Iterative prompt development improved reference standard quality and reduced physician review burden. These results also support a framework for human–LLM interaction spanning autonomous extraction, supervised screening, error identification, and consensus building. Our findings suggest that clinical data extraction is best viewed as an interactive process in which experts and LLMs contribute complementary strengths across multiple stages of extraction.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.