Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Medical Informatics

Date Submitted: Aug 7, 2026
Open Peer Review Period: Aug 20, 2026 - Oct 15, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Policy-Anchored Evaluation of Large Language Model Disposition Recommendations for ED Chest Pain: Single-Institution Feasibility Study (RAPID-ED Phase 0)

  • Carlo Lutz; 
  • Alan Teigman; 
  • Jennifer Zapata; 
  • Kyle Shibuya; 
  • Deborah White; 
  • Benjamin Friedman; 
  • Michael Jones; 
  • Shitij Arora

ABSTRACT

Background:

Large language models (LLMs) are increasingly deployed as clinical decision support systems (CDSS), but the accuracy of their emergency department (ED) disposition recommendations is unclear. We propose measuring not just accuracy but also direction of error and adherence to institutional protocols.

Objective:

To assess the accuracy of policy-anchored LLM disposition recommendations for chest pain presentations to the ED, and to characterize concordance, direction of discordance, and institutional disposition-vocabulary adherence across 5 artificial intelligence (AI) configurations.

Methods:

In a single-institution feasibility study, 31 ED chest pain cases (21 real, 10 synthetic borderline-risk vignettes) were scored on a 5-tier ordinal disposition scale (discharge, observation, floor, telemetry, intensive care/cath lab). An emergency physician (EP) adjudicating against the institutional chest pain policy served as the reference standard (Reviewer 1) and was compared to a second EP’s review to establish ceiling inter-rater reliability. Five AI configurations were asked to provide level of care and campus disposition: an ungrounded general-purpose model (ChatGPT-Edu, GPT-5.5); 3 retrieval-augmented generation (RAG) -like configurations grounded in the 2021 national chest pain guideline and institutional disposition policy, differing in knowledge-delivery prompting; and OpenEvidence, a literature-grounded RAG LLM. The primary metric was linear-weighted Gwet AC2, with 95% CIs by bootstrap. Discordant cases were classified as over- or under-triage; agreement with actual disposition was also measured.

Results:

Physician adjudication agreed on 30/31 cases; inter-rater reliability was AC2 0.98, 95% CI 0.95-1.00. Among LLM configurations, per-case in-chat delivery achieved the highest agreement (AC2 0.53, 95% CI 0.39-0.68) and single-batch delivery the lowest (0.45), with OpenEvidence at 0.49; the ungrounded baseline was lowest overall (0.31, 95% CI 0.10-0.52). Reviewer 1 agreed with actual disposition (n=21) at AC2 0.81 (0.64-0.96), exceeding every configuration. GPT base over-triaged (22 over vs 3 under; telemetry or intensive care in 77% of cases vs 26% for the reference), whereas RAG-batch and OpenEvidence tended to under-triage (6 over vs 14 under).

Conclusions:

No LLM configuration approached the level of physician-physician agreement. Institutional grounding improved concordance but reversed the direction of residual error, from over-triage to under-triage. This error was driven predominantly by an observation pathway not actually present at the institution. Adherence to the institutional disposition vocabulary is a distinct evaluation axis from accuracy; deployment should constrain model output to the local disposition menu.


 Citation

Please cite as:

Lutz C, Teigman A, Zapata J, Shibuya K, White D, Friedman B, Jones M, Arora S

Policy-Anchored Evaluation of Large Language Model Disposition Recommendations for ED Chest Pain: Single-Institution Feasibility Study (RAPID-ED Phase 0)

JMIR Preprints. 07/08/2026:109054

DOI: 10.2196/preprints.109054

URL: https://preprints.jmir.org/preprint/109054

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.