Currently submitted to: Journal of Medical Internet Research
Date Submitted: Sep 22, 2026
Open Peer Review Period: Sep 23, 2026 - Nov 18, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Large Language Models Across the Acute Stroke Care Pathway: A Scoping Review of Applications, Evidence Maturity, and Implementation Readiness
ABSTRACT
Background:
Large language models (LLMs) are being tested at nearly every step of acute stroke care, yet evidence is fragmented across specialties and study types, and the maturity of that evidence along the time-critical clinical pathway has not been mapped.
Objective:
We aimed to map the applications of generative LLMs across the acute stroke pathway—from prehospital recognition to discharge communication—and to grade the maturity of the evidence and its readiness for implementation.
Methods:
We conducted a scoping review following the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) framework, with a protocol registered on the Open Science Framework (DOI: 10.17605/OSF.IO/XDWHS). PubMed, Europe PMC (including preprints), and Google Scholar were searched on August 25, 2026 (records from November 1, 2022 onward). Eligible studies evaluated a generative LLM on a task situated within the hyperacute-to-discharge stroke pathway. Three reviewers independently screened records and charted data using a piloted form, with disagreements resolved by consensus. Each study was mapped to one of five pathway stages and graded on a three-tier evidence-maturity scale (Tier 1: simulation or benchmark; Tier 2: retrospective real-world data; Tier 3: prospective or real-world deployment).
Results:
Of 601 records, 71 studies (published 2023-2026) were included. Applications concentrated in documentation and communication (n=19), imaging-related text tasks (n=18), emergency department triage and diagnosis (n=15), reperfusion decision support (n=13), and prehospital recognition (n=5). Evidence maturity was low: 29 studies were simulation or benchmark evaluations, 39 used retrospective real-world data, and only 3 reported prospective or deployment-stage evidence. Performance was highly polarized—structured data extraction from radiology reports reached 93.5%-94.8% accuracy, whereas zero-shot interpretation of CT images showed poor sensitivity (recall 0.14) and consumer chatbots under-triaged 52% of emergency scenarios. Research output was geographically concentrated (United States, China, Germany, and Türkiye contributed 47/71 studies), GPT-family models dominated (used in 80% of studies), and no included study evaluated regulatory approval or medico-legal liability.
Conclusions:
Generative LLMs show credible near-term value for documentation- and extraction-type tasks in acute stroke care, but remain unproven for triage, image interpretation, and reperfusion decision-making, where prospective, regulated evaluation is essentially absent. Research should now shift from benchmark performance to deployment studies, reporting standards, and regulatory science. Clinical Trial: Protocol is registered on the Open Science Framework (DOI: 10.17605/OSF.IO/XDWHS)
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.