Currently submitted to: JMIR Medical Informatics
Date Submitted: Jul 14, 2026
Open Peer Review Period: Aug 7, 2026 - Oct 2, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Clinical Concept-Based Evaluation of Automatic Speech Recognition and Large Language Model Pipelines in Trauma Care
ABSTRACT
Background:
Trauma workflows involve rapid and frequent verbal communication, multidisciplinary handovers, and extensive clinical documentation. Recent advances in automatic speech recognition (ASR) and large language models (LLMs) have shown potential for supporting clinical documentation, but many studies focus only on generic transcription accuracy rather than usefulness within clinically meaningful workflows, such as trauma handover.
Objective:
This study aimed to explore the feasibility of ASR with LLM pipelines to support trauma documentation workflows in a New Zealand trauma system.
Methods:
A modular ASR + LLM pipeline was developed in which ASR models transcribed clinical speech and LLMs generated structured clinical outputs. Multiple ASR and LLM combinations were evaluated separately and then combined into a pipeline, using public and curated clinical speech datasets.
Results:
AssemblyAI achieved the strongest overall ASR performance, with WER values of approximately 0.13 and consistently high concept-level F1 scores across datasets. However, clinically important transcription errors remained present in all ASR outputs. LLM-based refinement improved transcript readability and corrected a range of terminology and contextual errors. The greatest improvement was that GPT-5.5 increased F1 scores for clinical terms by up to 0.13. Refined transcripts were subsequently used to generate structured checklist items and support reference transcript curation, demonstrating potential applications within trauma documentation workflows.
Conclusions:
The ASR + LLM pipeline demonstrated potential to support trauma documentation workflows within a New Zealand trauma system across diverse clinical speech datasets and multiple AI models evaluated in this study. By converting handover speech into clearer transcripts and generating structured documentation outputs, the pipeline may help clinicians with various documentation tasks. Rather than operating autonomously or replacing clinical decision-making, the pipeline is intended to serve as an assistant that supports clinician review, reduces documentation burden, and improves clarity in communication and handover between care teams.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.