Currently submitted to: JMIR Medical Informatics
Date Submitted: Sep 18, 2026
Open Peer Review Period: Sep 27, 2026 - Nov 22, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Structuring Paper-Based Japanese Emergency Medical Services Verification Records With Multimodal Large Language Models: Retrospective Evaluation Study
ABSTRACT
Background:
Paper-based emergency medical services (EMS) verification records are widely used for prehospital quality assurance but are difficult to reuse for epidemiological analysis, quality improvement, and downstream clinical data integration. Multimodal large language models (LLMs) may enable direct extraction of structured data from scanned paper forms; however, safe use requires field-level error analysis, identification of high-risk fields, and explicit human verification safeguards.
Objective:
This study aimed to develop a pipeline for converting scanned Japanese EMS post-event verification records into structured data sets and to characterize extraction performance, error profiles, and human verification priorities for two cloud-based multimodal LLMs and one locally deployed vision-language model under their respective deployment settings.
Methods:
We conducted a retrospective evaluation study using 630 ambulance activity verification records from five fire departments in Japan, collected from January to December 2024. Gemini 2.5 Pro, Claude Opus 4.6, and Qwen2.5-VL 32B processed the same scanned images using a common prompt and JavaScript Object Notation (JSON) schema covering 183 fields. Two board-certified emergency physicians created the reference standard. We calculated macro accuracy, micro accuracy, restricted-cell micro accuracy, and five-way error counts. Cell-based Wilson confidence intervals were descriptive and did not account for clustering within documents or fields.
Results:
The evaluation included 115,290 document-field cells per model. Preconsensus observed agreement was 94.1% (structured-field Cohen kappa=0.94). Micro accuracy was 89.2% (95% CI 89.0%-89.4%) for Gemini 2.5 Pro, 84.2% (95% CI 84.0%-84.4%) for Claude Opus 4.6, and 80.0% (95% CI 79.8%-80.2%) for Qwen2.5-VL 32B. Restricted-cell micro accuracy was 84.3% (95% CI 84.0%-84.5%), 77.0% (95% CI 76.7%-77.3%), and 70.0% (95% CI 69.7%-70.3%), respectively. Accuracy was highest for numeric and time-and-distance fields and lowest for free-text and resuscitation-related fields. Restricted-cell accuracy in Resuscitation & Treatment was 50.4%, 33.9%, and 11.5%, respectively. Hallucination rates were 1.7% (95% CI 1.6%-1.8%), 2.6% (95% CI 2.4%-2.7%), and 1.7% (95% CI 1.5%-1.8%); omission rates were 7.2% (95% CI 7.0%-7.3%), 12.2% (95% CI 11.9%-12.4%), and 18.2% (95% CI 18.0%-18.5%).
Conclusions:
Multimodal LLMs generated draft structured data from paper-based Japanese EMS verification records, but the results do not support unsupervised clinical or quality-assurance use. Correct empty-cell matches can inflate overall accuracy. Restricted-cell and error-type analyses identify priorities for source-image verification, including free-text, checkbox, circled-selection, and resuscitation-related fields. The proposed workflow requires prospective evaluation of reviewer workload and residual error before deployment.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.