Artificial Intelligence–Based Structured Information Extraction From Synthetic Nursing Handover Transcripts: Comparative Evaluation of Large Language Models
ABSTRACT
Background:
Clinical handover is a safety-critical point in healthcare, where a clinician's understanding of the patient is shaped by the information communicated before responsibility for care is transferred. Artificial intelligence has the potential to improve the reliability and completeness of clinical handover by helping clinicians verify that key information has been communicated, identify explicit information gaps, and prompt clarification before responsibility is transferred.
Objective:
This study evaluated the performance of several large language models and prompt optimisation strategies for structured information extraction across tasks that could be integrated into AI-assisted clinical handover.
Methods:
Two registered nurses independently annotated a dataset of 203 synthetic handover transcripts to produce consensus labels for information extraction tasks. Tasks included: 1) labelling spans of text into SBAR (Situation, Background, Assessment, Recommendation) categories; 2) content detection to determine if specific pieces of information were communicated; and 3) labelling spans of text where clinicians expressed uncertainty, including a sub-task for identifying that a relevant fact was unknown, unavailable, or not yet confirmed. Accuracy, precision, recall and F1 scores for baseline and Genetic-Pareto (GEPA) optimised prompts were compared for GPT-5.2, GPT-5-nano, and MedGemma 27B large language models. GEPA iteratively evaluates candidate prompts against task performance and uses model-generated feedback to revise them. Additionally, the LangExtract structured information extraction large language model framework was evaluated for span-extraction tasks.
Results:
The GPT-5.2 optimised model achieved micro-F1 0.85 (95% CI 0.83-0.88) for content detection, an absolute improvement of +0.08 compared with the matched baseline. GPT-5-nano also performed strongly after optimisation for content detection (micro-F1 0.81 (95% CI 0.78-0.84)), suggesting that this structured task was not limited to the highest-capacity model. For SBAR span extraction, GPT-5.2 with prompt optimisation achieved micro-F1 0.76 (95% CI 0.72-0.79), improving by +0.24 compared with baseline and exceeding LangExtract; GPT-5-nano also improved to micro-F1 0.69 (95% CI 0.66-0.72). Broad uncertainty-span extraction remained comparatively weak despite prompt optimisation (micro-F1 0.41 (95% CI 0.33-0.48); absolute improvement +0.06). In contrast, explicit unknown-fact extraction was stronger with GPT-5.2 (micro-F1 0.84 (95% CI 0.63-1.00)), GPT-5-nano (micro-F1 0.84 (95% CI 0.63-1.00)), and MedGemma 27B (micro-F1 0.80 (95% CI 0.63-1.00)). Genetic-Pareto optimised prompts outperformed the LangExtract approach across each span-extraction task.
Conclusions:
Large language models demonstrated sufficient performance to support selected clinical handover tasks, particularly verification of key information and extraction of structured SBAR content. These capabilities could help reduce information omissions and improve the consistency of handover communication when used as clinician-reviewed prompts rather than autonomous summaries. Although broad uncertainty detection remained insufficiently reliable, models performed substantially better at identifying explicit gaps in knowledge. The most useful near-term function may therefore be identifying what has not been said, rather than restructuring what has, by prompting clarification before communication failures contribute to patient harm.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.