Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Jan 11, 2026
Date Accepted: Jun 8, 2026
Quality Evaluation of Large Language Model-Assisted Generation of Initial Senior Physician Ward Round Records for Acute Poisoning Patients: A Cross-Sectional Study
ABSTRACT
Background:
Large language models (LLMs) have demonstrated potential in certain medical text generation tasks. Senior physician ward round records are critical medical documents in the diagnostic and treatment process, with their quality reflecting the accuracy and continuity of clinical decision-making. The quality of senior physician rounds notes generated by large language models (LLMs) in the specific context of acute poisoning remains unclear.
Objective:
This study focused on acute poisoning patients and systematically compared the differences in medical record writing quality among the domestic model DeepSeek, the international model ChatGPT, and human physicians across multiple dimensions to clarify their clinical application value.
Methods:
A retrospective analysis was conducted on 256 poisoning cases from the emergency department ward of Taihe Hospital, Hubei University of Medicine. The study utilized DeepSeek V3.2 Exp and ChatGPT-5.1 to generate senior physician ward round records based on standardized prompts, which were then compared with actual records from the original medical charts. Three senior emergency physicians conducted blinded evaluations, scoring overall quality across five dimensions using a Likert 1-5 scale: case characteristics, current diagnosis, differential diagnosis, treatment plan, and prognosis assessment. Error frequencies in medical records were also documented. Additionally, potential harm was assessed using a modified AHRQ harm scale.
Results:
The DeepSeek model achieved the highest mean total score (24.14 ± 0.90), significantly higher than ChatGPT (23.30±1.42, P< 0.001) and the physician group (23.86±0.86, P=0.015). DeepSeek demonstrated optimal performance in both differential diagnosis (4.98±0.10) and prognosis assessment (4.73±0.42), while matching the physician group's performance (4.96± 0.15) in case characteristics (4.90±0.23) (P>0.05). Subgroup analysis showed that for drug poisoning and pesticide poisoning, DeepSeek's mean total scores (24.23±0.75, 23.92 ±1.14) were significantly higher than ChatGPT (23.34±1.33, 22.78 ± 1.33; both P<0.001). In biological toxin poisoning, DeepSeek (23.97±0.96) and the physician group (24.26±0.62) had similar scores, both significantly higher than ChatGPT (22.53±1.86, P<0.001). All between-group differences were statistically significant (P<0.05). The overall potential harm scores for errors were low across all three groups (<1 point), with no statistically significant differences (P = 0.38).
Conclusions:
LLMs demonstrated satisfactory performance in generating initial senior physician ward round records for acute poisoning cases, particularly outperforming the physician group in case characteristics integration and treatment plan formulation, showing potential for assisting clinical documentation.LLMs can serve as effective tools to help physicians reduce documentation workload, but physician review and oversight remain essential. Clinical Trial: none
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.