Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Medical Education

Date Submitted: Aug 20, 2026
Open Peer Review Period: Aug 20, 2026 - Oct 15, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Process-Level Assessment of Diagnostic Reasoning With a Clinical Simulation Video and Free-Text Responses: Nationwide Cross-Sectional Study of Resident Physicians

  • Kiyoshi Shikino; 
  • Yuji Nishizaki; 
  • Koshi Kataoka; 
  • Sho Fukui; 
  • Kentaro Sakamaki; 
  • Yu Yamamoto; 
  • Taro Shimizu; 
  • So Sakamoto; 
  • Ryo Morishima; 
  • Tadamasa Wakabayashi; 
  • Hiroyuki Kobayashi; 
  • Yasuharu Tokuda

ABSTRACT

Background:

Multiple-choice assessments may not fully capture how learners generate diagnostic hypotheses, justify them using multimodal clinical information, and translate them into initial management. Clinical simulation video (CSV) items with free-text responses may provide process-level information, but coding such responses at a national scale is resource-intensive.

Objective:

The study examines whether the following three process markers derived from a CSV free-text item are associated with General Medicine In-Training Examination (GM-ITE) total scores: diagnostic hypothesis generation, diagnostic justification, and initial management.

Methods:

This cross-sectional study used data from the 2023 GM-ITE. A patient-reenactment CSV item depicting hypoglycemia as a stroke mimic was administered as an add-on and excluded from the official score. Eligible participants were examinees who completed the item and consented to its use in research. Free-text responses were normalized using a constrained generative AI-assisted workflow and independently re-coded by at least two authors using an expert-informed codebook. Associations between the three process markers and GM-ITE total scores were evaluated. Joint patterns of hypoglycemia listing and hypoglycemia-directed management were examined using four mutually exclusive groups, and an exploratory diagnosis-by-management interaction was tested.

Results:

A total of 6,584 participants were included in the analytic cohort. Hypoglycemia was listed first by 248 (3.8%) participants and within the top three by 1,363 (20.7%). Unadjusted GM-ITE total scores were 48.3, 46.8, and 44.4 among participants who listed hypoglycemia first, second or third, and not within the top three, respectively (P<.001). Among participants who listed hypoglycemia, scores were 43.9, 46.2, 46.1, and 47.6 across clue-category counts of 0, 1, 2, and 3 or more, respectively (P=.001). Hypoglycemia-directed management was associated with a higher score among participants who listed hypoglycemia (47.2 vs 45.5; mean difference=1.68, 95% CI 0.47-2.88), but with a lower score among those who did not list hypoglycemia (42.0 vs 44.5; mean difference=-2.48, 95% CI -3.41 to -1.56; interaction P<.001). In adjusted sensitivity analyses, listing hypoglycemia first (adjusted coefficient=3.19, 95% CI 2.27-4.11), each additional clue category (adjusted coefficient=0.82, 95% CI 0.50-1.14), and hypoglycemia-directed management (adjusted coefficient=1.43, 95% CI 1.02-1.85) remained associated with higher GM-ITE total scores. The diagnosis-by-management interaction remained statistically significant after adjustment (adjusted interaction coefficient=2.90, 95% CI 1.53-4.26; P<.001).

Conclusions:

Process markers derived from a CSV free-text response were associated with an established measure of clinical competence. The alignment between an explicitly generated diagnostic hypothesis and the selected management action may provide additional information beyond either component alone. These findings suggest the feasibility of using human-verified, AI-assisted normalization to analyze free-text responses in large cohorts.


 Citation

Please cite as:

Shikino K, Nishizaki Y, Kataoka K, Fukui S, Sakamaki K, Yamamoto Y, Shimizu T, Sakamoto S, Morishima R, Wakabayashi T, Kobayashi H, Tokuda Y

Process-Level Assessment of Diagnostic Reasoning With a Clinical Simulation Video and Free-Text Responses: Nationwide Cross-Sectional Study of Resident Physicians

JMIR Preprints. 20/08/2026:110021

DOI: 10.2196/preprints.110021

URL: https://preprints.jmir.org/preprint/110021

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.