Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Medical Education

Date Submitted: Jul 27, 2026
Open Peer Review Period: Jul 29, 2026 - Sep 23, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Predicting Diagnostic Accuracy from Student Clinical Simulation Transcripts Using Term Frequency Inverse Document Frequency Analysis

  • Christine Yang Zhou; 
  • James Bowen; 
  • Sandy Chan; 
  • Matthew Kelleher; 
  • Danielle E Weber; 
  • Seth Overla; 
  • Sally Santen; 
  • Laurah Turner

ABSTRACT

Background:

Assessment of clinical reasoning (CR) traditionally relies on resource-intensive methods that indirectly infer reasoning quality and evaluate only a limited number of encounters. Artificial intelligence (AI)-based simulation platforms generate analyzable transcripts of learner clinical encounters, offering a potential scalable complement to existing CR assessment.

Objective:

This study examined whether linguistic features within student-generated transcripts from an AI-based clinical simulation platform (2-Sigma) are associated with diagnostic accuracy, and which portion of the encounter (history-taking, physical examination, diagnostic testing, or intervention) contributes most to that association.

Methods:

This was an exploratory, retrospective study of 179 second-year medical students completing simulated patient encounters across 9 cases (1,515 student-case sessions). Student messages were pre-classified into action categories (history-taking, physical examination, diagnostic testing, intervention, other); "other" messages, including final diagnosis submissions, were excluded to avoid data leakage. Text was converted to term frequency–inverse document frequency (TF-IDF) features and used to train L2-regularized logistic regression models with 5-fold stratified cross-validation. Eight pooled (cross-case) and 13 per-case model variants (including word-count-only and hybrid models) were compared using AUROC, AUPRC, Brier score, and Youden's J-optimized sensitivity/specificity. Singular value decomposition tested whether performance depended on broad language patterns versus rare vocabulary.

Results:

All 8 pooled models discriminated diagnostic accuracy above chance (p<0.0005). The full four-category model performed best (AUROC 0.766, 95% CI 0.743-0.788); among single categories, Diagnostics Only (AUROC 0.745) outperformed History Only (AUROC 0.710) and Exam Only (AUROC 0.620). Adding Diagnostics text to History improved discrimination (ΔAUROC +0.029, p<0.001), but the reverse did not (ΔAUROC −0.009, p=0.17). Of 114 per-case models, 89 (78.1%) achieved significant discrimination, with substantial case-level heterogeneity (AUROC range 0.622-0.887) depending on whether a case had a defining diagnostic test. Dimensionality reduction to 50-200 components preserved performance within 0.044 AUROC of the full model.

Conclusions:

TF-IDF analysis of simulated encounter transcripts showed moderate, statistically significant discrimination between diagnostically successful and unsuccessful encounters, with diagnostic-testing language providing the strongest and most asymmetric signal relative to history-taking. The asymmetry suggests a potential detectable inflection point between information-gathering and hypothesis-testing phases of reasoning, warranting validation across training levels and clinical settings.


 Citation

Please cite as:

Zhou CY, Bowen J, Chan S, Kelleher M, Weber DE, Overla S, Santen S, Turner L

Predicting Diagnostic Accuracy from Student Clinical Simulation Transcripts Using Term Frequency Inverse Document Frequency Analysis

JMIR Preprints. 27/07/2026:108078

DOI: 10.2196/preprints.108078

URL: https://preprints.jmir.org/preprint/108078

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.