Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Medical Education

Date Submitted: Nov 28, 2025
Date Accepted: Apr 21, 2026

The final, peer-reviewed published version of this preprint can be found here:

Benefits of Chain-of-Thought Prompting for Clinical Record Rubric Evaluation in Undergraduate Medical Education: Experimental Evaluation Study With Medical Faculty

Nogales A, Denizon S, Mateos A, Cervera J, Pandelet G, Aranguren E, García-Tejedor AJ, Cervera E

Benefits of Chain-of-Thought Prompting for Clinical Record Rubric Evaluation in Undergraduate Medical Education: Experimental Evaluation Study With Medical Faculty

JMIR Med Educ 2026;12:e88652

DOI: 10.2196/88652

PMID: 42492497

Benefits of Chain of Thought Prompting for Clinical Record Rubric Evaluation in Undergraduate Medicine Education: An Experimental Evaluation Study with Medical Faculty

  • Alberto Nogales; 
  • Sophia Denizon; 
  • Alonso Mateos; 
  • Javier Cervera; 
  • Gonzalo Pandelet; 
  • Enrique Aranguren; 
  • Alvaro J. García-Tejedor; 
  • Emilio Cervera

ABSTRACT

Background:

Large Language Models from the field of Artificial Intelligence have been one of the tools that have a great and real impact on the daily lives of people. In this regard, they emerge as an aid in specific fields such as education helping educators in cumbersome tasks such as periodical evaluations.

Objective:

This paper is focused in analyzing the benefits of Large Language Models and in particular Chain of Thought strategy for the task of evaluating the writing in Spanish of medical records by students. The aim is two-fold: first we try to save time and re-sources and second, we use the reasonings of the Chain of Thought to evaluate the ru-brics and its interpretations.

Methods:

The proposed solution studies the application of two models such as Llama 3.1 and Claude 3.5 in combination with two strategies One-Shot and Chain of Thought to evaluate how medicine students write medical records in Spanish. First, machine learn-ing metrics are applied to measure the performance of the solutions. Then, different statistical analysis is obtained at clinical record and item level. Finally, the differences between proposed models and evaluators are studied in-depth.

Results:

Machine learning metrics indicate that, at first, Claude got an Accuracy over 86% and no substantial differences among them. However, an in-depth analysis high-lights the potential of Chain of Thought both for identifying evaluator errors and for addressing issues in the rubrics. Applying this analysis, we have found 317 human errors which, after a correction, have set the Accuracy in 94.6%.

Conclusions:

Chain of Thought demonstrates strong potential for supporting the evalua-tion of clinical records written in Spanish by medical students and providing feedback to them. More importantly, it shows significant promise in assisting professors by as-sessing the quality of their rubrics and identifying possible errors.


 Citation

Please cite as:

Nogales A, Denizon S, Mateos A, Cervera J, Pandelet G, Aranguren E, García-Tejedor AJ, Cervera E

Benefits of Chain-of-Thought Prompting for Clinical Record Rubric Evaluation in Undergraduate Medical Education: Experimental Evaluation Study With Medical Faculty

JMIR Med Educ 2026;12:e88652

DOI: 10.2196/88652

PMID: 42492497

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.