Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Formative Research

Date Submitted: Aug 29, 2026
Open Peer Review Period: Aug 30, 2026 - Oct 25, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Feasibility of a Pathway-Aware, Physician-Rated Framework for Evaluating Large Language Models in Urolithiasis Management

  • YiYang Lin; 
  • Yi Shao; 
  • Weiming Mou; 
  • Ziyang Zheng; 
  • Ruixuan Zhu; 
  • Yan Zhuang; 
  • Yang Tao; 
  • Jiaxu Song; 
  • Deng Li

ABSTRACT

Background:

Formative evaluations of digital health technologies remain underreported, yet they determine whether an evaluation instrument and a decision-support tool are ready for prospective deployment.

Objective:

Most medical LLM evaluations rely on isolated tasks and fail to capture longitudinal specialty decision-making. We developed and evaluated a urolithiasis-specific, pathway-aware, physician-rated framework covering diagnosis, treatment planning, procedural decision-making, and postoperative management, and assessed its feasibility and inter-rater reliability. We applied the framework to six contemporary LLMs as a demonstration of stage-dependent performance, error patterns, and deployment trade-offs.

Methods:

Six LLMs—GPT-5.2, GPT-5.2 Thinking, GPT-5 Mini, Gemini 3 Pro, DeepSeek R1, and Grok 4.1—answered 55 structured items for each of 100 real-world urolithiasis cases from one academic center. Two senior urologists independently rated all 33,000 responses (66,000 ratings) against the clinical record using a 5-point Likert scale. We analyzed inter-rater reliability, between-model differences, error taxonomy, response latency, and estimated cost.

Results:

Inter-rater reliability was high (Gwet’s AC1=0.8655; ICC=0.965). Model performance differed overall (Kruskal–Wallis H=458.45, P<.001), with GPT-5 Mini highest (mean 4.9528). Reasoning enhancement produced its largest gain in stent-management scenarios (+1.97% for GPT-5.2 Thinking vs GPT-5.2). Information omission was the predominant failure mode (61.8% of low-scoring responses), concentrated in treatment planning and postoperative management. Latency and cost estimates demonstrated deployment-oriented comparability and task-stratified trade-offs.

Conclusions:

This pathway-aware, physician-rated framework is feasible for pre-deployment formative evaluation, with acceptable inter-rater reliability and deployment-relevant comparators. The results suggest task-stratified model selection rather than a single best model. Prospective, multi-institutional validation linking framework scores to patient outcomes is the next required step before clinical deployment. Clinical Trial: Not applicable


 Citation

Please cite as:

Lin Y, Shao Y, Mou W, Zheng Z, Zhu R, Zhuang Y, Tao Y, Song J, Li D

Feasibility of a Pathway-Aware, Physician-Rated Framework for Evaluating Large Language Models in Urolithiasis Management

JMIR Preprints. 29/08/2026:110729

DOI: 10.2196/preprints.110729

URL: https://preprints.jmir.org/preprint/110729

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.