Currently submitted to: JMIR Formative Research
Date Submitted: Aug 29, 2026
Open Peer Review Period: Aug 30, 2026 - Oct 25, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Feasibility of a Pathway-Aware, Physician-Rated Framework for Evaluating Large Language Models in Urolithiasis Management
ABSTRACT
Background:
Formative evaluations of digital health technologies remain underreported, yet they determine whether an evaluation instrument and a decision-support tool are ready for prospective deployment.
Objective:
Most medical LLM evaluations rely on isolated tasks and fail to capture longitudinal specialty decision-making. We developed and evaluated a urolithiasis-specific, pathway-aware, physician-rated framework covering diagnosis, treatment planning, procedural decision-making, and postoperative management, and assessed its feasibility and inter-rater reliability. We applied the framework to six contemporary LLMs as a demonstration of stage-dependent performance, error patterns, and deployment trade-offs.
Methods:
Six LLMs—GPT-5.2, GPT-5.2 Thinking, GPT-5 Mini, Gemini 3 Pro, DeepSeek R1, and Grok 4.1—answered 55 structured items for each of 100 real-world urolithiasis cases from one academic center. Two senior urologists independently rated all 33,000 responses (66,000 ratings) against the clinical record using a 5-point Likert scale. We analyzed inter-rater reliability, between-model differences, error taxonomy, response latency, and estimated cost.
Results:
Inter-rater reliability was high (Gwet’s AC1=0.8655; ICC=0.965). Model performance differed overall (Kruskal–Wallis H=458.45, P<.001), with GPT-5 Mini highest (mean 4.9528). Reasoning enhancement produced its largest gain in stent-management scenarios (+1.97% for GPT-5.2 Thinking vs GPT-5.2). Information omission was the predominant failure mode (61.8% of low-scoring responses), concentrated in treatment planning and postoperative management. Latency and cost estimates demonstrated deployment-oriented comparability and task-stratified trade-offs.
Conclusions:
This pathway-aware, physician-rated framework is feasible for pre-deployment formative evaluation, with acceptable inter-rater reliability and deployment-relevant comparators. The results suggest task-stratified model selection rather than a single best model. Prospective, multi-institutional validation linking framework scores to patient outcomes is the next required step before clinical deployment. Clinical Trial: Not applicable
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.