Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Mar 11, 2026
Date Accepted: Aug 18, 2026

The final, peer-reviewed published version of this preprint can be found here:

Evaluating a Guideline-Integrated Clinical Interaction Framework Vs a Standard Large Language Model Interaction for Dietary Recommendations in Recurrent Urolithiasis: In Silico Study

Wang X, Li J, Hu Y, Chen Y, Zhong Y, Zhu F, Yuan Y, Yang F, Ye J

Evaluating a Guideline-Integrated Clinical Interaction Framework Vs a Standard Large Language Model Interaction for Dietary Recommendations in Recurrent Urolithiasis: In Silico Study

J Med Internet Res 2026;28:e95162

DOI: 10.2196/95162

PMID: 42696728

Evaluating a Guideline-Integrated Clinical Interaction Framework Versus a Standard Large Language Model Interaction for Dietary Recommendations in Recurrent Urolithiasis: An In Silico Study

  • Xiaofeng Wang; 
  • Jun Li; 
  • Yudong Hu; 
  • Yujie Chen; 
  • Yong Zhong; 
  • Faming Zhu; 
  • Ye Yuan; 
  • Fan Yang; 
  • Jin Ye

ABSTRACT

Background:

Personalized dietary counseling is central to recurrence prevention in patients with urolithiasis, particularly after 24-hour urine metabolic evaluation. However, translating quantitative metabolic abnormalities into patient-facing, guideline-concordant, and safe dietary recommendations can be challenging in routine clinical practice. Large language models (LLMs) may assist with this task, but unguided responses may overlook key metabolic priorities or case-specific safety constraints.

Objective:

This study evaluated whether a guideline-integrated, safety-aware clinical interaction framework (StoneAgent) could generate higher-quality personalized dietary recommendations than a standard LLM configuration for recurrent urolithiasis. We also assessed whether any performance advantage was retained when the same clinical scenarios were presented as patient query-style inputs.

Methods:

We conducted an in silico comparative study using 30 synthetic clinical vignettes representing common, mixed, and safety-relevant metabolic stone scenarios. For the primary experiment, StoneAgent and a standard LLM configuration were compared using structured vignette inputs. For the robustness experiment, each vignette was reformulated into 2 patient query-style variants (Query A and Query B), which preserved the same clinical content in more natural conversational language. A vignette-specific expert reference standard was developed from guideline-informed specialist consensus. Three independent reviewers rated blindly outputs on a 5-point Likert scale for metabolic specificity, guideline adherence, and actionability, and safety was assessed as a binary outcome. For the patient query-style experiment, Query A and Query B were aggregated to the vignette level for paired comparison.

Results:

In the structured-input experiment, StoneAgent achieved higher performance than the standard LLM across metabolic specificity, guideline adherence, and actionability, with median case-level scores of 5.00 (IQR 5.00-5.00) versus 3.00 (IQR 2.75-3.92) for metabolic specificity, 5.00 (IQR 5.00-5.00) versus 3.67 (IQR 3.08-4.00) for guideline adherence, and 5.00 (IQR 5.00-5.00) versus 3.00 (IQR 2.67-3.33) for actionability (all P<.001). Safety pass rates were 30/30 for StoneAgent and 25/30 for the standard LLM (exact McNemar P=.063). In the patient query-style robustness experiment, StoneAgent retained a directional advantage after case-level aggregation, with mean scores of 4.44 versus 3.62 for metabolic specificity, 4.51 versus 3.63 for guideline adherence, and 4.11 versus 3.40 for actionability. Safety pass rates were 30/30 for StoneAgent and 27/30 for the standard LLM (exact McNemar P=.25). The performance gap was more conservative under patient query-style inputs than under structured inputs, but the overall pattern remained consistent across domains.

Conclusions:

In this in silico study, a guideline-integrated, safety-aware clinical interaction framework generated higher quality dietary recommendations for recurrent urolithiasis than a standard general-purpose LLM configuration under structured vignette inputs, and this advantage was retained under patient query-style inputs. These findings suggest that explicit clinical framing, guideline grounding, and safety-oriented response scaffolding may improve the reliability of specialty counseling tasks involving metabolic stone prevention. Further validation is needed using real patient-authored queries and prospective clinical workflows.


 Citation

Please cite as:

Wang X, Li J, Hu Y, Chen Y, Zhong Y, Zhu F, Yuan Y, Yang F, Ye J

Evaluating a Guideline-Integrated Clinical Interaction Framework Vs a Standard Large Language Model Interaction for Dietary Recommendations in Recurrent Urolithiasis: In Silico Study

J Med Internet Res 2026;28:e95162

DOI: 10.2196/95162

PMID: 42696728

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.