Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Mar 11, 2026
Date Accepted: Aug 18, 2026
Evaluating a Guideline-Integrated Clinical Interaction Framework Versus a Standard Large Language Model Interaction for Dietary Recommendations in Recurrent Urolithiasis: An In Silico Study
ABSTRACT
Background:
Personalized dietary counseling is central to recurrence prevention in patients with urolithiasis, particularly after 24-hour urine metabolic evaluation. However, translating quantitative metabolic abnormalities into patient-facing, guideline-concordant, and safe dietary recommendations can be challenging in routine clinical practice. Large language models (LLMs) may assist with this task, but unguided responses may overlook key metabolic priorities or case-specific safety constraints.
Objective:
This study evaluated whether a guideline-integrated, safety-aware clinical interaction framework (StoneAgent) could generate higher-quality personalized dietary recommendations than a standard LLM configuration for recurrent urolithiasis. We also assessed whether any performance advantage was retained when the same clinical scenarios were presented as patient query-style inputs.
Methods:
We conducted an in silico comparative study using 30 synthetic clinical vignettes representing common, mixed, and safety-relevant metabolic stone scenarios. For the primary experiment, StoneAgent and a standard LLM configuration were compared using structured vignette inputs. For the robustness experiment, each vignette was reformulated into 2 patient query-style variants (Query A and Query B), which preserved the same clinical content in more natural conversational language. A vignette-specific expert reference standard was developed from guideline-informed specialist consensus. Three independent reviewers rated blindly outputs on a 5-point Likert scale for metabolic specificity, guideline adherence, and actionability, and safety was assessed as a binary outcome. For the patient query-style experiment, Query A and Query B were aggregated to the vignette level for paired comparison.
Results:
In the structured-input experiment, StoneAgent achieved higher performance than the standard LLM across metabolic specificity, guideline adherence, and actionability, with median case-level scores of 5.00 (IQR 5.00-5.00) versus 3.00 (IQR 2.75-3.92) for metabolic specificity, 5.00 (IQR 5.00-5.00) versus 3.67 (IQR 3.08-4.00) for guideline adherence, and 5.00 (IQR 5.00-5.00) versus 3.00 (IQR 2.67-3.33) for actionability (all P<.001). Safety pass rates were 30/30 for StoneAgent and 25/30 for the standard LLM (exact McNemar P=.063). In the patient query-style robustness experiment, StoneAgent retained a directional advantage after case-level aggregation, with mean scores of 4.44 versus 3.62 for metabolic specificity, 4.51 versus 3.63 for guideline adherence, and 4.11 versus 3.40 for actionability. Safety pass rates were 30/30 for StoneAgent and 27/30 for the standard LLM (exact McNemar P=.25). The performance gap was more conservative under patient query-style inputs than under structured inputs, but the overall pattern remained consistent across domains.
Conclusions:
In this in silico study, a guideline-integrated, safety-aware clinical interaction framework generated higher quality dietary recommendations for recurrent urolithiasis than a standard general-purpose LLM configuration under structured vignette inputs, and this advantage was retained under patient query-style inputs. These findings suggest that explicit clinical framing, guideline grounding, and safety-oriented response scaffolding may improve the reliability of specialty counseling tasks involving metabolic stone prevention. Further validation is needed using real patient-authored queries and prospective clinical workflows.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.