Currently submitted to: Journal of Medical Internet Research
Date Submitted: Sep 30, 2026
Open Peer Review Period: Sep 30, 2026 - Nov 25, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Neuro-Symbolic Knowledge Graphs Improve Motivational Interviewing Fidelity in Clinical Artificial Intelligence Tools: Development and Simulation Study
ABSTRACT
Background:
Large language models (LLMs) are increasingly used in health counseling, but adherence to evidence-based approaches such as motivational interviewing (MI) is inconsistent, and their reasoning is difficult to inspect or audit. Supervised fine-tuning (SFT) trains a model to imitate expert responses but does not explicitly encode the sequential clinical decisions MI requires. Knowledge graphs (KGs), which represent clinical knowledge as explicit, inspectable rules and relationships, may offer a more transparent way to govern LLM counseling.
Objective:
This study aimed to develop a neuro-symbolic KG encoding MI clinical logic and to evaluate, in simulated sessions, whether KG guidance improves MI fidelity compared with a base LLM and SFT, alone and in combination.
Methods:
We built a 3-layer KG comprising an MI ontology (34 techniques, stages of change, and patient talk types), a deterministic rule engine (34 decision rules governing technique selection and phase progression), and 13 hard safety constraints, and integrated it into CHIA (Chatbot for HIV Prevention and Action), a preexposure prophylaxis (PrEP) counseling platform. Using a 2×2 factorial design crossing SFT (GPT-4.1 base vs fine-tuned on 2400 expert MI examples) with KG guidance (disabled vs enabled), we simulated 8-turn sessions with 20 LLM-generated patient personas (3 runs per persona per arm; 239/240 sessions completed). A blinded LLM judge (Claude Sonnet 4.5) scored Motivational Interviewing Treatment Integrity (MITI) 4.2 global scores (Empathy, Partnership, Cultivating Change Talk, and Softening Sustain Talk) and flagged premature planning. We used 2-way repeated-measures ANOVA and generalized estimating equations with false discovery rate correction, and computed linguistic metrics including lexical diversity. A blinded MI expert and a second LLM judge (Gemini 2.5 Pro) rescored a stratified subsample of 40 sessions.
Results:
The base LLM with KG guidance (Base+KG) achieved the highest mean Empathy (4.90, SD 0.30) and Partnership (4.88, SD 0.32). Compared with SFT alone, Base+KG showed higher Empathy (Cohen d=0.68, 95% CI 0.37-1.01) and Partnership (d=0.91, 95% CI 0.60-1.25). SFT had significant negative main effects on Empathy (F1,19=21.26; P<.001) and Partnership (F1,19=17.39; P<.001) and reduced lexical diversity (d=−1.47). The LLM judge flagged premature planning in 0/60 Base+KG sessions vs 1/60 Base, 4/59 SFT, and 6/60 SFT+KG sessions; however, the human expert identified more premature planning in every arm. Human-LLM agreement was moderate for Empathy (intraclass correlation coefficient [ICC] 0.55) and Partnership (ICC 0.42) and poor for Softening Sustain Talk, but all 3 raters produced consistent arm rankings.
Conclusions:
In this simulation study, KG guidance applied to a base LLM yielded higher MI fidelity than SFT, which degraded relational quality. Externalizing clinical reasoning into an inspectable KG is a promising, auditable approach to governing LLM counseling, pending confirmation with human participants and larger human-rated samples.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.