Currently submitted to: JMIR Formative Research
Date Submitted: Sep 21, 2026
Open Peer Review Period: Sep 21, 2026 - Nov 16, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Evidence-Grounded Simulated Patients for Testing Conversational Health AI: A Behavioral Simulation Framework
ABSTRACT
Background:
Conversational health AI is increasingly being developed for patient-facing applications, but evaluating stochastic, multi-turn systems before human studies remains challenging. Simulated patients could provide repeatable test environments, but their usefulness depends on whether the simulated patient itself reliably preserves the behavioral characteristics defining the test case.
Objective:
We evaluated an evidence-grounded simulated-patient framework for testing conversational health AI, using medication adherence as a behavioral test case. Specifically, we examined whether explicit behavioral representation and structured operationalization improve simulation fidelity beyond the same evidence supplied as prose or conventional patient prompting.
Methods:
Behavioral profiles for South Asian and Chinese simulated patients were derived from independent systematic reviews. Across 80 standardized multi-turn conversations (880 patient responses), we compared structured evidence-grounded simulation with the same evidence supplied as prose, ad hoc prompting, and web-informed patient construction. Automated measures assessed behavioral grounding, recovery of assigned behaviors, fidelity to assigned behavioral characteristics, unsupported barriers, conversation behavior, and correspondence with an independent corpus of published patient utterances.
Results:
Structured operationalization improved fidelity of generated behavior to the intended evidence-derived patient state. In the primary paired ablation, structured operationalization increased mean conversation-level behavioral groundedness compared with the same evidence supplied as prose (0.850 vs 0.732; mean paired difference 0.117, 95% CI 0.077-0.157; Cohen dz=1.376; P<.001). Structure did not improve every measure: prose produced better recovery of assigned COM-B dimensions (F1-score: 0.641 vs 0.503; P<.001), and other outcomes varied by evidence source.
Conclusions:
Evidence-grounded simulated patients provide a computationally evaluable approach for creating repeatable behavioral test cases for conversational health AI. Explicit structure improved behavioral grounding beyond evidence content alone but did not uniformly improve all fidelity measures. This computational stage can support pre-human testing while identifying where subsequent expert and patient validation remains necessary.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.