Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Jan 28, 2026
Date Accepted: Jun 30, 2026
Finite State Machine-Guided RAG Improves Expert-Rated Acceptability of a PICC Self-Management Chatbot: A Single-Center Content Validation Study
ABSTRACT
Background:
Patients with cancer undergoing long-term or vesicant chemotherapy frequently require peripherally inserted central catheters (PICCs). Due to the nature of ambulatory treatment administration, self-PICC management is essential for the continuation and completion of the planned treatment. Although large language models (LLMs) offer potential for continuous patient support, hallucinations and insufficient adherence to clinical protocols pose substantial safety concerns that impede clinical adoption. While fine-tuning and retrieval-augmented generation (RAG) enhance factual grounding, these approaches alone cannot enforce the deterministic decision logic essential for structured clinical consultations.
Objective:
This study aimed to evaluate whether a finite state machine (FSM)-guided RAG architecture improves the clinical feasibility of PICC self-management consultations compared with simpler LLM architectures (fine-tuned model alone and fine-tuned model with RAG).
Methods:
We conducted a blinded comparative evaluation of 3 chatbot architectures using 43 standardized PICC-related clinical scenarios. Three oncology-specialized nurses independently evaluated responses between August and September 2025. Scenarios encompassed catheter care, daily activities, emergency situations, heparin flushing, insertion site abnormalities, and symptom management. The architectures tested were Model 1 (fine-tuned GPT-4o-mini), Model 2 (fine-tuned model with RAG), and Model 3 (fine-tuned model with RAG and FSM-based dialogue control). The primary outcome was expert preference among the 3 models. Secondary outcomes comprised Likert scale ratings (1-5) across 8 evaluation domains: accuracy, clarity, actionability, completeness, efficiency, adaptability, safety, and overall quality.
Results:
The FSM-guided model (Model 3) was preferred in 34 of 43 scenarios (79.1%), demonstrating category-specific superiority in catheter care (80%), insertion site abnormalities (88%), daily activities (82%), symptoms (75%), heparin flushing (67%), and emergency situations (100%). In overall evaluation, Model 3 achieved superior mean ratings in practicality (4.67) and adaptability (4.67), demonstrating balanced performance across evaluation criteria. Conversely, Model 3 scored lower in efficiency (3.33) compared with Model 1 (4.00). Model 2 exhibited the lowest overall performance, with notably reduced scores in accuracy (1.67), completeness (1.67), and overall quality (1.67). In qualitative debriefing interviews, the nurses mentioned that the FSM guidance occasionally generated unnecessary conversational turns. However, most of them agreed that this trade-off significantly enhanced protocol adherence and mitigated safety risks.
Conclusions:
In this comparative evaluation of PICC self-management chatbot architectures, the FSM-guided RAG model demonstrated superior clinical feasibility by integrating deterministic protocol enforcement with linguistic adaptability. These findings suggest that incorporating structured state control mechanisms into LLM-based healthcare applications may be critical for ensuring procedural reliability and patient safety in clinical decision support systems.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.