Currently submitted to: JMIR AI
Date Submitted: Sep 26, 2026
Open Peer Review Period: Oct 5, 2026 - Nov 30, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Benign Symptom Framing and Chatbot Escalation Advice After Smoking Cessation: A Paired Turkish-Language Vignette Study
ABSTRACT
Background:
People seeking advice after quitting smoking may suggest that respiratory or chest symptoms reflect normal recovery. Consumer chatbots need to distinguish that explanation from the action required by the symptoms described. Reassuring language and safety-net advice do not necessarily establish an appropriate recommendation for current symptoms.
Objective:
To examine whether benign explanatory framing changes the frequency of insufficient escalation in Turkish-language chatbot advice after smoking cessation, and to describe differences between products and repeated responses.
Methods:
CLEAR-QUIT evaluated 24 synthetic scenarios, comprising eight emergency, eight urgent, and eight routine/self-care scenarios. ChatGPT, Gemini, Claude, Grok, and DeepSeek received neutral and benign-framed versions in two administrations, producing 480 responses. AI-assisted outcome codes were reviewed by an expert in smoking cessation using blinded materials. The primary comparison used 16 first-administration emergency/urgent scenarios, with 80 responses per framing. Insufficient escalation combined advice below the reference urgency, unclear urgency, or conflicting delay. We estimated equally weighted scenario-level risk differences with scenario-block bootstrap 95% confidence intervals and exact sign-flip tests. Post hoc product comparisons specified in a separate plan used Holm correction.
Results:
Insufficient escalation occurred in 29/80 neutral responses (36.3%) and 26/80 benign-framed responses (32.5%): risk difference −3.75 percentage points (95% CI −13.75 to 6.25; P=.637). Emergency-scenario failures were 0/40 versus 4/40; urgent-scenario failures were 29/40 versus 22/40. These strata were exploratory. Accepting broader urgent-timing language changed the pooled difference to 2.50 points (95% CI −6.25 to 11.25). Observed product risks ranged from 7/32 to 15/32, but no pairwise comparison met the Holm-adjusted threshold. Across all responses, 421/480 (87.7%) included safety-netting and none advised restarting smoking. Repeat-administration triage agreement was 210/240 (87.5%).
Conclusions:
This benchmark did not establish an overall increase in insufficient escalation under benign framing; its confidence interval allowed effects in either direction. Urgent-advice classification depended on the timing rule. Chatbot health-information evaluations should assess the recommended action and timeframe alongside supportive content. These results do not establish product equivalence or safety in clinical use.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.