Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Feb 25, 2026
Date Accepted: Jul 21, 2026
Effects of a Safety UI Bundle on Verification Intentions in Generative AI Chat Use Among Older Chinese Adults: A Randomized Vignette Survey
ABSTRACT
Background:
Generative AI chat systems are increasingly used for everyday information seeking, but plausible errors and omissions can mislead users when outputs are accepted without scrutiny. Interface-level safety cues may help users calibrate trust and engage in verification, yet evidence in older Chinese adults remains limited.
Objective:
To test whether adding a safety UI bundle to a generative AI chat interface increases verification intention among older Chinese adults and to examine selected secondary outcomes, including reliance intention, trust calibration, perceived trustworthiness, comprehension, usability/readability, cognitive load, and a behavioral proxy of verification.
Methods:
We conducted a cross-sectional survey with an embedded randomized UI vignette experiment between May 22, 2025 and September 3, 2025. Chinese adults aged ≥60 years were recruited through community sites, outpatient clinic waiting areas, and WeChat groups. Participants were randomized 1:1 to view screenshots of either a baseline chat UI or a safety UI bundle that included generic source-label cues and an uncertainty plus verification nudge. Each participant completed two scenarios (service/travel decision and general wellbeing related to sleep/fatigue), followed by measures of verification intention (primary), reliance intention, trust calibration index, comprehension (0–8), perceived trustworthiness, usability/readability, cognitive load (0–10), manipulation checks, and a behavioral proxy (expanding optional “source information”). Analyses used intention-to-treat regression models with covariate adjustment.
Results:
Of 214 consenting respondents who started the survey, 200 were included in analysis (100 per arm). The safety UI bundle increased verification intention compared with baseline (mean 4.72 vs 4.41 on a 7-point scale; adjusted β=0.293, 95% CI 0.128 to 0.457; P<.001). Reliance intention did not increase (mean 4.97 vs 5.03; adjusted β=-0.105, 95% CI -0.239 to 0.029; P=.13). Trust calibration improved in the safety UI arm (trust calibration index mean -0.29 vs 0.29; adjusted β=-0.567, 95% CI -1.005 to -0.129; P=.01). Expansion of optional source information was numerically higher in the safety UI arm, although the adjusted confidence interval included the null (42% vs 27%; adjusted OR=1.76, 95% CI 0.95 to 3.27; P=.07). Comprehension remained high and similar across arms (mean 6.33 vs 6.32; adjusted β=-0.132, 95% CI -0.428 to 0.163; P=.38). Perceived trustworthiness was modestly lower in the safety UI arm (mean 5.20 vs 5.39; adjusted β=-0.199, 95% CI -0.382 to -0.016; P=.03). Usability/readability was unchanged, and cognitive load did not increase. Manipulation checks indicated higher cue recognition in the safety UI arm.
Conclusions:
In a randomized static-vignette survey of older Chinese adults, a brief safety UI bundle was associated with higher verification intention and a trust-calibration index consistent with lower over-reliance risk, without detectable reductions in comprehension or usability/readability. Because the intervention was tested as a bundle using screenshots and generic source labels, findings should be interpreted as evidence for a practical interface-level strategy rather than proof that any single cue caused the observed effects.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.