Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Feb 25, 2026
Date Accepted: Jul 21, 2026

The final, peer-reviewed published version of this preprint can be found here:

Effects of a Safety User Interface Bundle on Verification Intentions in Generative AI Chat Use Among Older Chinese Adults: Randomized Vignette Survey

Yu J, Chen J, Ren A, Duan H, Meng H, Gao Z

Effects of a Safety User Interface Bundle on Verification Intentions in Generative AI Chat Use Among Older Chinese Adults: Randomized Vignette Survey

J Med Internet Res 2026;28:e94140

DOI: 10.2196/94140

PMID: 42600074

PMCID: 13475784

Effects of a Safety UI Bundle on Verification Intentions in Generative AI Chat Use Among Older Chinese Adults: A Randomized Vignette Survey

  • Jun'an Yu; 
  • Jun Chen; 
  • Anjie Ren; 
  • Hui Duan; 
  • Hua Meng; 
  • Zhuo Gao

ABSTRACT

Background:

Generative AI chat systems are increasingly used for everyday information seeking, but plausible errors and omissions can mislead users when outputs are accepted without scrutiny. Interface-level safety cues may help users calibrate trust and engage in verification, yet evidence in older Chinese adults remains limited.

Objective:

To test whether adding a safety UI bundle to a generative AI chat interface increases verification intention among older Chinese adults and to examine selected secondary outcomes, including reliance intention, trust calibration, perceived trustworthiness, comprehension, usability/readability, cognitive load, and a behavioral proxy of verification.

Methods:

We conducted a cross-sectional survey with an embedded randomized UI vignette experiment between May 22, 2025 and September 3, 2025. Chinese adults aged ≥60 years were recruited through community sites, outpatient clinic waiting areas, and WeChat groups. Participants were randomized 1:1 to view screenshots of either a baseline chat UI or a safety UI bundle that included generic source-label cues and an uncertainty plus verification nudge. Each participant completed two scenarios (service/travel decision and general wellbeing related to sleep/fatigue), followed by measures of verification intention (primary), reliance intention, trust calibration index, comprehension (0–8), perceived trustworthiness, usability/readability, cognitive load (0–10), manipulation checks, and a behavioral proxy (expanding optional “source information”). Analyses used intention-to-treat regression models with covariate adjustment.

Results:

Of 214 consenting respondents who started the survey, 200 were included in analysis (100 per arm). The safety UI bundle increased verification intention compared with baseline (mean 4.72 vs 4.41 on a 7-point scale; adjusted β=0.293, 95% CI 0.128 to 0.457; P<.001). Reliance intention did not increase (mean 4.97 vs 5.03; adjusted β=-0.105, 95% CI -0.239 to 0.029; P=.13). Trust calibration improved in the safety UI arm (trust calibration index mean -0.29 vs 0.29; adjusted β=-0.567, 95% CI -1.005 to -0.129; P=.01). Expansion of optional source information was numerically higher in the safety UI arm, although the adjusted confidence interval included the null (42% vs 27%; adjusted OR=1.76, 95% CI 0.95 to 3.27; P=.07). Comprehension remained high and similar across arms (mean 6.33 vs 6.32; adjusted β=-0.132, 95% CI -0.428 to 0.163; P=.38). Perceived trustworthiness was modestly lower in the safety UI arm (mean 5.20 vs 5.39; adjusted β=-0.199, 95% CI -0.382 to -0.016; P=.03). Usability/readability was unchanged, and cognitive load did not increase. Manipulation checks indicated higher cue recognition in the safety UI arm.

Conclusions:

In a randomized static-vignette survey of older Chinese adults, a brief safety UI bundle was associated with higher verification intention and a trust-calibration index consistent with lower over-reliance risk, without detectable reductions in comprehension or usability/readability. Because the intervention was tested as a bundle using screenshots and generic source labels, findings should be interpreted as evidence for a practical interface-level strategy rather than proof that any single cue caused the observed effects.


 Citation

Please cite as:

Yu J, Chen J, Ren A, Duan H, Meng H, Gao Z

Effects of a Safety User Interface Bundle on Verification Intentions in Generative AI Chat Use Among Older Chinese Adults: Randomized Vignette Survey

J Med Internet Res 2026;28:e94140

DOI: 10.2196/94140

PMID: 42600074

PMCID: 13475784

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.