Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR AI

Date Submitted: Jul 8, 2026
Open Peer Review Period: Jul 14, 2026 - Sep 8, 2026
(closed for review but you can still tweet)

NOTE: This is an unreviewed Preprint

Warning: This is a unreviewed preprint (What is a preprint?). Readers are warned that the document has not been peer-reviewed by expert/patient reviewers or an academic editor, may contain misleading claims, and is likely to undergo changes before final publication, if accepted, or may have been rejected/withdrawn (a note "no longer under consideration" will appear above).

Peer review me: Readers with interest and expertise are encouraged to sign up as peer-reviewer, if the paper is within an open peer-review period (in this case, a "Peer Review Me" button to sign up as reviewer is displayed above). All preprints currently open for review are listed here. Outside of the formal open peer-review period we encourage you to tweet about the preprint.

Citation: Please cite this preprint only for review purposes or for grant applications and CVs (if you are the author).

Final version: If our system detects a final peer-reviewed "version of record" (VoR) published in any journal, a link to that VoR will appear below. Readers are then encourage to cite the VoR instead of this preprint.

Settings: If you are the author, you can login and change the preprint display settings, but the preprint URL/DOI is supposed to be stable and citable, so it should not be removed once posted.

Submit: To post your own preprint, simply submit to any JMIR journal, and choose the appropriate settings to expose your submitted version as preprint.

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Quality Assessment of Artificial Intelligence Chatbot Responses to Post-Earthquake Health Frequently Asked Questions: A Cross-Platform, Cross-Language, and Cross-Mode Comparative Study

  • Haokun Wang; 
  • Youhua Lu; 
  • Kaixiang Nan; 
  • Qi Zhang; 
  • Leijie Qiu; 
  • Jiwen Wang

ABSTRACT

Background:

Large language model (LLM)-based chatbots may help disseminate scalable, evidence-informed health information during postearthquake crises, when health care infrastructure and professional consultation may be disrupted. However, the quality, readability, reliability, and short-term stability of artificial intelligence (AI)-generated guidance across platforms, languages, and reasoning modes remain insufficiently characterized.

Objective:

This study aimed to compare the quality, readability, and temporal stability of responses generated by four mainstream LLM-based chatbots to expert-reviewed postearthquake health frequently asked questions (FAQs), with attention to language and reasoning-mode differences.

Methods:

An expert-reviewed bank of 25 postearthquake health FAQs was submitted to four platforms (ChatGPT-5.5, Gemini 3.1 Flash, DeepSeek-V4, and Doubao) in English and Chinese under standard and platform-specific reasoning-enhanced ("Thinking") modes. In total, 600 single-turn responses were independently evaluated by two blinded disaster medicine experts using the Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P), Global Quality Score (GQS), and modified DISCERN (mDISCERN). Readability indices were calculated for English and Chinese responses. Temporal stability was assessed using two one-sided tests (TOST) for equivalence and Bland-Altman analysis, and semantic stability was assessed using cosine similarity and ROUGE-L text overlap.

Results:

Pooled interrater reliability was excellent in the primary Chinese standard-mode Round 1 subset (all pooled intraclass correlation coefficients [ICCs] for the ICC[2,k] model >0.90), and updated platform-stratified estimates were good to excellent across metrics (ICC[2,k] range 0.892-0.980). In the primary comparison (Chinese, standard mode, Round 1), significant overall platform effects were observed for all evaluated quality metrics (all P<.001), with ChatGPT-5.5 and Gemini 3.1 Flash generally ranking above Doubao and DeepSeek. Language effects were bidirectional: Chinese responses scored higher for understandability and actionability, whereas English responses scored higher for information reliability as measured by mDISCERN. The reasoning-enhanced "Thinking" mode did not consistently improve response quality. Instead, it decreased practical actionability across several platforms (eg, GPT, mean difference -0.62; P<.001), suggesting a potential trade-off between completeness and actionability. Temporal stability was generally acceptable at the group level, but individual-level variability remained nonnegligible; English outputs met equivalence criteria more often, whereas limits of agreement were metric- and platform-dependent. English responses from all platforms exceeded the recommended eighth-grade reading level.

Conclusions:

LLM-based chatbots show promise for postdisaster health communication, but their usefulness depends on platform, target language, and reasoning-mode configuration. Reasoning-enhanced modes may reduce the concise actionability needed for emergency instructions. Validation against authoritative public health guidance and testing with lay users are needed before real-world deployment in public health emergencies.


 Citation

Please cite as:

Wang H, Lu Y, Nan K, Zhang Q, Qiu L, Wang J

Quality Assessment of Artificial Intelligence Chatbot Responses to Post-Earthquake Health Frequently Asked Questions: A Cross-Platform, Cross-Language, and Cross-Mode Comparative Study

JMIR Preprints. 08/07/2026:106543

DOI: 10.2196/preprints.106543

URL: https://preprints.jmir.org/preprint/106543

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.