Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Mar 30, 2026
Date Accepted: Aug 10, 2026

The final, peer-reviewed published version of this preprint can be found here:

Bilingual Performance of Large Language Models in Answering Consumer Health Questions in English and Chinese: Comparative Benchmark Study

Wu Z, Xu H, Feng F, Cheng R, Shi T, Wang S, Zhang S, Li Y

Bilingual Performance of Large Language Models in Answering Consumer Health Questions in English and Chinese: Comparative Benchmark Study

J Med Internet Res 2026;28:e96587

DOI: 10.2196/96587

Bilingual Performance of Large Language Models in Consumer Health Questions Answering in English and Chinese: Comparative Benchmark Study

  • Ziyan Wu; 
  • Honglin Xu; 
  • Futai Feng; 
  • Rongrong Cheng; 
  • Tianqi Shi; 
  • Siyu Wang; 
  • Shulan Zhang; 
  • Yongzhe Li

ABSTRACT

Background:

Large language models (LLMs) increasingly serve as intermediaries for healthcare information. However, their cross-linguistic reliability and equity remain insufficiently characterized.

Objective:

To systematically quantify cross-linguistic performance disparities and identify model–language interaction patterns influencing the reliability and equity of LLM-mediated health information delivery.

Methods:

We conducted a controlled, bilingual evaluation of eleven widely deployed LLMs—including ChatGPT, Claude, Gemini, Grok, DeepSeek, Qwen, Doubao, Kimi, Hunyuan, ERNIE Bot, and ChatGLM—using 150 binary health questions from the TREC Health Misinformation Track. Models were assessed across different prompting strategies, evaluating accuracy, comprehensiveness, precision, and understandability. To quantitatively integrate the four metrics into a single composite score, we employed the Technique for Order Preference by Similarity to Ideal Solution (TOPSIS).

Results:

English and Chinese inputs showed comparable overall accuracy under the no-context condition (95.37% vs. 93.82%), language main effect (β=0.19, p=0.13), TOPSIS analysis identified ChatGPT and Qwen as Tier 1 performers universally. But significant model-specific language interactions emerged. Models like DeepSeek exhibited marked performance degradation in English across three quality metrics (all p < 0.05), whereas Kimi demonstrated a relative English advantage. Notably, error attribution prompt yielded the highest incremental gains (Δ53.13%), whereas expert prompt unexpectedly degraded performance.

Conclusions:

Although contemporary LLMs demonstrated robust bilingual accuracy, substantial language-dependent variations in communication quality persist, with critical implications for equitable global healthcare deployment. Cross-linguistic reliability should be recognized as a fundamental dimension of AI safety in digital medicine. These cross-linguistic performance disparities represent a critical dimension of AI health equity, as non-English speakers may receive lower-quality health information from widely deployed AI systems. Clinical Trial: NA


 Citation

Please cite as:

Wu Z, Xu H, Feng F, Cheng R, Shi T, Wang S, Zhang S, Li Y

Bilingual Performance of Large Language Models in Answering Consumer Health Questions in English and Chinese: Comparative Benchmark Study

J Med Internet Res 2026;28:e96587

DOI: 10.2196/96587

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.