Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Mar 30, 2026
Date Accepted: Aug 10, 2026
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Cross-Linguistic Performance and Equity of Large Language Models in Medical Question Answering: Evaluating Eleven Large Language Models Across English and Chinese Contexts
ABSTRACT
Background:
Large language models (LLMs) increasingly serve as intermediaries for healthcare information. However, their cross-linguistic reliability and equity remain insufficiently characterized.
Objective:
To systematically quantify cross-linguistic performance disparities and identify model–language interaction patterns influencing the reliability and equity of LLM-mediated health information delivery.
Methods:
We conducted a controlled, bilingual evaluation of eleven widely deployed LLMs—including ChatGPT, Claude, Gemini, Grok, DeepSeek, Qwen, Doubao, Kimi, Hunyuan, ERNIE Bot, and ChatGLM—using 150 binary health questions from the TREC Health Misinformation Track. Models were assessed across different prompting strategies, evaluating accuracy, comprehensiveness, precision, and understandability. To quantitatively integrate the four metrics into a single composite score, we employed the Technique for Order Preference by Similarity to Ideal Solution (TOPSIS).
Results:
English and Chinese inputs showed comparable overall accuracy under the no-context condition (95.37% vs. 93.82%), language main effect (β=0.19, p=0.13), TOPSIS analysis identified ChatGPT and Qwen as Tier 1 performers universally. But significant model-specific language interactions emerged. Models like DeepSeek exhibited marked performance degradation in English across three quality metrics (all p < 0.05), whereas Kimi demonstrated a relative English advantage. Notably, error attribution prompt yielded the highest incremental gains (Δ53.13%), whereas expert prompt unexpectedly degraded performance.
Conclusions:
Although contemporary LLMs demonstrated robust bilingual accuracy, substantial language-dependent variations in communication quality persist, with critical implications for equitable global healthcare deployment. Cross-linguistic reliability should be recognized as a fundamental dimension of AI safety in digital medicine. These cross-linguistic performance disparities represent a critical dimension of AI health equity, as non-English speakers may receive lower-quality health information from widely deployed AI systems. Clinical Trial: NA
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.