Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Multilingual Disparities in LLM-Based Symptom Detection for Global Disease Surveillance: Findings from High- and Low-Resource Languages
ABSTRACT
Background:
Symptom detection is essential in global disease surveillance to detect potential outbreaks, as symptoms are the first observable signs of infection. To reflect real-time ground truth conditions during pandemics, social media has emerged as a valuable data source. Moreover, effective digital disease surveillance systems must operate across diverse linguistic settings, but large language models (LLMs) have been shown to perform inconsistently across languages, with a tendency to have lower performance on low-resource languages. While multilingual approaches have been explored in various health-related NLP tasks, a critical gap remains in understanding whether LLM-based symptom detection can perform consistently across languages for global disease surveillance. Southeast Asia demonstrates this challenge, combining diverse languages and the potential for emerging infectious disease outbreaks, making it a case for evaluating how multilingual performance disparities manifest in symptom detection.
Objective:
This study aims to evaluate multilingual disparities in symptom detection using a large language model across languages and symptom types to better understand their implications for global disease surveillance.
Methods:
This study uses the MedWeb dataset, a multilingual pseudo–social media text dataset with multiple symptom labels. The data consist of 12 languages, covering diverse regions and language resource classifications. The symptoms included in this study are fever, headache, runny nose, cough, diarrhea, hay fever, influenza, and cold. We employed GPT-5 as the symptom detection system, representing a strong model for health-related tasks. The results are evaluated using the macro precision, recall, and F1-score.
Results:
Our findings show that performance varied across languages, language resource groups, and symptom types. Southeast Asian (SEA) languages generally achieved lower scores than non-SEA, with Japanese obtaining the highest F1-score (0.812) and Lao the lowest (0.716). High-resource languages achieved the most consistent performance, while low-resource languages obtained the lowest overall scores. At the symptom level, Diarrhea, Headache, Cough, and Fever showed stable detection across languages, while Hay fever, Runny nose, Influenza, and Cold exhibited greater variability. Hay fever showed the widest variability, forming two distinct performance clusters, aligned with language resource and region classification. The error analysis revealed four misclassification patterns: explicitly mentioned symptoms, cross-lingual variation, symptom overgeneralization, and context misinterpretation.
Conclusions:
Due to the performance disparities, to achieve more reliable and equitable LLM-based symptom detection from social media text for global disease surveillance, it would benefit from broader representation of training data for low-resource languages, improved cultural-linguistic sensitivity, and stronger contextual understanding of symptom-related expressions.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.