Currently submitted to: Transfer Hub (manuscript eXchange)
Date Submitted: Jul 20, 2026
Open Peer Review Period: Jul 31, 2026 - Sep 25, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Safety, Accuracy, Empathy, Reliability, and Readability of Responses From Large Language Models to Public-Facing Questions About Nutrition in Hepatocellular Carcinoma: A Comparative Study
ABSTRACT
Background:
Patients with hepatocellular carcinoma (HCC) and their caregivers increasingly use large language models (LLMs) to obtain nutrition-related information. However, the safety and quality of such information remain uncertain, particularly because HCC frequently coexists with cirrhosis, sarcopenia, ascites, hepatic encephalopathy, treatment-related adverse effects, and potential food–drug interactions.
Objective:
To compare the safety, accuracy, empathy, reliability, information quality, and readability of responses generated by five publicly accessible chatbots to public-facing questions about nutrition in HCC.
Methods:
This cross-sectional comparative study evaluated ChatGPT, Gemini, Microsoft Copilot, Doubao, and DeepSeek using 62 standardized English-language questions about HCC-related nutrition. Each question was submitted to each model in a new conversation through its official web interface between June 2 and 6, 2026. Five evaluators independently assessed safety, accuracy, and empathy and evaluated reliability and information quality using DISCERN, Ensuring Quality Information for Patients (EQIP), the Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Score (GQS). Six automated indices were used to assess readability. Model comparisons were conducted using paired question-level analyses.
Results:
A total of 310 responses were analyzed. Inter-rater agreement was high, with a Fleiss κ of 0.862 for safety and ICC(2,1) values ranging from 0.846 to 0.888 for the other rater-scored outcomes. Overall, 263 responses were classified as safe and 47 as potentially unsafe. Safe response rates were 91.9% for ChatGPT, 85.5% for Gemini, 83.9% for Copilot, 79.0% for DeepSeek, and 83.9% for Doubao. The overall difference in safety was not statistically significant. Accuracy, empathy, reliability, information quality, and readability differed significantly across chatbots. Potentially unsafe responses most commonly involved incomplete escalation advice, insufficient qualification of general advice, overgeneralized dietary or supplement recommendations, and inadequate attention to food safety or food–drug interaction risks.
Conclusions:
Across this standardized set of public-facing HCC nutrition questions, most responses generated by the evaluated chatbots were classified as safe; however, potentially unsafe or insufficiently qualified advice was identified in clinically relevant scenarios. LLM-generated nutrition information may supplement general patient education but should not replace individualized guidance from oncologists, hepatologists, pharmacists, or dietitians. Future system development should prioritize clear advice on when red-flag symptoms require urgent clinical assessment, clinically appropriate qualification, source transparency, and plain-language communication.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.