Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Formative Research

Date Submitted: Jul 25, 2026
Open Peer Review Period: Aug 4, 2026 - Sep 29, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Large Language Models for Nutritional Counseling: A Comparative Evaluation Across Four Clinical Scenarios.

  • Adam Hruška; 
  • Anna Jílková; 
  • Andrea Maťhová; 
  • Barbora Lampová; 
  • Daria Musiienko

ABSTRACT

Background:

Large language models (LLMs) are increasingly used by the public for health information, yet their reliability for individualized nutrition counseling remains uncertain.

Objective:

This study aims to benchmark five freely available LLM families across four clinically relevant nutrition scenarios, evaluating their performance on accuracy, personalization, comprehensiveness, user-friendliness, safety, and overall quality.

Methods:

In a repeated-measures expert evaluation design, standardized personas representing obesity, osteoporosis, inflammatory bowel disease, and endometriosis were used to generate nutrition consultations from ChatGPT, Gemini, DeepSeek, LeChat, and LLama in both 2025 and 2026 versions. Ten independent experts rated each output on a 1–5 Likert-type rubric, yielding 400 evaluations in total. Agreement was assessed with ICCs and Kendall’s W. The model comparisons used Friedman and paired Wilcoxon tests, with supplementary linear mixed-effects models adjusting for expert and persona clustering.

Results:

Model performance showed a stable hierarchy across analyses: ChatGPT and Gemini formed the top tier, DeepSeek occupied an intermediate position, consistently followed by LeChat, whereas LLama performed worst. Aggregated inter-rater reliability was excellent across all criteria, although individual-rater agreement was only fair to moderate. Between 2025 and 2026, most models improved, with the largest gains observed for ChatGPT and DeepSeek. Personalization remained the weakest domain, and all models showed systematic Western-centric dietary defaults, limiting cultural and contextual fit.

Conclusions:

Contemporary LLMs can generate broadly useful nutrition advice under benchmark conditions, but they remain insufficiently reliable for autonomous individualized counseling. Their main limitations are incomplete personalization, variable safety, and weak cultural adaptation, supporting a cautious role as adjunct tools rather than replacements for qualified nutrition professionals.


 Citation

Please cite as:

Hruška A, Jílková A, Maťhová A, Lampová B, Musiienko D

Large Language Models for Nutritional Counseling: A Comparative Evaluation Across Four Clinical Scenarios.

JMIR Preprints. 25/07/2026:107909

DOI: 10.2196/preprints.107909

URL: https://preprints.jmir.org/preprint/107909

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.