Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Medical Informatics

Date Submitted: Feb 13, 2026
Date Accepted: Jul 13, 2026

The final, peer-reviewed published version of this preprint can be found here:

Discordance Between Textual Reasoning and Visual Interpretation in Large Language Models for Low Back Pain: Cross-Sectional Quantitative Evaluation and Exploratory Multimodal Stress Test

Zhang Z, Chen L, Lv Z, Lv H, Sheng W, Wei Z, Wang B, Shen Y, Tian Y, Hu J, Shen Z, Lv L

Discordance Between Textual Reasoning and Visual Interpretation in Large Language Models for Low Back Pain: Cross-Sectional Quantitative Evaluation and Exploratory Multimodal Stress Test

JMIR Med Inform 2026;14:e93522

DOI: 10.2196/93522

PMID: 42622617

PMCID: 13494679

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Discordance Between Textual Reasoning and Visual Interpretation in Large Language Models for Low Back Pain: Mixed-Methods Evaluation of Safety Risks

  • Ziyu Zhang; 
  • Longhao Chen; 
  • Zhizhen Lv; 
  • Hanzhe Lv; 
  • Wei Sheng; 
  • Zicheng Wei; 
  • Binghao Wang; 
  • Ying Shen; 
  • Yu Tian; 
  • Jingwen Hu; 
  • Zhifang Shen; 
  • Lijiang Lv

ABSTRACT

Background:

Background:

Large language models (LLMs) are rapidly evolving from text-based agents to multimodal systems capable of interpreting medical images. While the textual reasoning of these models has improved significantly, the safety implications of this shift—specifically regarding the alignment between visual interpretation and textual advice in low back pain (LBP) management—remain under-explored.

Objective:

Objective:

The objective of this study was to evaluate the performance of 10 state-of-the-art LLMs (including reasoning-enhanced models like ChatGPT-5 Plus and Qwen3-Thinking) in LBP contexts, specifically assessing the "asymmetric evolution" between their textual accuracy and visual diagnostic safety.

Methods:

Methods:

We employed a mixed-methods design. First, a quantitative evaluation was conducted using 25 standardized textual inquiries based on international guidelines (e.g., NICE, ACP), assessing 1,816 individual recommendations from 10 models. Outcomes included clinical accuracy (coded by a multidisciplinary panel), readability (Flesch Reading Ease [FRE], SMOG), understandability (PEMAT), and safety disclaimer coverage. Second, to probe multimodal safety risks, we conducted a qualitative adversarial stress test ($N=5$) using clinical cases with deliberate clinical-radiological mismatches (e.g., ankylosing spondylitis). Multimodal performance was scored on diagnostic consistency (1–4 scale) and descriptive accuracy.

Results:

Results:

In textual tasks, models achieved a high overall accuracy of 88.4% (1605/1816). ChatGPT-5 Plus demonstrated near-perfect performance (98.96% accuracy) with zero severe errors. The use of "Reasoning/Thinking" modes (e.g., Qwen3-Thinking) reduced the rate of serious errors to 1.64%, compared to 4.91% in standard models. However, textual advice remained difficult to read (mean SMOG 11.54 ± 1.94; 11th–12th grade level), although understandability scores were robust (PEMAT mean 85.6%). In contrast, multimodal performance exhibited a significant negative divergence. In pure imaging tasks, the average diagnostic score was poor (1.70/4.00), with models failing to identify core pathologies. Crucially, safety disclaimer coverage dropped precipitously from 82.8% in textual tasks to 28.0% in multimodal interactions. Models frequently exhibited "textual masking," hallucinating visual findings to align with the provided clinical history rather than the actual imaging evidence.

Conclusions:

Conclusions:

Current LLMs exhibit a dangerous capability mismatch: expert-level textual reasoning coexists with unreliable visual interpretation. This asymmetry creates a "high-confidence trap," where users may mistakenly extend their trust in the model's textual logic to its visual analysis. While reasoning models represent a significant leap in textual safety, the lack of cross-modal alignment and the systemic absence of safety disclaimers in visual tasks suggest that current multimodal features are not yet ready for direct patient use in LBP diagnostics. Clinical Trial: not applicable


 Citation

Please cite as:

Zhang Z, Chen L, Lv Z, Lv H, Sheng W, Wei Z, Wang B, Shen Y, Tian Y, Hu J, Shen Z, Lv L

Discordance Between Textual Reasoning and Visual Interpretation in Large Language Models for Low Back Pain: Cross-Sectional Quantitative Evaluation and Exploratory Multimodal Stress Test

JMIR Med Inform 2026;14:e93522

DOI: 10.2196/93522

PMID: 42622617

PMCID: 13494679

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.