Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Feb 13, 2026
Date Accepted: Jul 13, 2026
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Discordance Between Textual Reasoning and Visual Interpretation in Large Language Models for Low Back Pain: Mixed-Methods Evaluation of Safety Risks
ABSTRACT
Background:
Background:
Large language models (LLMs) are rapidly evolving from text-based agents to multimodal systems capable of interpreting medical images. While the textual reasoning of these models has improved significantly, the safety implications of this shift—specifically regarding the alignment between visual interpretation and textual advice in low back pain (LBP) management—remain under-explored.
Objective:
Objective:
The objective of this study was to evaluate the performance of 10 state-of-the-art LLMs (including reasoning-enhanced models like ChatGPT-5 Plus and Qwen3-Thinking) in LBP contexts, specifically assessing the "asymmetric evolution" between their textual accuracy and visual diagnostic safety.
Methods:
Methods:
We employed a mixed-methods design. First, a quantitative evaluation was conducted using 25 standardized textual inquiries based on international guidelines (e.g., NICE, ACP), assessing 1,816 individual recommendations from 10 models. Outcomes included clinical accuracy (coded by a multidisciplinary panel), readability (Flesch Reading Ease [FRE], SMOG), understandability (PEMAT), and safety disclaimer coverage. Second, to probe multimodal safety risks, we conducted a qualitative adversarial stress test ($N=5$) using clinical cases with deliberate clinical-radiological mismatches (e.g., ankylosing spondylitis). Multimodal performance was scored on diagnostic consistency (1–4 scale) and descriptive accuracy.
Results:
Results:
In textual tasks, models achieved a high overall accuracy of 88.4% (1605/1816). ChatGPT-5 Plus demonstrated near-perfect performance (98.96% accuracy) with zero severe errors. The use of "Reasoning/Thinking" modes (e.g., Qwen3-Thinking) reduced the rate of serious errors to 1.64%, compared to 4.91% in standard models. However, textual advice remained difficult to read (mean SMOG 11.54 ± 1.94; 11th–12th grade level), although understandability scores were robust (PEMAT mean 85.6%). In contrast, multimodal performance exhibited a significant negative divergence. In pure imaging tasks, the average diagnostic score was poor (1.70/4.00), with models failing to identify core pathologies. Crucially, safety disclaimer coverage dropped precipitously from 82.8% in textual tasks to 28.0% in multimodal interactions. Models frequently exhibited "textual masking," hallucinating visual findings to align with the provided clinical history rather than the actual imaging evidence.
Conclusions:
Conclusions:
Current LLMs exhibit a dangerous capability mismatch: expert-level textual reasoning coexists with unreliable visual interpretation. This asymmetry creates a "high-confidence trap," where users may mistakenly extend their trust in the model's textual logic to its visual analysis. While reasoning models represent a significant leap in textual safety, the lack of cross-modal alignment and the systemic absence of safety disclaimers in visual tasks suggest that current multimodal features are not yet ready for direct patient use in LBP diagnostics. Clinical Trial: not applicable
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.