Diagnostic Performance of a Locally Deployed Vision-Language Model for Bone Tumor Diagnosis Using Smartphone-Captured Images: An Exploratory Retrospective Study
ABSTRACT
Background:
Vision-language models (VLMs) show promise in medical imaging, yet their performance on high-noise smartphone-captured images—common in primary care referrals—remains untested. Furthermore, it remains controversial whether Retrieval-Augmented Generation (RAG) using external expert guidelines actually improves diagnostic accuracy for rare bone tumors.
Objective:
This study evaluates the diagnostic efficacy of a locally deployed, open-source VLM (Qwen3-VL) on high-noise bone tumor images. We investigate how clinical persona prompts and RAG integration affect the model's cognitive boundaries.
Methods:
This retrospective study included 42 patients with biopsy-proven primary bone tumors and tumor-like lesions. To simulate real-world conditions, we captured the original DICOM images from a monitor using a handheld smartphone without stabilization, organically capturing ambient glare and Moiré patterns typical of real-world teleconsultations. Using a 2×2 factorial design, we compared the diagnostic performance of the base model versus the RAG-integrated model under two distinct system personas: "Radiologist" and "Orthopedic Oncologist." Primary outcomes were Top-1 and Top-3 diagnostic accuracy. We used the McNemar test for inter-group comparisons and conducted a qualitative analysis of AI hallucinations using the model's Chain-of-Thought (CoT) logs.
Results:
Without RAG, the Top-3 accuracy showed no significant difference between the radiologist and orthopedic oncologist personas (35.7% vs. 33.3%, p = 0.763). After RAG integration, the radiologist persona experienced a statistically significant drop in Top-1 accuracy (from 28.6% to 14.3%, p = 0.034). Conversely, the orthopedic oncologist persona demonstrated strong clinical resilience against RAG-induced noise, with no significant changes in Top-1 or Top-3 accuracy (p = 0.655 and p > 0.999, respectively). Qualitative CoT analysis revealed a severe "Text-over-Vision" defect during multimodal fusion. The models were easily misled by text priors such as patient age, trauma history, and RAG-retrieved data, leading to demographic anchoring, premature visual closure, and rare-disease hallucinations.
Conclusions:
For the complex morphological differentiation of bone tumors, blindly attaching text-based knowledge bases via RAG fails to improve VLM accuracy and instead triggers "reverse hallucinations" by encouraging the model to take cognitive shortcuts. Compared to a standard radiologist persona, an orthopedic surgical persona embedding multidimensional differential logic creates an effective "clinical moat" against AI text-bias. Future development in medical AI requires extensive "dirty data" from real-world settings to perform deep visual fine-tuning on the underlying vision encoders.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.