Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR AI

Date Submitted: Apr 29, 2026
Date Accepted: Aug 20, 2026

The final, peer-reviewed published version of this preprint can be found here:

Diagnostic Performance of a Locally Deployed Vision Language Model for Bone Tumor Diagnosis Using Smartphone-Captured Images: Exploratory Retrospective Study

Shi Y, Tong ZC

Diagnostic Performance of a Locally Deployed Vision Language Model for Bone Tumor Diagnosis Using Smartphone-Captured Images: Exploratory Retrospective Study

JMIR AI 2026;5:e99757

DOI: 10.2196/99757

PMID: 42766848

Diagnostic Performance of a Locally Deployed Vision-Language Model for Bone Tumor Diagnosis Using Smartphone-Captured Images: An Exploratory Retrospective Study

  • Yin Shi; 
  • Zhi-Chao Tong

ABSTRACT

Background:

Vision-language models (VLMs) show promise in medical imaging, yet their performance on high-noise smartphone-captured images—common in primary care referrals—remains untested. Furthermore, it remains controversial whether Retrieval-Augmented Generation (RAG) using external expert guidelines actually improves diagnostic accuracy for rare bone tumors.

Objective:

This study evaluates the diagnostic efficacy of a locally deployed, open-source VLM (Qwen3-VL) on high-noise bone tumor images. We investigate how clinical persona prompts and RAG integration affect the model's cognitive boundaries.

Methods:

This retrospective study included 42 patients with biopsy-proven primary bone tumors and tumor-like lesions. To simulate real-world conditions, we captured the original DICOM images from a monitor using a handheld smartphone without stabilization, organically capturing ambient glare and Moiré patterns typical of real-world teleconsultations. Using a 2×2 factorial design, we compared the diagnostic performance of the base model versus the RAG-integrated model under two distinct system personas: "Radiologist" and "Orthopedic Oncologist." Primary outcomes were Top-1 and Top-3 diagnostic accuracy. We used the McNemar test for inter-group comparisons and conducted a qualitative analysis of AI hallucinations using the model's Chain-of-Thought (CoT) logs.

Results:

Without RAG, the Top-3 accuracy showed no significant difference between the radiologist and orthopedic oncologist personas (35.7% vs. 33.3%, p = 0.763). After RAG integration, the radiologist persona experienced a statistically significant drop in Top-1 accuracy (from 28.6% to 14.3%, p = 0.034). Conversely, the orthopedic oncologist persona demonstrated strong clinical resilience against RAG-induced noise, with no significant changes in Top-1 or Top-3 accuracy (p = 0.655 and p > 0.999, respectively). Qualitative CoT analysis revealed a severe "Text-over-Vision" defect during multimodal fusion. The models were easily misled by text priors such as patient age, trauma history, and RAG-retrieved data, leading to demographic anchoring, premature visual closure, and rare-disease hallucinations.

Conclusions:

For the complex morphological differentiation of bone tumors, blindly attaching text-based knowledge bases via RAG fails to improve VLM accuracy and instead triggers "reverse hallucinations" by encouraging the model to take cognitive shortcuts. Compared to a standard radiologist persona, an orthopedic surgical persona embedding multidimensional differential logic creates an effective "clinical moat" against AI text-bias. Future development in medical AI requires extensive "dirty data" from real-world settings to perform deep visual fine-tuning on the underlying vision encoders.


 Citation

Please cite as:

Shi Y, Tong ZC

Diagnostic Performance of a Locally Deployed Vision Language Model for Bone Tumor Diagnosis Using Smartphone-Captured Images: Exploratory Retrospective Study

JMIR AI 2026;5:e99757

DOI: 10.2196/99757

PMID: 42766848

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.