Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Feb 3, 2026
Date Accepted: Jun 19, 2026

The final, peer-reviewed published version of this preprint can be found here:

Model and Task-Aware Test-Time Scaling Strategies for Large Language and Vision-Language Models in Medicine: Evaluation Study

Oh G, Kim S, Park S, Kim BH

Model and Task-Aware Test-Time Scaling Strategies for Large Language and Vision-Language Models in Medicine: Evaluation Study

J Med Internet Res 2026;28:e90693

DOI: 10.2196/90693

PMID: 42490549

Model and Task-Aware Test-Time Scaling Strategies for Large Language and Vision-Language Models in Medicine: Evaluation Study

  • Gyutaek Oh; 
  • Seoyeon Kim; 
  • Sangjoon Park; 
  • Byung-Hoon Kim

ABSTRACT

Background:

Test-time scaling has emerged as a promising method to enhance the reasoning capabilities of large language models (LLMs) and vision-language models (VLMs) during inference without additional training. While interest in medical AI is growing, the effectiveness of these strategies for VLMs and the optimal configurations for varying medical contexts remain underexplored.

Objective:

This study aims to conduct a comprehensive investigation of test-time scaling in the medical domain to identify model-aware and task-aware strategies. The study evaluates the impact of scaling on both LLMs and VLMs across different model sizes and task complexities, and assesses the robustness of these strategies against user-driven factors such as misleading information.

Methods:

The study evaluated a diverse set of models, including general and medical-specific LLMs and VLMs. Experiments utilized five textual medical benchmarks comprising over 5,500 questions, and two multimodal benchmarks comprising 7,000 samples. Performance was measured using accuracy and coverage under three scaling conditions: increasing token budgets (up to 8,192 tokens), iterative sequential scaling, and parallel scaling. Robustness was tested by embedding misleading hints (varying by tone and expertise) into prompts.

Results:

For non-reasoning LLMs, accuracy saturated quickly with token usage often remaining under 500 tokens regardless of budget increases. In contrast, reasoning models demonstrated significant performance gains on complex tasks as token budgets increased. Regarding scaling strategies, parallel scaling outperformed sequential scaling on easier tasks, where additional sequential steps led to accuracy declines; conversely, sequential scaling or increased token budgets proved more effective for complex reasoning and calculation tasks. For VLMs, test-time scaling showed limited effectiveness, with only the QVQ model exhibiting consistent accuracy improvements on the challenging dataset. Finally, while optimal scaling strategies improved robustness against misleading "expert" prompts, they did not fully restore performance to baseline levels.

Conclusions:

Longer reasoning is not universally beneficial in the medical domain; concise reasoning with parallel scaling is optimal for simpler tasks, while extended chain-of-thought via sequential scaling or increased budgets is required for complex problem-solving. Current medical VLMs and non-reasoning LLMs show limited benefit from simple token budget expansion. Effective implementation of test-time scaling requires strategies tailored to specific model characteristics and task difficulty levels to ensure reliability and robustness.


 Citation

Please cite as:

Oh G, Kim S, Park S, Kim BH

Model and Task-Aware Test-Time Scaling Strategies for Large Language and Vision-Language Models in Medicine: Evaluation Study

J Med Internet Res 2026;28:e90693

DOI: 10.2196/90693

PMID: 42490549

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.