Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Formative Research

Date Submitted: Apr 11, 2026
Date Accepted: Jul 1, 2026

The final, peer-reviewed published version of this preprint can be found here:

Evaluating Retrieval-Augmented Large Language Models on Anesthesiology Board-Style Questions: Benchmark Study

Phuong NQ, Ruan SJ, Chen Pf

Evaluating Retrieval-Augmented Large Language Models on Anesthesiology Board-Style Questions: Benchmark Study

JMIR Form Res 2026;10:e97902

DOI: 10.2196/97902

PMID: 42579838

Evaluating Retrieval-Augmented Large Language Models on Anesthesiology Board-Style Questions : A Benchmark Study

  • Nguyen Quang Phuong; 
  • Shanq-Jang Ruan; 
  • Pei-fu Chen

ABSTRACT

Background:

Open-source, mid-scale large language models (LLMs) have emerged as scalable, privacy-preserving alternatives to ultra-large foundation models (eg, GPT-4) in health care systems. Techniques such as retrieval-augmented generation (RAG) enable sub-100-billion-parameter models to address highly specialized medical domains such as anesthesiology. However, studies evaluating RAG architectures on complex medical examinations remain scarce, highlighting the need for rigorous benchmarking to bridge the gap between raw parametric knowledge and clinically relevant application.

Objective:

This study aimed to systematically evaluate retrieval-augmented generation (RAG) pipelines for answering anesthesiology board-style questions, quantify the effects of key design choices including hyperparameter settings, embedding models, source complexity, and chunking strategies, and compare the performance of reasoning-oriented models with that of conventional LLMs.

Methods:

We conducted large-scale benchmarking using American Board of Anesthesiology-style multiple-choice questions to compare multiple RAG-enabled configurations with matched standalone LLM baselines. Configurations were first optimized on a 46-item diagnostic set and then validated on a 350-item corpus. Additional experiments on three 100-question subsets derived from the 350-item corpus were used to assess the effects of source selection, source complexity, information density, and chunking strategy on answer accuracy. Models including LLaMA-3-8B-Instruct, LLaMA-3.1-8B-Instruct, LLaMA-3.2-3B-Instruct, LLaMA-3.3-70B-Instruct, Qwen2.5-7B and Qwen2.5-72B, and Qwen3-8B and Qwen3-32B reasoning models were evaluated under this framework. Self-RAG with adaptive retrieval techniques was also implemented and evaluated. Cochran’s Q and McNemar’s test were used to assess performance differences across configurations and model pairs.

Results:

The RAG framework significantly increased the number of correct answers. System stability peaked under highly deterministic sampling configurations (temperature = 0.1, top-p = 0.1). High-capacity general-text embeddings and applying context-preserving semantic chunking further improved accuracy. Standard RAG provided only modest gains over nonaugmented baselines, improving accuracy from 50.29% to 56.57%, and Self-RAG yielded similarly limited gains of up to 4.85 percentage points. Overall, the Qwen family outperformed the LLaMA series. The 32-billion-parameter reasoning model Qwen-3-32B achieved an 89% correct ratio under complex distractor-heavy retrieval conditions and up to 96% with direct context, significantly outperforming the much larger 72-billion-parameter conventional model Qwen-2.5-72B-Instruct (84%). Smaller reasoning models also showed greater robustness to noise or suboptimal retrieved documents than larger conventional LLMs. Within the LLaMA family, increasing parameter size to 70 billion did not produce proportional performance gains on this benchmark.

Conclusions:

RAG-based LLM systems improved performance on anesthesiology board-style questions, but gains depended strongly on retrieval design. Careful optimization of retrieval settings, embeddings, and chunking strategies improved robustness and answer accuracy. Reasoning-oriented models demonstrated that multi-step reasoning can, in some settings, compensate for larger parameter scale. These findings provide a practical foundation for developing locally deployable LLM systems for anesthesiology education.


 Citation

Please cite as:

Phuong NQ, Ruan SJ, Chen Pf

Evaluating Retrieval-Augmented Large Language Models on Anesthesiology Board-Style Questions: Benchmark Study

JMIR Form Res 2026;10:e97902

DOI: 10.2196/97902

PMID: 42579838

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.