Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Dec 20, 2025
Date Accepted: Jun 1, 2026

The final, peer-reviewed published version of this preprint can be found here:

Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Health Care: Scoping Review

Zhao Y, Miao Y, Guo R, Luo Y, Wang H, Wu Y

Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Health Care: Scoping Review

J Med Internet Res 2026;28:e90046

DOI: 10.2196/90046

PMID: 42546264

Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Healthcare: A Scoping Review

  • Yuhan Zhao; 
  • Yiqun Miao; 
  • Rongrong Guo; 
  • Yuan Luo; 
  • Huiying Wang; 
  • Ying Wu

ABSTRACT

Background:

Inference-time retrieval augmentation is increasingly used to improve the reliability, traceability, and verifiability of large language model (LLM) applications in healthcare. However, evaluation practices for these systems, including text-based retrieval-augmented generation (RAG) and graph-structured retrieval-augmented generation (GraphRAG), remain heterogeneous, limiting interpretability, comparability, and assessment of clinical readiness

Objective:

This scoping review aimed to systematically map evaluation methods used for inference-time retrieval-augmented and graph-structured retrieval-augmented LLM systems in healthcare and to characterize how evaluation constructs are defined, operationalized, and reported across system layers.

Methods:

We conducted a scoping review in accordance with PRISMA-ScR, with search reporting informed by PRISMA-S. Searches were performed on December 18, 2025, in PubMed (MEDLINE), Web of Science Core Collection, IEEE Xplore, ACM Digital Library, arXiv, and medRxiv, with backward and forward citation tracking of included studies. Eligible records described healthcare-relevant LLM systems using inference-time retrieval augmentation and reported at least 1 evaluation component. Data were charted on study characteristics, system design, retrieval-layer evaluation, evidence linkage, safety-related and GraphRAG-specific evaluation, and selected reporting and governance characteristics.

Results:

A total of 102 studies met the inclusion criteria. Clinical question answering was the most frequently represented application (73/102, 71.6%), followed by clinical decision support (38/102, 37.3%). Most evaluations were conducted in offline-only settings (75/102, 73.5%), whereas only 24/102 (23.5%) reported at least 1 workflow-facing, prospective, or deployment-level evaluation setting. Independent retrieval-layer evaluation was reported in 38/102 studies (37.3%). Grounding- and faithfulness-related evaluation was reported in 90/102 studies (88.2%), but verification units were frequently coarse or incompletely specified; response-level verification was most common (42/102, 41.2%), whereas 29/102 studies (28.4%) had unclear verification units. Human evaluation was reported in 64/102 studies (62.7%), but interrater reliability was reported in only 17/64 (26.6%). LLM-based judging was reported in 20/102 studies (19.6%), with bias-control measures reported in 8/20 (40.0%). Explicit safety-oriented evaluation was reported in 23/102 studies (22.5%), with an additional 8/102 (7.8%) reporting partial or indirect safety-related assessment. Among the 12/102 studies (11.8%) meeting the prespecified GraphRAG minimum criterion, graph construction quality evaluation was reported in 5/12 (41.7%). Across domains, recurrent gaps included limited higher-realism evaluation, incomplete retrieval-layer assessment, sparse contradiction handling, limited explicit safety endpoints, and uneven reporting of implementation-relevant study details.

Conclusions:

Evaluation of healthcare RAG and GraphRAG systems has expanded rapidly but remains methodologically fragmented. Unlike prior reviews focused on applications or overall system performance, this review maps evaluation across retrieval, evidence linkage, safety, and clinical-use-proximity layers. By integrating these findings into a Clinical Realism Continuum and a practice-oriented Minimum Evaluation Package, this review provides an actionable framework for implementers, researchers, and evaluators. These stage-specific recommendations may help health systems move beyond superficial benchmark performance toward safer, more interpretable, and more governance-aware evaluation in clinical workflows. Clinical Trial: OSF Registries MTF5X; https://osf.io/mtf5x


 Citation

Please cite as:

Zhao Y, Miao Y, Guo R, Luo Y, Wang H, Wu Y

Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Health Care: Scoping Review

J Med Internet Res 2026;28:e90046

DOI: 10.2196/90046

PMID: 42546264

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.