Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Dec 20, 2025
Date Accepted: Jun 1, 2026
Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Healthcare: A Scoping Review
ABSTRACT
Background:
Inference-time retrieval augmentation is increasingly used to improve the reliability, traceability, and verifiability of large language model (LLM) applications in healthcare. However, evaluation practices for these systems, including text-based retrieval-augmented generation (RAG) and graph-structured retrieval-augmented generation (GraphRAG), remain heterogeneous, limiting interpretability, comparability, and assessment of clinical readiness
Objective:
This scoping review aimed to systematically map evaluation methods used for inference-time retrieval-augmented and graph-structured retrieval-augmented LLM systems in healthcare and to characterize how evaluation constructs are defined, operationalized, and reported across system layers.
Methods:
We conducted a scoping review in accordance with PRISMA-ScR, with search reporting informed by PRISMA-S. Searches were performed on December 18, 2025, in PubMed (MEDLINE), Web of Science Core Collection, IEEE Xplore, ACM Digital Library, arXiv, and medRxiv, with backward and forward citation tracking of included studies. Eligible records described healthcare-relevant LLM systems using inference-time retrieval augmentation and reported at least 1 evaluation component. Data were charted on study characteristics, system design, retrieval-layer evaluation, evidence linkage, safety-related and GraphRAG-specific evaluation, and selected reporting and governance characteristics.
Results:
A total of 102 studies met the inclusion criteria. Clinical question answering was the most frequently represented application (73/102, 71.6%), followed by clinical decision support (38/102, 37.3%). Most evaluations were conducted in offline-only settings (75/102, 73.5%), whereas only 24/102 (23.5%) reported at least 1 workflow-facing, prospective, or deployment-level evaluation setting. Independent retrieval-layer evaluation was reported in 38/102 studies (37.3%). Grounding- and faithfulness-related evaluation was reported in 90/102 studies (88.2%), but verification units were frequently coarse or incompletely specified; response-level verification was most common (42/102, 41.2%), whereas 29/102 studies (28.4%) had unclear verification units. Human evaluation was reported in 64/102 studies (62.7%), but interrater reliability was reported in only 17/64 (26.6%). LLM-based judging was reported in 20/102 studies (19.6%), with bias-control measures reported in 8/20 (40.0%). Explicit safety-oriented evaluation was reported in 23/102 studies (22.5%), with an additional 8/102 (7.8%) reporting partial or indirect safety-related assessment. Among the 12/102 studies (11.8%) meeting the prespecified GraphRAG minimum criterion, graph construction quality evaluation was reported in 5/12 (41.7%). Across domains, recurrent gaps included limited higher-realism evaluation, incomplete retrieval-layer assessment, sparse contradiction handling, limited explicit safety endpoints, and uneven reporting of implementation-relevant study details.
Conclusions:
Evaluation of healthcare RAG and GraphRAG systems has expanded rapidly but remains methodologically fragmented. Unlike prior reviews focused on applications or overall system performance, this review maps evaluation across retrieval, evidence linkage, safety, and clinical-use-proximity layers. By integrating these findings into a Clinical Realism Continuum and a practice-oriented Minimum Evaluation Package, this review provides an actionable framework for implementers, researchers, and evaluators. These stage-specific recommendations may help health systems move beyond superficial benchmark performance toward safer, more interpretable, and more governance-aware evaluation in clinical workflows. Clinical Trial: OSF Registries MTF5X; https://osf.io/mtf5x
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.