Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Jan 27, 2026
Date Accepted: Jul 29, 2026

The final, peer-reviewed published version of this preprint can be found here:

Improving Reliability and Explainability of Medical Question Answering Through Atomic Fact-Checking in Retrieval-Augmented Large Language Models: Creation and Validation Study

Vladika J, Domres A, Nguyen M, Moser R, Nano J, Busch F, Adams L, Bressem KK, Bernhardt D, Combs S, Borm K, Matthes F, Peeken JC

Improving Reliability and Explainability of Medical Question Answering Through Atomic Fact-Checking in Retrieval-Augmented Large Language Models: Creation and Validation Study

J Med Internet Res 2026;28:e92090

DOI: 10.2196/92090

Improving Reliability and Explainability of Medical Question Answering through Atomic Fact-Checking in Retrieval-Augmented LLMs: Creation and Validation

  • Juraj Vladika; 
  • Annika Domres; 
  • Mai Nguyen; 
  • Rebecca Moser; 
  • Jana Nano; 
  • Felix Busch; 
  • Lisa Adams; 
  • Keno K Bressem; 
  • Denise Bernhardt; 
  • Stephanie Combs; 
  • Kai Borm; 
  • Florian Matthes; 
  • Jan C Peeken

ABSTRACT

Background:

Large language models (LLMs) exhibit extensive medical knowledge but are prone to hallucinations and show low fact-level explainability, limiting clinical adoption and regulatory compliance. Existing approaches, such as Retrieval Augmented Generation, partially address these issues by grounding answers in source documents, however, the aforementioned problems persist.

Objective:

We propose a novel atomic fact-checking framework, designed to enhance the reliability and explainability of LLMs in medical long-form question answering. By decomposing generated answers into discrete atomic facts, and verifying each against an authoritative knowledge base of medical guidelines, precise identification and correction of incorrect statements, alongside explicit linkage to supporting literature, are enabled.

Methods:

The Fact-Checking algorithm operates within a Retrieval Augmented Generation (RAG) framework: LLM-generated answers are decomposed into atomic facts (smallest, self-contained information units), each is assessed and corrected if FALSE. To determine an optimal strategy, a Validation-QA-Set on prostate cancer treatment was tested under varying instructions. An extensive evaluation, including multi-reader assessments by human medical experts and an automated open Q&A benchmark, AMEGA, was conducted on the final pipeline. In addition to another radiooncologic Test-QA-Set, anonymized real tumor board cases and a content-wise unrelated NeurologyQA-Set were used. Given their transparency and accessibility advantages, we compared various open-source models in pairs of generalist models and their medical finetuned counterparts.

Results:

The framework significantly reduced hallucinations and inaccuracies. Medical expert assessment and automated benchmarks demonstrated significant improvements in factual accuracy, achieving up to 40% overall answer improvement and 80% hallucination detection rate. Notably, the observed gain was strongest in real tumor-board questions, the most challenging dataset. Additionally, the framework achieved high explainability by tracing each atomic fact back to the most relevant chunks from the database, providing a granular, transparent explanation of the generated responses.

Conclusions:

To conclude, we present a novel atomic fact-checking algorithm that identifies fact inaccuracies and hallucinations. Correcting these findings improves the overall answer quality while achieving fact-wise explainability, paving the way for more trustworthy clinical use of LLMs.


 Citation

Please cite as:

Vladika J, Domres A, Nguyen M, Moser R, Nano J, Busch F, Adams L, Bressem KK, Bernhardt D, Combs S, Borm K, Matthes F, Peeken JC

Improving Reliability and Explainability of Medical Question Answering Through Atomic Fact-Checking in Retrieval-Augmented Large Language Models: Creation and Validation Study

J Med Internet Res 2026;28:e92090

DOI: 10.2196/92090

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.