Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Jan 27, 2026
Date Accepted: Jul 29, 2026
Improving Reliability and Explainability of Medical Question Answering through Atomic Fact-Checking in Retrieval-Augmented LLMs: Creation and Validation
ABSTRACT
Background:
Large language models (LLMs) exhibit extensive medical knowledge but are prone to hallucinations and show low fact-level explainability, limiting clinical adoption and regulatory compliance. Existing approaches, such as Retrieval Augmented Generation, partially address these issues by grounding answers in source documents, however, the aforementioned problems persist.
Objective:
We propose a novel atomic fact-checking framework, designed to enhance the reliability and explainability of LLMs in medical long-form question answering. By decomposing generated answers into discrete atomic facts, and verifying each against an authoritative knowledge base of medical guidelines, precise identification and correction of incorrect statements, alongside explicit linkage to supporting literature, are enabled.
Methods:
The Fact-Checking algorithm operates within a Retrieval Augmented Generation (RAG) framework: LLM-generated answers are decomposed into atomic facts (smallest, self-contained information units), each is assessed and corrected if FALSE. To determine an optimal strategy, a Validation-QA-Set on prostate cancer treatment was tested under varying instructions. An extensive evaluation, including multi-reader assessments by human medical experts and an automated open Q&A benchmark, AMEGA, was conducted on the final pipeline. In addition to another radiooncologic Test-QA-Set, anonymized real tumor board cases and a content-wise unrelated NeurologyQA-Set were used. Given their transparency and accessibility advantages, we compared various open-source models in pairs of generalist models and their medical finetuned counterparts.
Results:
The framework significantly reduced hallucinations and inaccuracies. Medical expert assessment and automated benchmarks demonstrated significant improvements in factual accuracy, achieving up to 40% overall answer improvement and 80% hallucination detection rate. Notably, the observed gain was strongest in real tumor-board questions, the most challenging dataset. Additionally, the framework achieved high explainability by tracing each atomic fact back to the most relevant chunks from the database, providing a granular, transparent explanation of the generated responses.
Conclusions:
To conclude, we present a novel atomic fact-checking algorithm that identifies fact inaccuracies and hallucinations. Correcting these findings improves the overall answer quality while achieving fact-wise explainability, paving the way for more trustworthy clinical use of LLMs.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.