Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR AI

Date Submitted: Aug 10, 2026
Open Peer Review Period: Aug 17, 2026 - Oct 12, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Answer Disagreement Detects Unreliable Large Language Model Responses in Medical Question Answering, but Deliberate Prompting Does Not Prevent Them: A Pre-Registered Evaluation of Seven Open-Weight Models

  • Bahadır Eryılmaz; 
  • Mikel Bahn; 
  • Hendrik Damm; 
  • Bohao Chu; 
  • Noëlle Bender; 
  • Jens Kleesiek; 
  • Felix Nensa

ABSTRACT

Background:

Large language models (LLMs) answering medical questions produce fluent but sometimes incorrect responses, and a clinically safe system must either avoid such errors or flag them. Deliberate multi-step prompting is widely proposed to prevent them, and sampling-based uncertainty estimation to detect them. Self-consistency prompting connects the two: it resamples a question several times, then discards the disagreement among those samples in favor of a majority vote.

Objective:

We tested whether deliberate prompting prevents unreliable answers, and whether the discarded disagreement detects them.

Methods:

We conducted a pre-registered computational evaluation (Open Science Framework) crossing 12 prompting techniques, organized on a System-1/System-2 dual-process ladder from direct answering to multi-pass verification, with 5 datasets and 7 open-weight LLMs at 10 random seeds per cell. Datasets were PubMedQA, MIMIC-IV-ED emergency-department disposition, LongHealth, BioASQ Task 13B, and MedHallu. All inference was local with models pinned by commit hash; no clinical text was sent to hosted application programming interfaces, as required by the PhysioNet Data Use Agreement. A separate pass resampled each question's zero-shot answer 10 times; we scored the entropy of the resulting answer distribution against correctness by area under the receiver operating characteristic curve (AUROC). Accuracy contrasts used a paired bootstrap (10,000 resamples) with Wilcoxon signed-rank tests and Holm-Bonferroni correction.

Results:

Deliberate prompting did not reliably improve accuracy. Of 176 corrected technique-versus-baseline contrasts across the 4 primary models, 52 were significantly better and 54 significantly worse; pooled bucket effects were +1.3 percentage points for single-pass System-2, −1.5 for multi-pass verification and −0.2 for integration, and the best technique added 2.9 points. The same 10-sample budget used for voting changed accuracy by +0.15 points, whereas the disagreement among those samples detected incorrect answers with AUROC 0.73 (95% CI 0.70–0.75; n=1559) on LongHealth, 0.69 (0.65–0.72; n=1040) on BioASQ, 0.58 (0.56–0.60; n=1990) on PubMedQA and 0.56 (0.53–0.58; n=1930) on MIMIC-IV-ED. Abstaining on the least confident half raised accuracy on the answered subset by 10.2, 10.3, 4.9 and 4.3 points respectively, each above a matched random-abstention baseline. Detectability tracked error arbitrariness: self-agreement for correct versus incorrect answers was 0.94 versus 0.77 on LongHealth but 0.83 versus 0.79 on MIMIC-IV-ED. An entailment-based estimator confirmed the free-text result (AUROC 0.692 entailment-based versus 0.686 rule-based).

Conclusions:

In the open-weight, zero-shot, small-sample regime typical of local deployment, deliberate prompting is not a reliable safety lever, whereas the disagreement self-consistency discards is an informative and essentially free reliability signal. Its limitation is systematic: it detects arbitrary errors and is blind to confident, reproducible ones, the errors of greatest clinical concern. Sampling-based uncertainty should therefore serve as a triage filter, never as a certificate of correctness.


 Citation

Please cite as:

Eryılmaz B, Bahn M, Damm H, Chu B, Bender N, Kleesiek J, Nensa F

Answer Disagreement Detects Unreliable Large Language Model Responses in Medical Question Answering, but Deliberate Prompting Does Not Prevent Them: A Pre-Registered Evaluation of Seven Open-Weight Models

JMIR Preprints. 10/08/2026:109219

DOI: 10.2196/preprints.109219

URL: https://preprints.jmir.org/preprint/109219

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.