Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR AI

Date Submitted: Sep 17, 2026
Open Peer Review Period: Sep 21, 2026 - Nov 16, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Reliability and construct validity of a multi-model LLM-as-judge evaluator battery for clinician-facing clinical question answering: a controlled adversarial benchmark study

  • Henry Bergman

ABSTRACT

Background:

Open-ended clinical question systems generate fluent free-text for which conventional accuracy metrics are insufficient. Complex questions may admit more than one defensible answer depending on clinical context, jurisdiction, guideline version and the intent of the person asking. Human expert review remains important but is costly, hard to scale and variable. Large language models (LLMs) are increasingly used as automated evaluators, but an LLM-as-judge should be characterised as a measurement instrument before its outputs are treated as evidence.

Objective:

To characterise the inter-judge reliability, within-judge stability, severity-rating heterogeneity, difficulty discrimination and construct validity of an LLM-as-judge battery used to evaluate clinician-facing clinical question answering.

Methods:

We evaluated a three-model FACTFULNESS ensemble (Claude Opus 4.7, Gemini 3.1 Pro, GPT-5.4, each with native web search) and a single-model COMPLETENESS evaluator (GPT-5.4) on an author-constructed adversarial bank of 1,424 clinical questions across five difficulty tiers and 16 construct families. Of 1,421 scored items, 1,173 (82.5%) carried complete three-judge verdicts. Agreement on binary factuality used Gwet AC1, Fleiss kappa and unanimity, overall and within 22 strata. Sensitivity analyses addressed the complete-case restriction, claim-set non-identity, repeated scoring and retrieved-evidence overlap.

Results:

Gwet AC1 was 0.344 (95% CI 0.302-0.386), Fleiss kappa 0.212, and 46.3% of items unanimous. AC1 fell from 0.60 at difficulty tier D1 to 0.27 at D5, ranged from 0.79 to 0.13 across construct families, and ranged from 0.183 to 0.525 between judge pairs with non-overlapping intervals. The serious-or-worse harm-severity rate differed 17.1-fold across judges (0.23%, 2.34% and 3.94%). Requiring two of three judges reduced HIGH+ classification from 13.79% to 2.89%, and no answer met CRITICAL by consensus (0/1,421). FACTFULNESS-target families were more often consensus-inaccurate than non-target families (36.8% vs 27.9%, P=0.002), whereas the COMPLETENESS check was null once refusal-driven security items were excluded (1.7% vs 4.3%, P=0.27). On rescoring byte-identical answers, mean within-judge AC1 was 0.703 against 0.265 between judges, and the consensus risk band was reproduced across all five replicates for 52.0% of items.

Conclusions:

Treated as a measurement instrument, the battery's reliability and severity grading depended strongly on judge identity and on the task being judged, and one of two planned construct checks behaved as intended. Consensus reduced extreme classifications, and aggregate estimates were reproducible enough to support period-on-period monitoring, whereas single-run item-level verdicts were not. Multi-model batteries are therefore usable for monitoring when consensus is reported with judge-specific dispersion, missingness and stratified reliability. Human-anchored calibration remains necessary before severity scale or risk labels can be interpreted as absolute clinical risk. Clinical Trial: NA


 Citation

Please cite as:

Bergman H

Reliability and construct validity of a multi-model LLM-as-judge evaluator battery for clinician-facing clinical question answering: a controlled adversarial benchmark study

JMIR Preprints. 17/09/2026:112207

DOI: 10.2196/preprints.112207

URL: https://preprints.jmir.org/preprint/112207

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.