Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: Journal of Medical Internet Research

Date Submitted: Aug 18, 2026
Open Peer Review Period: Aug 19, 2026 - Oct 14, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

CausalNHANES: An Automated Benchmark for Evaluating Large Language Models' Causal Reasoning in Epidemiology

  • Jing Gao; 
  • Dong Lin; 
  • Jianfeng Hu

ABSTRACT

Background:

Evaluating Large Language Models (LLMs) on epidemiological causal inference tasks traditionally relies on expensive human expert scoring, which suffers from high cost, poor reproducibility, and scalability bottlenecks. Automated benchmarks are urgently needed to rigorously assess whether LLMs truly understand causal inference principles — the cornerstone of epidemiological methodology.

Objective:

We aimed to develop and validate CausalNHANES, a fully automated benchmark that evaluates LLMs' causal reasoning capabilities using the National Health and Nutrition Examination Survey (NHANES) dataset as the empirical substrate, without requiring human expert scoring.

Methods:

We constructed a 3+1 evaluation framework (three core causal reasoning tasks plus one safety gate) comprising (1) directed acyclic graph construction, (2) confounder identification, (3) causal inference from descriptive statistics, and (4) counterfactual reasoning. Evaluation relied on two objective computational anchors: literature consensus anchoring via PubMed-based retrieval and logical rigidity anchoring via formal causal logic rule verification. Cross-model consensus was computed post hoc as a supplementary qualitative probe for error-pattern clustering, not as a quantitative scoring component. We systematically evaluated 9 LLMs across 80 test scenarios (60 primary tasks plus 20 negative-controls). The confounder identification layer operated as a binary pass/fail sanity check and was excluded from the composite score. Innovation was redefined as strategy finesse multiplied by a logic-rigidity admission term (squared penalty for violations) and negative-control pass rate. The negative-control scenarios served as a prerequisite disqualification filter: any model hallucinating a causal relationship in a negative-control scenario was excluded from the main leaderboards and listed separately in a disqualified cohort. Because every evaluated model failed at least one negative-control, no model qualified for the ordinal tier leaderboards. This universal failure became the primary analytical focus.

Results:

The formal experiment evaluated 9 state-of-the-art LLMs across 720 scenario-model pairs (180 negative-control scenarios). Strikingly, all 9 models failed at least one negative-control scenario, yielding a zero-percent qualification rate for the main leaderboards. No model achieved Platinum tier (negative-control pass plus zero logical violations). ERNIE-4.5 and Doubao approached Gold-tier eligibility (negative-control pass with low violation rates) but were ultimately disqualified by hallucinated causal claims in negative-control tasks. Kimi-K2.6 exhibited the highest logical violation rate (49.0%). The pattern was consistent across model families and geographic origins: Chinese-origin and Western-origin models were equally susceptible to hallucinating causal relationships where none exist. Notably, Western flagship models (GPT-4o, Claude-Sonnet-5, Gemini-3.6-Flash) also failed negative-controls, confirming that the vulnerability is architectural rather than geographically or commercially bounded.

Conclusions:

CausalNHANES reveals a sobering finding: despite rapid advances in general reasoning, current state-of-the-art LLMs are not yet reliable for epidemiological causal inference. The universal failure on negative-controls—a task requiring only the ability to state 'no causal relationship'—suggests that models lack genuine causal reasoning capability and instead rely on pattern-matching heuristics that overfit to epidemiological vocabulary. This failure-mode analysis, rather than a performance ranking, constitutes the primary scientific contribution. The framework itself remains fully reproducible, cost-efficient, and scalable, providing a methodological paradigm for safety-first AI evaluation in high-stakes epidemiological and digital health domains.


 Citation

Please cite as:

Gao J, Lin D, Hu J

CausalNHANES: An Automated Benchmark for Evaluating Large Language Models' Causal Reasoning in Epidemiology

JMIR Preprints. 18/08/2026:109941

DOI: 10.2196/preprints.109941

URL: https://preprints.jmir.org/preprint/109941

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.