Currently submitted to: Journal of Medical Internet Research
Date Submitted: Sep 23, 2026
Open Peer Review Period: Sep 23, 2026 - Nov 18, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Sycophancy and Pressure Resistance in Medical Large Language Models: Systematic Review, Taxonomy, and Minimum Evaluation Framework
ABSTRACT
Background:
Medical large language models can perform well on static questions yet alter clinically relevant outputs when users introduce false premises, misleading evidence, authority claims, or repeated disagreement.
Objective:
We systematically reviewed how sycophancy and pressure resistance have been evaluated in patient- and clinician-facing medical settings.
Methods:
The registered review record reported searches of MEDLINE/PubMed, Embase, IEEE Xplore, arXiv, and CENTRAL through 31 May 2026 and yielded a frozen corpus of 31 benchmark and simulation studies. Studies were organized post hoc as direct sycophancy or pressure-resistance evidence (n=16), direct-challenge evidence (n=1), or adjacent interactive-robustness evidence (n=14). Fourteen were archival publications, 1 was accepted, 2 were nonarchival workshop reports, and 14 were preprints as of 4 September 2026. We extracted user role, counterfactual, comparator, analytic unit, denominator, model snapshot, native endpoint, uncertainty, and publication status.
Results:
Direct evaluations documented user-concordant output shifts under pressure in controlled settings. However, construct definitions, comparators, conditioning rules, analytic units, turn structures, model versions, and endpoints were incompatible; no common estimand existed, and pooling was not methodologically defensible. No study estimated routine-care incidence or patient outcomes.
Conclusions:
We derived a taxonomy that distinguishes progressive correction from regressive capitulation and propose a minimum evaluation framework covering construct definition, counterfactual, user role, comparator, model identity, conversation protocol, outcome, denominator, analysis, and openness. The evidence supports pressure resistance as a distinct target for benchmark design and lifecycle regression testing but does not establish a universal failure rate or protection conferred by release year, medical specialization, or a reasoning label.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.