Currently submitted to: Journal of Medical Internet Research
Date Submitted: Aug 3, 2026
Open Peer Review Period: Aug 4, 2026 - Sep 29, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Beyond Sycophancy: Vendor- and Interface-Dependent Revision Recommendations in Large Language Model Methods Review
ABSTRACT
Background:
Large language models (LLMs) have been demonstrated to exhibit sycophantic behavior, agreeing with users rather than corrected them, which may reinforce user belief over accuracy. Whether this compromises LLM appraisal of flawed research methods is unknown.
Objective:
This study evaluated whether LLMs endorse methodologically flawed clinical research excerpts as sound, whether they correctly pass sound excerpts, and whether stated user confidence or authority alters these behaviors. Sycophantic false reassurance was hypothesized to occur and to increase with user confidence.
Methods:
Twenty clinical research methods excerpts were developed: ten containing a single prespecified major methodological flaw and ten matched flaw-corrected twins. Excerpts were submitted to two commercial LLM vendors (OpenAI, Anthropic) as independent structured review requests. The primary outcome was false reassurance, defined as endorsing a flawed excerpt as sound. The secondary outcome was clean overflagging, defined as recommending major revision of a flaw-corrected excerpt. The application program interface (API) experiment crossed 20 excerpts by 2 vendors, 7 prompt framings, 3 stated expertise levels, and 5 replicates (4200 responses). A targeted 200-response substudy was collected through the consumer ChatGPT and Claude interfaces with expertise fixed at intermediate.
Results:
False reassurance was rare (3/2100 flawed responses, 0.14%; 0/1050 Anthropic, 3/1050 OpenAI, 0.29%) and did not occur in the commercial product substudy (0/40). Clean overflagging occurred in 14.5% (152/1050) of Anthropic and 48.8% (509/1044) of OpenAI clean responses (OR 5.70, 95% CI: 2.92 – 11.10). This difference persisted after sequential exclusion of the most frequently flagged excerpts. Stated expertise was not associated with either outcome. The overall prompt-framing effect was not significant (p = .26), though authority-pressure (OR 1.81, 95% CI: 1.10 – 2.97, p = .02) and high-user-confidence (OR 2.74, 95% CI: 1.05 – 7.14, p = .04) were associated with higher odds of clean overflagging than neutral framing. Clean overflagging was substantially higher through consumer products than the API for both vendors.
Conclusions:
Sycophantic false reassurance was essentially absent and asserting user confidence or authority did not induce it. The dominant behavior was instead recommending revision of methodologically sound excerpts, which varied markedly by vendor and by whether the model was accessed through a consumer product or an API. Given identical excerpts produced different verdicts across products and surfaces, model-generated revision recommends should prompt human evaluation of the stated concern rather than automatic acceptance that a flaw exists. Clinical Trial: NA
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.