Currently submitted to: Journal of Medical Internet Research
Date Submitted: Aug 18, 2026
Open Peer Review Period: Aug 19, 2026 - Oct 14, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Evaluative Stance Toward Artificial Intelligence in the Medical Literature: Infoveillance Study of 16,749 High-Quartile Journal Abstracts, 2021-2026
ABSTRACT
Background:
Medical artificial intelligence (AI) publications report performance while also framing AI as beneficial, uncertain, or risky, which has not been measured at scale. The literature is itself an information environment clinicians rely on, yet health AI sentiment has been measured mainly in public online discourse, not in medical journals.
Objective:
We aimed to measure the evaluative stance toward AI in high-quartile medical journal abstracts from January 2021 to April 2026, and to characterize it across time, concern themes, failure mechanisms, specialties, first-author country, and publication format.
Methods:
We conducted a large language model (LLM)-assisted computational content analysis of PubMed records from top-two-quartile (Q1/Q2) SCImago journals. Of 97,492 eligible records, an AI-engagement prefilter retained 16,759 discourse or evaluative records, and stance classification yielded 16,749 valid labels on an ordered scale of Alarm, Caution, Neutral, Cautious Optimism, and Advocacy. Critical stance was Alarm plus Caution; favourable stance was Cautious Optimism plus Advocacy. Critical records were further classified for concern theme and, on two axes, for failure mechanism, which distinguished fabrication from factual error, and model type. Every LLM step was validated against blinded human coding (prefilter Cohen kappa 0.51; stance quadratic-weighted kappa 0.79, 95% CI 0.72-0.84; specialty kappa 0.75; model type kappa 0.84; failure mode kappa 0.57).
Results:
Cautious Optimism was the majority stance (62.8%) and 30.8% of records were critical. Advocacy fell monotonically from 2.9% (2021) to 0.6% (partial 2026; Spearman rho -1.00), while critical share rose modestly from 25.4% (95% CI 22.9-28.1) to 32.6% (30.9-34.4), rising within both prefilter genres. Among critical abstracts, patient safety rose 11.0 percentage points and errors 30.9, whereas regulation fell 22.0 in prevalence share. Within the hallucination-gated subset (n=2226), factual error was the majority mechanism every year (56% to 90%), while fabrication-involved papers reached about 21% to 23% from 2023 onward; fabrication estimates are an upper bound. Critical rate varied roughly threefold across 18 specialty and domain categories (51.3% in Mental Health and Psychiatry to 16.8% in Cardiology), was inversely associated with US Food and Drug Administration cleared-device availability (rho -0.65, two-sided p=.004), and varied by first-author country (34.5% United States, 16.2% China). Reviews were the least critical (24.9%) and most favourable (73.1%) format.
Conclusions:
Medical-AI abstracts moved away from unqualified promotion toward more qualified assessment, not toward broad opposition. Two implications follow for how the literature is written and read. That reviews were the least critical and most favourable format may matter for readers, guideline developers, and educators. Critical reporting should distinguish factual error from fabrication rather than treating hallucination as one category, which bears on how journals, reviewers, and reporting standards describe language model failure. Throughout our investigation, we measure published discourse, not AI capability, nor whether these stances are correct.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.