Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Apr 13, 2026
Date Accepted: Sep 9, 2026
Safety-Oriented Benchmarking of Large Language Models in Initial ASCCP Risk-Based Management of Abnormal Cervical Screening Results: Scenario-Based Benchmark Study
ABSTRACT
Background:
Large language models (LLMs) are increasingly being considered for clinical decision support, yet their safety in risk-based cervical screening management remains insufficiently characterized.
Objective:
This study evaluated the guideline concordance and clinical safety of three LLMs in the initial ASCCP risk-based management of abnormal cervical screening results.
Methods:
Sixty synthetic clinical scenarios reflecting initial abnormal screening management in immunocompetent, non-pregnant women aged 25–65 years were developed using a predefined scenario coverage matrix. GPT-5.3, Gemini 3 Flash, and DeepSeek V3.2 were tested under two prompt conditions: Baseline Clinical Prompt and Structured Guideline-Directed Prompt. Each scenario was run in three independent repetitions per model and prompt arm (1,080 total observations). Responses were evaluated by two blinded obstetrics and gynecology specialists using a prespecified rubric. The primary endpoint was the unsafe major error-free rate.
Results:
Under structured prompting, the unsafe major error-free rate was 100.0% for GPT-5.3 (95% CI, 97.9–100.0), 98.9% for Gemini 3 Flash (95% CI, 96.0–99.7), and 75.0% for DeepSeek V3.2 (95% CI, 68.2–80.8). Structured prompting independently improved both safety (OR = 3.76, p < 0.001) and exact concordance (OR = 8.27, p < 0.001). Error rates increased substantially with scenario complexity, rising from 3.1% in low-complexity to 29.0% in high-complexity scenarios. The most frequent error subtypes were under-management, genotype misinterpretation, and history neglect. Inter-rater agreement was substantial to excellent (weighted κ = 0.839).
Conclusions:
LLM safety in initial ASCCP risk-based management varies markedly by model, prompt strategy, and scenario complexity. Structured prompting significantly improves both safety and guideline concordance, but even the best-performing model remained vulnerable in complex, history-dependent scenarios. LLMs may have value as clinician-supervised decision support tools but are not yet suitable for autonomous clinical use in cervical screening management.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.