Previously submitted to: JMIR Cancer (no longer under consideration since Aug 14, 2026)
Date Submitted: Jul 15, 2026
Open Peer Review Period: Jul 17, 2026 - Aug 14, 2026
(closed for review but you can still tweet)
NOTE: This is an unreviewed Preprint
Warning: This is a unreviewed preprint (What is a preprint?). Readers are warned that the document has not been peer-reviewed by expert/patient reviewers or an academic editor, may contain misleading claims, and is likely to undergo changes before final publication, if accepted, or may have been rejected/withdrawn (a note "no longer under consideration" will appear above).
Peer review me: Readers with interest and expertise are encouraged to sign up as peer-reviewer, if the paper is within an open peer-review period (in this case, a "Peer Review Me" button to sign up as reviewer is displayed above). All preprints currently open for review are listed here. Outside of the formal open peer-review period we encourage you to tweet about the preprint.
Citation: Please cite this preprint only for review purposes or for grant applications and CVs (if you are the author).
Final version: If our system detects a final peer-reviewed "version of record" (VoR) published in any journal, a link to that VoR will appear below. Readers are then encourage to cite the VoR instead of this preprint.
Settings: If you are the author, you can login and change the preprint display settings, but the preprint URL/DOI is supposed to be stable and citable, so it should not be removed once posted.
Submit: To post your own preprint, simply submit to any JMIR journal, and choose the appropriate settings to expose your submitted version as preprint.
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
The Decision-Support Value of Large Language Models in Rectal Cancer Diagnosis and Treatment: A Multidimensional Comparative Study Across Diverse Clinical Scenarios
ABSTRACT
Background:
The integration of artificial intelligence into oncology promises to augment clinical decision-making, yet the practical utility of LLMs in complex subspecialties like rectal cancer is not well-defined. Current benchmarks often overlook the critical differences between general and medical-optimized models in handling intricate treatment algorithms involving staging and immunotherapy biomarkers.
Objective:
To assess the diagnostic and therapeutic decision-making capabilities of seven LLMs using a cohort of simulated rectal cancer cases, and to delineate the specific limitations and contextual dependencies that must be addressed prior to clinical deployment.
Methods:
We developed 60 virtual cases covering various stages and molecular profiles of rectal cancer. Seven LLMs (five general, two medical) were assessed. Their generated treatment plans were blindly reviewed by senior oncologists and scored for staging accuracy, treatment appropriateness, and execution details.
Results:
Inter-rater reliability was excellent (Kappa coefficient range 0.68–0.84), ensuring assessment robustness. Analysis revealed that general-purpose models like GPT-5.2 and Gemini-3.1 significantly outperformed specialized medical-enhanced models in staging accuracy tasks. For treatment decision appropriateness, DeepSeek-V3.2 and Gemini-3 demonstrated the best and most stable performance, while GPT-5.2 led in the standardization of treatment execution. Subgroup analysis uncovered critical limitations: all models faced significant challenges in formulating treatment strategies for locally advanced cases (especially stage III C), and their performance was inconsistent regarding immunotherapy decisions related to dMMR subtypes, highlighting a pronounced "uneven proficiency" and high context-dependency of model capabilities.
Conclusions:
Leading general-purpose LLMs show potential for certain aspects of rectal cancer diagnosis and treatment, though their performance varies across tasks, disease stages, and molecular subtypes. Future work should focus on optimizing and validating these models for complex clinical scenarios. Integration into practice must proceed cautiously to ensure safe, effective, and equitable AI-assisted decision-making.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.