Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR AI

Date Submitted: Aug 29, 2026
Open Peer Review Period: Sep 4, 2026 - Oct 30, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Not Whether to Operate but How Far? Seven Open-Source Large Language Models Compared With Preoperative Surgical Conference Decisions in Colorectal Cancer: Comparative Evaluation Study

  • Takeshi Yanagita; 
  • Tsuyoshi Saito; 
  • Shinichiro Tsuru; 
  • Ryoji Umeki; 
  • Naoki Sawamura; 
  • Shuji Kurata; 
  • Futoshi Teranishi

ABSTRACT

Background:

Large language models (LLMs) are increasingly proposed as clinical decision-support tools. Most evaluations have used proprietary models reached through a commercial interface, which is hard to reconcile with the handling of protected health information. Open-source models can run inside the hospital, but they have rarely been compared under an identical protocol, and the decision on how far to extend an operation has rarely been examined.

Objective:

To evaluate how closely seven open-source LLMs reproduce the decisions of preoperative surgical conferences on whether a guideline-recommended procedure should be modified for the individual patient, and whether an extended-reasoning mode or a prompt reproducing surgical reasoning improves that agreement.

Methods:

In total, 109 consecutive patients with colorectal cancer from 2 hospitals were included. A standardized English vignette was prepared from the information available when the treatment plan was decided. The reference standard was the decision approved at the preoperative conference, classified as appropriate or as 1 of 4 modification types, and derived by a single rule from routinely recorded operative data. Seven models were served locally with Ollama. A baseline prompt was compared with a structured prompt containing a three-gate framework derived from surgical reasoning and fixed before the runs. The reasoning mode was set through the think field, and the chain-of-thought length was recorded for every response. The primary metric was Cohen kappa with a bootstrap 95% CI.

Results:

The conference modified the recommended procedure in 44 of 109 patients (40.4%). With the baseline prompt, kappa ranged from 0.243 to 0.501 and accuracy from 64.2% to 76.1%. The highest kappa was obtained with Qwen3.6 in the think mode (95% CI 0.328-0.665), and 20 of the 28 runs stayed in the fair range or below. The models answered "appropriate" more often than the reference standard in 9 of 14 runs. Within the 44 modified cases, the correct modification type was given in 13.6% to 31.8%. The reasoning mode could be controlled in only 2 of the 7 models, and the text directives commonly used for this purpose worked in none. Where it could be controlled, extended reasoning did not improve agreement, and the best result of Gemma 4 came without any chain of thought. The structured prompt raised kappa in the 3 weakest models and lowered it in the 2 strongest. GPT-OSS returned a different answer in 20 of 109 cases between 2 runs of the same fixed setting.

Conclusions:

Locally deployed open-source LLMs agreed only moderately at best with preoperative surgical conference decisions, and about one quarter of the patients were judged differently even in the best condition. Neither extended reasoning nor a surgeon-derived structured prompt improved agreement. These models are not yet suitable for supporting the decision on the extent of surgery in colorectal cancer.


 Citation

Please cite as:

Yanagita T, Saito T, Tsuru S, Umeki R, Sawamura N, Kurata S, Teranishi F

Not Whether to Operate but How Far? Seven Open-Source Large Language Models Compared With Preoperative Surgical Conference Decisions in Colorectal Cancer: Comparative Evaluation Study

JMIR Preprints. 29/08/2026:110741

DOI: 10.2196/preprints.110741

URL: https://preprints.jmir.org/preprint/110741

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.