Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Apr 10, 2026
Date Accepted: Jul 2, 2026
Performance of Large Language Models for Oncology Nursing Decision Support: Cross-Sectional Study
ABSTRACT
Background:
Background:
Large language models (LLMs) are increasingly used in healthcare, with emerging applications in clinical decision support and nursing education. However, evidence on their performance in nursing contexts, particularly in oncology nursing, remains limited. Given the complexity and high-risk nature of oncology care, it is important to evaluate the performance and clinical relevance of LLM-generated responses in oncology nursing contexts.
Objective:
Aim: To compare the performance of LLMs in oncology nursing decision-support tasks using standardized examination questions and case-based clinical scenarios, and to explore their potential applicability and current limitations in oncology nursing practice.
Methods:
Methods:
A total of 33 case-based questions derived from 10 oncology nursing clinical scenarios in a nationally used training manual, along with standardized examination-oriented questions from a commercially published preparation book for the Chinese Intermediate Nursing Professional Qualification Examination, were used to evaluate the performance of five LLMs (DeepSeek, Qwen, SparkDesk, WiseDiag, and ChatGPT). All models generated responses using a standardized prompt. Two oncology nurses with more than five years of clinical experience independently rated the case-based responses using three evaluation dimensions: correctness, clarity, and conciseness (3C). Inter-rater reliability was assessed using quadratic weighted Cohen’s kappa, intraclass correlation coefficients, and Spearman’s rank correlation coefficient. Differences among models were analyzed using the Kruskal–Wallis test with Dunn’s post hoc test. In addition, examination performance was evaluated based on total score, accuracy rate, and completion efficiency.
Results:
Results:
Inter-rater reliability analyses indicated moderate agreement between evaluators. The median 3C scores were as follows: DeepSeek 11.50 (10.50, 12.00), Qwen 11.00 (10.50, 12.00), SparkDesk 10.50 (9.50, 11.50), WiseDiag 10.00 (9.50, 11.50), and ChatGPT 10.00 (9.00, 11.50). The Kruskal–Wallis test indicated statistically significant differences among models (H = 11.416, P < 0.05), with post hoc analysis showing a significant difference only between DeepSeek and ChatGPT (P < 0.05). In examination-based tasks, all models achieved passing performance, with accuracy rates ranging from 77% to 93%. In terms of response completion, DeepSeek and ChatGPT completed all tasks in a single interaction, whereas other models required multiple interactions due to output interruptions.
Conclusions:
Conclusions:
LLMs demonstrate heterogeneous performance in oncology nursing tasks, with relatively stronger performance on structured knowledge assessments and standardized examination tasks. Their potential application appears most relevant in supporting information retrieval, knowledge organization, and patient education tasks. However, their performance remains limited in complex and high-risk clinical scenarios requiring individualized assessment and dynamic clinical judgment. Current evidence supports the use of LLMs as supportive tools within oncology nursing practice, with outputs requiring interpretation alongside professional clinical judgment.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.