Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Feb 24, 2026
Date Accepted: Jun 8, 2026
A Supervised Fine-tuned Large Language Model for Lifestyle Management in Prostate Cancer Patients: Development and Evaluation Study
ABSTRACT
Background:
Lifestyle interventions for patients with prostate cancer (PCa) have been shown to improve treatment adherence and quality of life. However, there remains a lack of large language models (LLMs) capable of delivering individualized and professional lifestyle recommendations under clearly defined medical safety boundaries and controlled evidence sources.
Objective:
This study aimed to develop and evaluate a supervised fine-tuned large language model—PCaPLMM_SFT (Prostate Cancer Patient Lifestyle Management Model via Supervised Fine-Tuning)—to support health literacy improvement and lifestyle self-management among PCa patients.
Methods:
We searched English-language literature primarily from PubMed (February 2015–February 2025) to build a structured lifestyle-management knowledge base covering diet, physical activity, weight management, medication adherence, and psychological support. We used a retrieval-augmented generation (RAG) pipeline to generate patient-style question–answer (QA) pairs from retrieved knowledge slices. Bilingual English–Chinese QA data were generated from English-language source evidence through patient-oriented reformulation and RAG-based answer generation, and independent English and Chinese test sets were constructed to assess bilingual QA performance. We trained Baichuan2-7B-Chat using a two-stage strategy, consisting of continued pre-training followed by SFT with low-rank adaptation (LoRA). Model outputs were evaluated in two double-blind rounds by referee LLMs (Qwen3-Max and DeepSeek-R1) and compared with GPT-3.5-Turbo and the base Baichuan2-7B-Chat using 2,500 queries across five lifestyle scenarios. Additionally, three domain experts conducted a blinded review of 50 QA samples (10 per scenario). We used the Mann–Whitney U test with effect size r and Benjamini–Hochberg false discovery rate correction, and examined consistency using intraclass correlation coefficients (ICCs).
Results:
Based on 2,211 included publications, we constructed the PCaPLMM_SFT-Train dataset. The knowledge base yielded >150,000 structured knowledge slices. After two rounds of review, we obtained 42,330 single-turn QA pairs and 3,008 multi-turn dialogues, and the SFT phase utilized 45,338 structured QA samples. In the dual-round referee LLM assessment, PCaPLMM_SFT consistently outperformed Baichuan2-7B-Chat across dimensions and showed comparable or superior performance to GPT-3.5-Turbo across five lifestyle scenarios. Consistency analyses indicated moderate-to-good agreement between referee models across rounds, supporting the robustness of the comparative evaluation.
Conclusions:
PCaPLMM_SFT demonstrates the feasibility of constructing a medical lifestyle-focused LLM by integrating structured medical knowledge, QA-style training data, and a multi-layer evaluation system. This framework provides a reproducible methodological foundation for evidence-based health education and lifestyle management and establishes groundwork for future evaluation in real-world health management settings.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.