A Supervised Fine-Tuned Large Language Model for Lifestyle Management in Patients With Prostate Cancer: Development and Evaluation Study.
Lifestyle interventions for patients with prostate cancer have been shown to improve treatment adherence and quality of life. However, there remains a lack of large language models (LLMs) capable of delivering individualized and professional lifestyle recommendations under clearly defined medical safety boundaries and controlled evidence sources.
This study aimed to develop and evaluate a supervised fine-tuned LLM-PCaPLMM_SFT (Prostate Cancer Patient Lifestyle Management Model via Supervised Fine-Tuning)-to support health literacy improvement and lifestyle self-management among patients with prostate cancer.
We searched English-language literature primarily from PubMed (February 2015 to February 2025) to build a structured lifestyle management knowledge base covering diet, physical activity, weight management, medication adherence, and psychological support. We used a retrieval-augmented generation pipeline to generate patient-style question-answer (QA) pairs from retrieved knowledge slices. Bilingual English-Chinese QA data were generated from English-language source evidence through patient-oriented reformulation and retrieval-augmented generation-based answer generation, and independent English and Chinese test sets were constructed to assess bilingual QA performance. We trained Baichuan2-7B-Chat using a 2-stage strategy, consisting of continued pretraining, followed by supervised fine-tuning with low-rank adaptation. Model outputs were evaluated in 2 double-blind rounds by referee LLMs (Qwen3-Max and DeepSeek-R1) and compared with GPT-3.5-Turbo and the base Baichuan2-7B-Chat using 2500 queries across 5 lifestyle scenarios. Additionally, 3 domain experts conducted a blinded review of 50 QA samples (10 per scenario). We used the Mann-Whitney U test with effect size r, and Benjamini-Hochberg false discovery rate correction, and examined consistency using intraclass correlation coefficients.
Based on 2211 included publications, we constructed the PCaPLMM_SFT-Train dataset. The knowledge base yielded >150,000 structured knowledge slices. After 2 rounds of review, we obtained 42,330 single-turn QA pairs and 3008 multiturn dialogues, and the supervised fine-tuning phase used 45,338 structured QA samples. In the dual-round referee LLM assessment, PCaPLMM_SFT consistently outperformed Baichuan2-7B-Chat across dimensions and showed comparable or superior performance to GPT-3.5-Turbo across 5 lifestyle scenarios. Consistency analyses indicated moderate to good agreement between referee models across rounds, supporting the robustness of the comparative evaluation.
PCaPLMM_SFT demonstrates the feasibility of constructing a medical lifestyle-focused LLM by integrating structured medical knowledge, QA-style training data, and a multilayer evaluation system. This framework provides a reproducible methodological foundation for evidence-based health education and lifestyle management and establishes groundwork for future evaluation in real-world health management settings.
This study aimed to develop and evaluate a supervised fine-tuned LLM-PCaPLMM_SFT (Prostate Cancer Patient Lifestyle Management Model via Supervised Fine-Tuning)-to support health literacy improvement and lifestyle self-management among patients with prostate cancer.
We searched English-language literature primarily from PubMed (February 2015 to February 2025) to build a structured lifestyle management knowledge base covering diet, physical activity, weight management, medication adherence, and psychological support. We used a retrieval-augmented generation pipeline to generate patient-style question-answer (QA) pairs from retrieved knowledge slices. Bilingual English-Chinese QA data were generated from English-language source evidence through patient-oriented reformulation and retrieval-augmented generation-based answer generation, and independent English and Chinese test sets were constructed to assess bilingual QA performance. We trained Baichuan2-7B-Chat using a 2-stage strategy, consisting of continued pretraining, followed by supervised fine-tuning with low-rank adaptation. Model outputs were evaluated in 2 double-blind rounds by referee LLMs (Qwen3-Max and DeepSeek-R1) and compared with GPT-3.5-Turbo and the base Baichuan2-7B-Chat using 2500 queries across 5 lifestyle scenarios. Additionally, 3 domain experts conducted a blinded review of 50 QA samples (10 per scenario). We used the Mann-Whitney U test with effect size r, and Benjamini-Hochberg false discovery rate correction, and examined consistency using intraclass correlation coefficients.
Based on 2211 included publications, we constructed the PCaPLMM_SFT-Train dataset. The knowledge base yielded >150,000 structured knowledge slices. After 2 rounds of review, we obtained 42,330 single-turn QA pairs and 3008 multiturn dialogues, and the supervised fine-tuning phase used 45,338 structured QA samples. In the dual-round referee LLM assessment, PCaPLMM_SFT consistently outperformed Baichuan2-7B-Chat across dimensions and showed comparable or superior performance to GPT-3.5-Turbo across 5 lifestyle scenarios. Consistency analyses indicated moderate to good agreement between referee models across rounds, supporting the robustness of the comparative evaluation.
PCaPLMM_SFT demonstrates the feasibility of constructing a medical lifestyle-focused LLM by integrating structured medical knowledge, QA-style training data, and a multilayer evaluation system. This framework provides a reproducible methodological foundation for evidence-based health education and lifestyle management and establishes groundwork for future evaluation in real-world health management settings.
Authors
Jiang Jiang, Yang Yang, Zheng Zheng, Yin Yin, Zhang Zhang, Lan Lan, Wu Wu, Lin Lin, Jiang Jiang, Chen Chen
View on Pubmed