Two-Stage Extraction of Clinical Course Information From Psychiatric Discharge Summaries Using Fine-Tuned Large Language Models: Cross-Validation Study.

Clinical course information, including disease onset, episode recurrence, and hospitalization history, is essential for psychiatric care and research. However, these data are embedded in unstructured clinical narratives with substantial linguistic variability, making manual extraction labor-intensive and rule-based extraction difficult to scale. Fine-tuned large language models (LLMs) may flexibly extract such information from privacy-sensitive psychiatric records.

This study aimed to evaluate the performance of LLMs for automatically extracting temporal and clinical course information from psychiatric discharge summaries.

We analyzed 500 psychiatric discharge summaries from the Integrated Medical Database of National Taiwan University Hospital. A psychiatrist and a natural language processing researcher manually annotated clinical events and temporal information. Four open-source LLMs (LLaMA, MentaLLaMA, OpenBioLLM, and Mistral) were fine-tuned using low-rank adaptation and evaluated using 10-fold cross-validation. A 2-stage framework was developed in which sentence-level extraction of clinical events and temporal information was followed by chart-level prediction of 4 clinical course features: first-episode onset time, episode count, number of psychiatric hospitalizations, and most recent hospitalization. Performance was assessed using precision, recall, F1-score, accuracy, mean absolute error (MAE), and bootstrap significance testing.

A total of 12,947 sentences were extracted from 500 discharge summaries, yielding 7177 clinical event annotations and 4842 temporal annotations. At the sentence level, Mistral achieved the highest F1-scores for clinical event extraction, including symptom/episode detection (0.925, 95% CI 0.920-0.930), hospitalization detection (0.944, 95% CI 0.934-0.952), and remission/response detection (0.867, 95% CI 0.851-0.883). Mistral also achieved the highest F1-scores for most temporal information categories, including age expressions (0.983, 95% CI 0.973-0.993), relative time expressions (0.953, 95% CI 0.944-0.963), duration expressions (0.976, 95% CI 0.962-0.987), and vague temporal expressions (0.888, 95% CI 0.871-0.905). At the chart level, the proposed 2-stage framework showed the clearest benefit for first-episode onset prediction. Using the 2-stage framework, Mistral achieved the highest point estimate for onset accuracy (0.772), although its performance did not significantly differ from that of LLaMA using the same framework (0.744; bootstrap P=.26). For onset prediction, the 2-stage framework significantly outperformed the direct and joint extraction approaches across all evaluated models (all bootstrap P≤.002). For other chart-level features, the best-performing approach varied: using the 2-stage framework, Mistral achieved the highest episode count accuracy (0.624; MAE=0.428); using the direct approach, LLaMA achieved the highest hospitalization count accuracy (0.692; MAE=0.379); and using the 2-stage framework, OpenBioLLM achieved the highest most recent hospitalization accuracy (0.868; F1-score=0.617).

Fine-tuned, locally deployable, open-source LLMs can extract temporal and longitudinal disease course information from psychiatric discharge summaries. The 2-stage framework was most beneficial for first-episode onset prediction and performed competitively across other chart-level features, supporting the use of LLMs to transform heterogeneous psychiatric narratives into structured data for research and decision support.
Mental Health
Access
Care/Management
Advocacy

Authors

Chen Chen, Tseng Tseng, Dai Dai, Su Su, Wang Wang, Chien Chien, Huang Huang, Wu Wu, Chen Chen
View on Pubmed
Share
Facebook
X (Twitter)
Bluesky
Linkedin
Copy to clipboard