Enhancing Psychiatry Training Using an Agentic AI Simulated Consultation Tool: Prospective Cohort Study.
Canadian psychiatry residents must demonstrate consultation competency, assessed using the standardized assessment of a clinical encounter report (STACER). However, opportunities to practice these skills and receive constructive assessment remain limited in clinical settings.
This study aimed to evaluate the technical feasibility of an agentic AI system designed to support psychiatry residents' consultation competence through simulated patient encounters with a patient agent and structured feedback from a rater agent.
We conducted a two-phase technical feasibility prospective single-arm cohort study of the STACER Agentic System, a large language model-based platform integrating a patient agent and a rater agent. Phase 1 involved automated evaluation of the patient agent using a psychiatrist agent across 227 synthetic major depressive disorder cases. Performance was assessed using DeepEval metrics (correctness, clarity, medical faithfulness, turn relevance, and role adherence) with descriptive statistics and 95% CIs. Phase 2 involved a preliminary user study with 14 convenience-sampled participants: a total of 5 members of the clinical research team and 9 psychiatry residents from the University of Alberta. Participants completed simulated diagnostic interviews and case presentations. Performance was evaluated using STACER-based scoring by the rater agent and 2 psychiatrists. Interrater reliability was assessed using intraclass correlation coefficients (α=.05). Participants rated realism, behavioral consistency, psychiatric nuance, and feedback utility using Likert scales and free-text answers.
The patient agent demonstrated high behavioral (51/56, 91.07%) and symptom fidelity (105/110, 95.45%), with strong automated performance (medical faithfulness mean 0.99, 95% CI 0.99-1.00; turn relevance 0.99, 95% CI 0.986-0.992). Participants rated simulations as psychiatrically plausible and diagnostically useful, particularly for depressive symptom representation, although rapport building was moderate (mean 2.78, SD 1.56 to mean 3.00, SD 1.41, out of 5.00) due to limited nonverbal cues. The rater agent generated structured STACER-aligned feedback with high intrarater consistency, especially at the section subtotal level. Interrater reliability with psychiatrists was poor at the item level (intraclass correlation coefficient range=0.25-0.49) but improved to good-to-excellent agreement at the section level for psychiatry resident sessions (intraclass correlation coefficient range=0.89-0.93). The rater agent's scores fell between those of the 2 psychiatrists for the clinical research team and were lower than both human raters for psychiatry residents.
The STACER Agentic System demonstrates the technical feasibility of using agentic AI to simulate psychiatric consultations and deliver STACER-aligned formative feedback. By combining adaptive multiturn psychiatric simulation with competency-based evaluation, it shows promise in supporting cognitive aspects of consultation, though it remains limited in facilitating relational skills such as rapport building. These findings suggest agentic AI could expand scalable, low-risk opportunities for deliberate practice and formative feedback in competency-based psychiatric education. Further controlled studies are needed to evaluate educational effectiveness and integration into residency training.
This study aimed to evaluate the technical feasibility of an agentic AI system designed to support psychiatry residents' consultation competence through simulated patient encounters with a patient agent and structured feedback from a rater agent.
We conducted a two-phase technical feasibility prospective single-arm cohort study of the STACER Agentic System, a large language model-based platform integrating a patient agent and a rater agent. Phase 1 involved automated evaluation of the patient agent using a psychiatrist agent across 227 synthetic major depressive disorder cases. Performance was assessed using DeepEval metrics (correctness, clarity, medical faithfulness, turn relevance, and role adherence) with descriptive statistics and 95% CIs. Phase 2 involved a preliminary user study with 14 convenience-sampled participants: a total of 5 members of the clinical research team and 9 psychiatry residents from the University of Alberta. Participants completed simulated diagnostic interviews and case presentations. Performance was evaluated using STACER-based scoring by the rater agent and 2 psychiatrists. Interrater reliability was assessed using intraclass correlation coefficients (α=.05). Participants rated realism, behavioral consistency, psychiatric nuance, and feedback utility using Likert scales and free-text answers.
The patient agent demonstrated high behavioral (51/56, 91.07%) and symptom fidelity (105/110, 95.45%), with strong automated performance (medical faithfulness mean 0.99, 95% CI 0.99-1.00; turn relevance 0.99, 95% CI 0.986-0.992). Participants rated simulations as psychiatrically plausible and diagnostically useful, particularly for depressive symptom representation, although rapport building was moderate (mean 2.78, SD 1.56 to mean 3.00, SD 1.41, out of 5.00) due to limited nonverbal cues. The rater agent generated structured STACER-aligned feedback with high intrarater consistency, especially at the section subtotal level. Interrater reliability with psychiatrists was poor at the item level (intraclass correlation coefficient range=0.25-0.49) but improved to good-to-excellent agreement at the section level for psychiatry resident sessions (intraclass correlation coefficient range=0.89-0.93). The rater agent's scores fell between those of the 2 psychiatrists for the clinical research team and were lower than both human raters for psychiatry residents.
The STACER Agentic System demonstrates the technical feasibility of using agentic AI to simulate psychiatric consultations and deliver STACER-aligned formative feedback. By combining adaptive multiturn psychiatric simulation with competency-based evaluation, it shows promise in supporting cognitive aspects of consultation, though it remains limited in facilitating relational skills such as rapport building. These findings suggest agentic AI could expand scalable, low-risk opportunities for deliberate practice and formative feedback in competency-based psychiatric education. Further controlled studies are needed to evaluate educational effectiveness and integration into residency training.
Authors
Rueda Rueda, Al-Shamali Al-Shamali, Cote Cote, Roy Roy, Janssen-Aguilar Janssen-Aguilar, Joseph Joseph, Teferra Teferra, Kamaleddin Kamaleddin, Burback Burback, Winkler Winkler, Kapralos Kapralos, Torres Torres, Sharma Sharma, Krishnan Krishnan, Dumas Dumas, Greenshaw Greenshaw, Hudon Hudon, Sockalingam Sockalingam, Zhang Zhang, Dubrowski Dubrowski, Bhat Bhat
View on Pubmed