Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training
作者: Luming Yang, Haoxian Liu, Siqing Li, Rong Jia, Yue Xiao, Guanhua Chen, Li Lu
分类: cs.MA, cs.AI, cs.HC
发布日期: 2026-09-10
💡 一句话要点
提出多代理大语言模型系统以提升临床访谈训练效果
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 临床教育 大语言模型 多代理系统 标准化病人 医学培训 人工智能 临床访谈 教育技术
📋 核心要点
- 现有的标准化病人训练方法资源消耗大,难以大规模推广,限制了临床教育的有效性。
- 本文提出了一种多代理大语言模型系统,通过模拟病人对话和提供反馈,提升医学生的访谈技能。
- 实验结果显示,多代理系统在沟通和同理心表达方面显著提高了学生的考试成绩,验证了其有效性。
📝 摘要(中文)
临床教育需要培养医学生在不确定条件下进行安全且连贯的病人访谈。传统的标准化病人训练资源消耗大且难以扩展。本文开发了一种以支架为导向的多代理大语言模型(LLM)AI标准化病人(AI-SP)训练平台。该系统包含用于模拟对话的病人代理、提供苏格拉底式提示的导师代理,以及监控临床进展的逐轮评估代理。随机对照研究显示,尽管最终诊断准确性无显著差异,但多代理系统在沟通、同理心表达和历史采集行为方面显著提升了考试成绩。研究结果表明,专门的LLM代理能够提高模拟临床访谈的质量,而不会人为提高考试结果。为支持未来研究,本文发布了多专家注释的数据集,旨在促进基于教育理论的AI-SP系统的发展。
🔬 方法详解
问题定义:本文旨在解决传统标准化病人训练方法资源消耗大、难以扩展的问题,导致医学生在临床访谈技能培养上的不足。
核心思路:提出了一种多代理大语言模型系统,通过模拟病人、导师和评估者的角色,提供动态反馈和支持,帮助学生在不确定环境下进行有效的临床访谈。
技术框架:系统由三个主要模块组成:病人代理用于模拟对话,导师代理提供引导性提示,评估代理监控学生的临床进展。整个流程通过多轮对话进行,确保学生在每个阶段都能获得及时反馈。
关键创新:该系统的创新在于结合了多代理的协作机制,能够在不泄露诊断信息的情况下,通过逐轮反馈提升学生的访谈能力,与传统方法相比,提供了更灵活和个性化的学习体验。
关键设计:系统设计中采用了逐轮评估机制,确保每次对话都能针对学生的表现进行反馈,使用了基于标准化临床考试(OSCE)的评分标准来评估学生的表现。
🖼️ 关键图片
📊 实验亮点
实验结果显示,使用多代理AI标准化病人系统的学生在最终考试中的表现显著优于对照组,尤其在沟通、同理心表达和历史采集行为方面,提升幅度明显,验证了该系统的有效性。
🎯 应用场景
该研究的潜在应用领域包括医学教育、临床技能培训和人工智能辅助的教育工具。通过提升医学生的临床访谈能力,该系统可以在医疗培训中发挥重要作用,促进更高质量的患者护理和沟通。未来,该平台可能扩展至其他医学领域或专业技能培训中。
📄 摘要(原文)
Clinical education must prepare medical students to conduct safe and coherent patient interviews under conditions of uncertainty. Traditional standardized patient (SP) training is resource-intensive and difficult to scale. We developed a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) training platform1. The system includes a patient agent for simulated dialog, a tutor agent providing Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. In a randomized controlled study (N = 100 medical students), participants were assigned to either a multi-agent (MA) scaffolding condition or a control condition. All students completed two learning sessions under their assigned condition followed by an examination conducted in a patient only environment. Performance was assessed using a standardized Objective Structured Clinical Examination (OSCE) based rubric. While no significant difference was observed in final diagnostic accuracy between groups, the multi-agent AI standardized patient system improved final examination scores compared to the control group utilizing structured progressive information disclosure; the most substantial and consistent improvements were observed in communication, the expression of empathy, and specific history-taking behaviors. These findings suggest that specialized LLM agents enhance the process quality of simulated clinical interviews without artificially inflating examination outcomes. To support future research, we release a multi-expert annotated dataset comprising transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes. This resource aims to facilitate the development of pedagogically grounded AI-SP systems and advance research on AI-supported clinical reasoning training.