MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings
作者: Ziyi Wang, Yuhang Wu, Dongxu Piao, Xingyu Liu, Tianhui Zhou, Miao Liu
分类: cs.CL, cs.CV
发布日期: 2026-07-21
💡 一句话要点
提出MeetingToM以解决多方会议中的心智理论推理问题
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 心智理论 多模态大型语言模型 社交行为推理 伪共识 多方会议 评估基准 非语言线索 群体动态
📋 核心要点
- 现有的多模态ToM基准主要集中在可验证信号的问答,缺乏对潜在社会状态和群体动态的全面覆盖。
- 本文提出MeetingToM基准,专注于多方会议中的复杂社会行为推理,评估不同层次的心智理论能力。
- 实验结果显示,当前的MLLMs在非语言线索整合和真实共识推断方面存在显著不足,强调了该领域的挑战。
📝 摘要(中文)
心智理论(ToM)是推断他人信念、意图和知识状态的能力,对社会互动至关重要,但对于当前的多模态大型语言模型(MLLMs)来说,尤其是在多方会议中,仍然具有挑战性。现有的多模态ToM基准主要集中在基于视频的问答,覆盖面有限。本文提出了MeetingToM,一个用于自然多方会议中复杂社会行为推理的基准,专注于会议特有现象如伪共识。该基准分层组织,评估不同社交粒度的ToM,包括个体心理状态预测、双人理解和群体共识推理。我们提供了统一的评估协议,并对代表性MLLMs进行了系统分析,揭示了在整合非语言线索、推断隐藏态度和区分真实共识与伪共识方面的持续局限性。
🔬 方法详解
问题定义:本文旨在解决多方会议中心智理论推理的不足,现有方法主要关注显性信号,无法有效处理潜在的社会状态和群体动态。
核心思路:提出MeetingToM基准,通过分层评估不同社交粒度的ToM能力,特别关注伪共识现象,以更好地理解复杂的社会互动。
技术框架:整体架构包括三个主要模块:个体心理状态预测、双人理解和群体共识推理,采用统一的评估协议进行系统分析。
关键创新:MeetingToM的创新在于其分层评估机制,能够深入分析多方会议中的复杂社交行为,与现有方法相比,提供了更全面的评估视角。
关键设计:在设计中,采用了多模态数据融合技术,重点关注非语言线索的整合,损失函数设置以优化对伪共识的识别能力。
🖼️ 关键图片
📊 实验亮点
实验结果显示,当前的多模态大型语言模型在MeetingToM基准上的表现存在显著不足,尤其是在非语言线索整合和伪共识识别方面,性能提升幅度有限,强调了该领域的研究挑战。
🎯 应用场景
该研究的潜在应用领域包括会议记录分析、社交机器人和人机交互等。通过提升多模态模型在复杂社交场景中的推理能力,能够增强机器对人类社交行为的理解,从而在实际应用中提供更智能的交互体验。
📄 摘要(原文)
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal ToM benchmarks mainly focus on video-grounded question answering over overt, externally verifiable signals, and provide limited coverage of latent social states and group dynamics. We introduce MeetingToM, a benchmark for complex social behavior reasoning in naturalistic multi-party meetings. MeetingToM targets meeting-specific phenomena such as \textbf{pseudo-consensus}, where apparent agreement masks private dissent under social pressure. The benchmark is hierarchically organized to evaluate ToM at increasing levels of social granularity, including (i) subject-level mental state prediction, (ii) dyadic-level addressee understanding, and (iii) group-level consensus reasoning. We provide a unified evaluation protocol and conduct systematic analyses of representative MLLMs, revealing persistent limitations in integrating non-verbal cues, inferring hidden attitudes, and distinguishing genuine consensus from pseudo-consensus. Our results highlight key challenges and establish MeetingToM as a testbed for advancing meeting-grounded ToM in multimodal models.