How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks
作者: Tahiya Chowdhury
分类: cs.CL
发布日期: 2026-07-20
备注: Accepted, to appear in Proceedings of ACM International Conference on Multimodal Interaction 2026, 13 pages, 6 figures, 4 tables
💡 一句话要点
提出三维评估框架以提升多模态对话状态信号的可靠性
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 多模态信号 对话状态 认知负荷 声学特征 语言特征 交互特征 视频会议 特征评估
📋 核心要点
- 现有方法在多模态对话状态测量中存在特征可靠性不足和任务上下文泛化能力差的问题。
- 论文提出了一个三维评估框架,评估特征的预测准确性、跨任务可推广性和重测可靠性,以提升多模态信号的可靠性。
- 实验结果表明,语言特征在认知负荷预测中表现最佳,但在跨任务评估中失效,而交互特征则提供了唯一可靠的信号。
📝 摘要(中文)
本研究测量多模态行为中的对话状态,如认知负荷和对话权力,要求特征不仅具有预测性,还需在任务上下文中可靠。我们提出了一个三维评估框架,评估预测准确性、跨任务可推广性和重测可靠性,应用于从视频会议平台中提取的交互、声学和语言特征。结果显示,没有单一特征家族在所有三个维度上占主导地位。语言特征在认知负荷的预测准确性上表现最佳,但在跨任务评估中崩溃,显示出对任务特定词汇的敏感性。此外,声学可靠性在控制说话者身份后降级,确认标准韵律特征测量的是声学特征而非对话状态。交互特征提供了唯一真正可靠的信号,未受说话者归一化的影响。我们的发现强调了说话者归一化和多维评估作为对话系统中上下文感知、稳健的多模态特征选择的前提。
🔬 方法详解
问题定义:本研究旨在解决多模态信号在对话状态测量中的可靠性和泛化能力不足的问题。现有方法往往忽视了特征在不同任务上下文中的表现,导致预测结果不稳定。
核心思路:我们提出了一个三维评估框架,专注于评估特征的预测准确性、跨任务可推广性和重测可靠性,以确保所选特征在不同情境下的有效性。
技术框架:该框架包括三个主要模块:1) 交互特征提取,2) 声学特征分析,3) 语言特征评估。每个模块都针对特定的对话状态进行特征提取和评估。
关键创新:本研究的关键创新在于提出了一个综合评估框架,强调了说话者归一化的重要性,并揭示了声学特征在控制说话者身份后可靠性下降的现象,这与传统方法形成鲜明对比。
关键设计:在特征提取过程中,采用了标准的声学和语言特征提取算法,并在评估中引入了重测可靠性分析,确保特征在不同时间点的一致性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,语言特征在认知负荷预测中准确率最高,但在跨任务评估中表现不佳,接近零的声学可靠性在控制说话者身份后显现。交互特征则提供了唯一稳定的信号,表明在对话状态测量中需重视特征的上下文适应性。
🎯 应用场景
该研究的潜在应用领域包括智能对话系统、情感分析和人机交互等。通过提升多模态信号的可靠性,能够为更复杂的对话系统提供支持,进而改善用户体验和交互效果。未来,这一框架可能推动对话系统在教育、医疗和客服等领域的广泛应用。
📄 摘要(原文)
Measuring conversational states such as cognitive load and conversational power from multimodal behavior requires characteristic features that are not only predictive but also reliable across task contexts. We present a three-dimensional evaluation framework assessing predictive accuracy, cross-task generalizability, and test-retest reliability, applied to interactional, acoustic, and linguistic features extracted from dyadic conversations during collaborative tasks performed over a video-conferencing platform (AVCAffe dataset; 53 dyads, 9 tasks). Our results show that no single feature family dominates all three dimensions. Linguistic features show the highest predictive accuracy for cognitive load but collapse under cross-task evaluation, revealing sensitivity to task-specific vocabulary. Additionally, acoustic reliability, often reported as evidence of feature stability, degrades once speaker identity is controlled, confirming that standard prosodic features measure vocal characteristics rather than conversational state. Interaction features provide the only genuinely reliable signal, unchanged after speaker normalization. Interestingly, classifying power role remained near chance baseline across all conditions, indicating limitations of task-level aggregated behavior for predicting power role in conversation. Our findings reveal three insights: (1) linguistic features predict best but generalize poorly across task contexts; (2) acoustic reliability collapses to near-zero once speaker identity is controlled, challenging standard evaluation practice; and (3) interaction features provide the only genuinely reliable signal, with floor dominance predicting within-dyad cognitive load asymmetry. These results argue for speaker normalization and multi-dimensional evaluation as prerequisites for context-aware, robust multimodal feature selection in conversational systems.