DYAD: A Multimodal Dataset of Co-Located Human Assistance

📄 arXiv: 2609.09023v1 📥 PDF

作者: Akhil Ajikumar, Mahya Qorbani, Sakib Reza, Sean Andrist, Mohsen Moghaddam

分类: cs.RO

发布日期: 2026-09-08

备注: 8 pages, 2 figures


💡 一句话要点

提出DYAD数据集以解决人机协作中的多模态交互问题

🎯 匹配领域: 支柱六:视频提取与匹配 (Video Extraction) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 多模态数据集 人机协作 协助行为分析 智能助手 任务状态识别

📋 核心要点

  1. 现有的数据集未能有效地将共处助手的语言和物理干预与执行者的请求和任务状态关联起来,导致协助行为的理解不足。
  2. DYAD数据集通过记录人类在齿轮组装过程中的协助行为,提供了一个多模态的同步记录,能够更好地分析协助过程中的互动。
  3. 实验结果表明,DYAD在模式事件的识别上表现优异,特别是在因果元数据和触发映射的准确性上,显示出显著的提升。

📝 摘要(中文)

本文介绍了DYAD(双人协助数据集),这是一个同步的多模态记录,旨在捕捉人类在齿轮组装过程中的协助行为。现有的数据集往往只关注单一执行或远程指令,而DYAD则将共处助手的语言和物理干预与执行者的请求、任务状态、协助触发和结果进行关联。通过20个会话,DYAD记录了528个任务步骤和611个执行者请求,涵盖了851条有效的协助记录。该数据集的注释跨越了协助过程,并通过三个参考任务评估选定组件的表现。实验结果显示,DYAD在不同模式事件上的表现优于现有方法,揭示了协助过程中的重要信息。

🔬 方法详解

问题定义:本文旨在解决现有数据集中缺乏对共处助手的语言和物理干预与执行者请求之间的关联性的问题。现有方法往往只关注单一的执行过程或远程指令,无法全面捕捉协助行为的复杂性。

核心思路:DYAD数据集通过记录人类在实际任务中的协助行为,采用同步的多模态数据记录方式,能够有效地链接语言和物理干预与任务状态和执行者请求,从而提供更全面的协助行为分析。

技术框架:DYAD的数据收集过程包括20个会话,其中一名训练有素的助手在协助HoloLens 2佩戴者时遵循指导优先的策略。数据记录涵盖528个任务步骤和611个执行者请求,形成851条有效的协助记录。

关键创新:DYAD的创新之处在于其链接的互动结构,涵盖了求助、干预选择、执行和结果,提供了一个全面的协助过程视角,而不仅仅是单一的执行步骤。

关键设计:数据集中包含的注释跨越了协助过程,并通过三个参考任务进行评估,涉及因果步骤理解、预发模式预测和指导者响应生成等关键组件。

🖼️ 关键图片

fig_0
fig_1

📊 实验亮点

在829个有效模式事件的实验中,DYAD数据集的四种种子RGB均值达到了0.548 +/- 0.007的宏F1分数,因果元数据的表现达到了0.624,而特权触发映射的准确率高达0.915,显示出DYAD在协助行为分析中的显著优势。

🎯 应用场景

DYAD数据集的潜在应用领域包括人机协作、智能助手设计和机器人学习等。通过提供丰富的多模态数据,研究人员可以更好地理解和优化人机交互过程,从而提升智能助手的响应能力和协助效果,推动相关技术的发展。

📄 摘要(原文)

An embodied assistant working beside a person must track task state, recognize help seeking, choose how to intervene, and produce an appropriate response. Existing procedural datasets richly describe individual execution, while interactive datasets capture remote verbal instruction or undifferentiated co-working. They do not jointly link a co-located helper's verbal and physical interventions to performer requests, task state, assistance triggers, and outcomes. We introduce DYAD (DYadic Assistance Dataset), a synchronized multimodal record of human-human assistance during gearbox assembly. Across 20 sessions, one trained helper follows a guidance-first policy while assisting HoloLens 2 wearers. DYAD links 528 task-step intervals and 611 performer requests with 851 valid assistance records spanning verbal and physical help. DYAD's annotations span the assistance process; three reference tasks evaluate selected components rather than an end-to-end system: causal step understanding, pre-onset mode anticipation, and instructor response generation. On 829 eligible mode events, the strongest four-seed RGB mean is 0.548 +/- 0.007 macro-F1; causal metadata reaches 0.624 and a privileged trigger mapping 0.915, revealing information not recovered from pre-onset RGB. DYAD's contribution is not scale, but a linked interaction structure spanning help seeking, intervention choice, execution, and outcome under egocentric and workspace sensing.