Robot Learning to Communicate through Projected Visual Abstractions

📄 arXiv: 2607.22434v1 📥 PDF

作者: Danyang Yan, Boyuan Wang, Jiaxun Liu, Boyuan Chen

分类: cs.RO, cs.AI

发布日期: 2026-07-24

备注: Our project website is at:https://generalroboticslab.com/shadow


💡 一句话要点

提出动态阴影表达方法以解决机器人沟通问题

🎯 匹配领域: 支柱一:机器人控制 (Robot Control)

关键词: 机器人沟通 动态阴影表达 视觉抽象 手部配置优化 人机交互

📋 核心要点

  1. 现有机器人主要依赖物理形态进行表达,缺乏通过视觉抽象进行沟通的能力。
  2. 提出一种动态阴影表达系统,利用软皮肤和学习的阴影自模型,实现手部配置与阴影外观的映射。
  3. 实验结果表明,该系统在多种场景下(如手语和动物模仿)成功实现了动态阴影表达,具有良好的视觉效果。

📝 摘要(中文)

人类通过身体的抽象形式(如阴影、轮廓和反射)进行沟通,而机器人仍主要依赖物理形态表达。为了使机器人能够通过投影视觉抽象进行沟通,本文提出了一种机器人系统,能够利用21自由度的灵巧手和学习的阴影自模型进行动态阴影表达。该系统通过软皮肤减少光泄漏,生成视觉上连续的轮廓,并通过任务无关的自我探索学习手部配置与投影阴影外观之间的映射。机器人根据目标阴影图像或视频,通过基于梯度的搜索优化手部配置,并通过碰撞感知模拟精炼解决方案,以获得物理可行的运动。实验展示了机器人在手语手势、手影木偶和动物动作模仿中的阴影表达能力,建立了机器人操控自身投影视觉抽象进行沟通和视觉叙事的框架。

🔬 方法详解

问题定义:本文旨在解决机器人在沟通中缺乏通过视觉抽象(如阴影)表达的能力。现有方法主要依赖物理形态,无法灵活表达情感或信息。

核心思路:提出一种基于动态阴影表达的机器人系统,通过软皮肤减少光泄漏,并利用学习的阴影自模型实现手部配置与阴影外观的映射,从而使机器人能够通过阴影进行有效沟通。

技术框架:系统包括三个主要模块:1) 21自由度的灵巧手,2) 学习的阴影自模型,3) 基于梯度的优化算法。机器人通过自我探索学习手部配置与阴影的关系,并在给定目标阴影时优化手部动作。

关键创新:最重要的创新在于引入了动态阴影表达和学习的阴影自模型,使机器人能够在不依赖物理形态的情况下,通过阴影进行沟通。这一方法与传统的机器人表达方式有本质区别。

关键设计:系统设计中采用了软皮肤以减少光泄漏,确保阴影的视觉连续性。同时,使用了碰撞感知模拟来优化手部动作,确保运动的物理可行性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,机器人在手语手势、手影木偶和动物动作模仿中成功实现了动态阴影表达,视觉效果显著。与基线方法相比,优化后的阴影表达在视觉连贯性和运动流畅性上有明显提升,展示了该系统的有效性和实用性。

🎯 应用场景

该研究的潜在应用领域包括人机交互、教育、娱乐和机器人表演等。通过使机器人能够通过阴影进行沟通,可以增强其在社交场合的表现力,提升人机协作的自然性和有效性。未来,该技术可能在智能家居、服务机器人等领域发挥重要作用。

📄 摘要(原文)

Humans routinely communicate through abstractions of their bodies, including shadows, silhouettes, and reflections. Yet robots remain largely confined to expressing themselves through their physical morphology. Enabling robots to communicate through such projected visual abstractions requires reasoning not only about bodily motion but also about how that motion is transformed into an external representation perceived by an observer. Among these abstractions, shadows provide a particularly compelling example because they emerge directly from the robot's embodiment while remaining visually distinct from the body itself. Here, we present a robotic system capable of dynamic shadow expression using a 21-degree-of-freedom dexterous hand with compliant soft skin and a learned shadow self-model. The soft-skinned embodiment reduces light leakage to produce visually continuous silhouettes, while the differentiable self-model learns the mapping between hand configurations and projected shadow appearance through task-agnostic self-exploration. Given a target shadow image or video, the robot optimizes its hand configurations through gradient-based search over 1 the learned self-model and refines the solution through collision-aware simulation to obtain physically feasible motions. For dynamic shadow performance, we further introduce expressive-region objectives, temporal smoothness regularization, and keyframe-based optimization to preserve visually important motion cues while reducing optimization complexity. We demonstrate robotic shadow expression across sign-language gestures, hand-shadow puppetry, and animal motion imitation in both simulation and physical experiments. These results establish a framework for enabling robots to manipulate projected visual abstractions of themselves for communication and visual storytelling.