RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

📄 arXiv: 2607.17977v1 📥 PDF

作者: Kehan Li, Bohan Hou, Minghao Zhu, Tianyi Zhang, Zesen Cheng, Zhikai Wang, Sicong Leng, Xin Li, Xiao Lin, Biying Yao, Minghua Zeng, Jiangpin Liu, Ronghao Dang, Jiayan Guo, Siteng Huang, Haoyu Zhao, Heng Ping, Yaxi Zhao, Kexiang Wang, Tong Lu, Shengke Xue, Jiahao Tang, Yulei Wang, Zejing Wang, Jianwei Gao, Shijian Lu, Chengju Liu, Jianfei Yang, Mingxiu Chen, Deli Zhao

分类: cs.RO

发布日期: 2026-07-20

备注: KL,BH,MZ,TZ,ZC,ZW,SL,XL,XL,BY,MZ,JL,RD contribute equally. Project Lead: Kehan Li and Xin Li project: https://alibaba-damo-academy.github.io/RynnBrain github: https://github.com/alibaba-damo-academy/RynnBrain huggingface: https://huggingface.co/collections/Alibaba-DAMO-Academy/rynnbrain-11 modelscope: https://modelscope.cn/collections/DAMO_Academy/RynnBrain-11


💡 一句话要点

提出RynnBrain 1.1以提升机器人感知与推理能力

🎯 匹配领域: 支柱一:机器人控制 (Robot Control) 支柱七:动作重定向 (Motion Retargeting) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 具身基础模型 机器人感知 空间推理 3D基础 多任务训练

📋 核心要点

  1. 现有的具身基础模型在感知、推理和操作的通用性和能力上存在不足,难以满足复杂环境中的需求。
  2. RynnBrain 1.1通过引入接触点预测和3D基础,增强了模型在机器人操作中的表现,支持更复杂的空间推理和规划任务。
  3. 实验结果显示,RynnBrain 1.1在多个基准测试中表现优异,尤其是122B-A10B模型在VSI-Bench等评估中超越了所有对比模型。

📝 摘要(中文)

我们提出了RynnBrain 1.1,这是一个涵盖2B、9B和122B-A10B规模的具身基础模型系列。该模型采用统一的时空和物理基础框架进行训练,支持具身感知、空间推理、定位和规划。与RynnBrain 1.0相比,RynnBrain 1.1进一步引入了跨模型系列的接触点预测和2B及9B模型的原生3D基础,生成的表示和输出与机器人操作更直接对齐。此外,我们开发了RynnBrain-VLA,具有统一的跨具身动作空间和具身特定的掩蔽,并在Unitree G1、Astripot-S1和Tianji-Wuji上进行了部署。RynnBrain 1.1在具身认知、定位和3D基础方面取得了优异的结果,122B-A10B模型在VSI-Bench、MMSI和RefSpatial-Bench上超越了所有评估的专有和开源模型。实际机器人实验表明,RynnBrain初始化的策略在性能上优于基于Qwen的代表性通用VLA,而联合多任务和多具身训练在过程分数和成功率上也优于单任务训练。

🔬 方法详解

问题定义:本论文旨在解决现有具身基础模型在复杂环境中的感知、推理和操作能力不足的问题。现有方法在处理多模态信息和空间推理时存在局限性,难以有效支持机器人操作。

核心思路:RynnBrain 1.1通过引入接触点预测和原生3D基础,增强了模型的空间理解和操作能力。该设计使得模型生成的表示与机器人操作需求更为一致,从而提升了实际应用效果。

技术框架:RynnBrain 1.1采用统一的时空和物理基础框架进行训练,包含多个模块,如感知模块、推理模块和规划模块,支持多种具身任务的执行。

关键创新:RynnBrain 1.1的主要创新在于跨模型系列的接触点预测和具身特定的3D基础,这使得模型在处理复杂任务时表现出更高的灵活性和准确性,与现有方法相比具有显著的优势。

关键设计:在模型设计中,采用了特定的损失函数来优化接触点预测,并在网络结构上进行了调整,以适应不同规模模型的需求,确保在多任务训练中保持高效性和准确性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

RynnBrain 1.1在VSI-Bench、MMSI和RefSpatial-Bench等多个基准测试中表现优异,尤其是122B-A10B模型超越了所有评估的专有和开源模型。此外,实际机器人实验显示,RynnBrain初始化的策略在性能上优于基于Qwen的模型,联合多任务训练提升了成功率和过程分数。

🎯 应用场景

RynnBrain 1.1的研究成果在机器人操作、自动化制造、智能家居等领域具有广泛的应用潜力。通过提升机器人在复杂环境中的感知与推理能力,该模型能够更好地适应动态变化的任务需求,推动智能机器人技术的发展。未来,该技术可能在服务机器人、无人驾驶等领域发挥重要作用。

📄 摘要(原文)

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.