ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
作者: Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo
分类: cs.CV, cs.AI, cs.LG
发布日期: 2026-07-21
💡 一句话要点
提出ABot-World-0以实现实时长时间闭环交互
🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)
关键词: 长时间交互 视频世界模型 教师蒸馏 多源数据 动态建模 实时推理 虚拟现实 机器人控制
📋 核心要点
- 现有方法在长时间交互和动态世界建模方面存在控制性不足和一致性问题。
- 论文提出通过多源数据学习可控世界动态,并利用教师强制和ODE蒸馏技术进行模型训练。
- 实验结果表明,ABot-World-0在WorldRoamBench上实现了竞争性的可控性和连贯的长时间世界演变。
📝 摘要(中文)
我们提出了ABot-World-0,这是一个基于动作条件的视频世界模型,旨在实现实时的长时间闭环交互。该模型依托于多源数据基础设施,包括AAA游戏、仿真引擎和互联网视频,以学习可控的世界动态。WorldExplorer通过训练反馈引导代理驱动的数据收集,同时统一管道应用14项确定性质量检查、基于VLM的评估以及同步的动作和文本注释。我们通过教师强制和ODE蒸馏逐步将双向动作条件教师提炼为因果学生,并引入LongForcing以对齐长时间学生自回归展开与扩展视野教师,减轻累积的分布偏移和自回归漂移。原始键盘动作提供了统一的控制接口,而参考角色记忆则在第三人称展开中提供持久的外观线索以保持身份一致性。
🔬 方法详解
问题定义:本论文旨在解决现有方法在长时间交互和动态世界建模中的控制性不足和一致性问题,尤其是在复杂环境下的闭环交互能力。
核心思路:论文提出了一种基于动作条件的视频世界模型,利用多源数据基础设施学习可控的世界动态,并通过教师强制和ODE蒸馏技术提升模型的表现。
技术框架:整体架构包括数据收集模块(WorldExplorer)、质量检查和评估模块、教师-学生蒸馏模块以及实时推理模块。数据收集由代理驱动,质量检查确保数据的准确性和一致性。
关键创新:最重要的技术创新在于引入LongForcing机制,能够有效对齐长时间自回归展开与扩展视野教师,从而减轻分布偏移和自回归漂移的问题。
关键设计:在模型设计中,采用了轻量级的VAE解码器、有效的注意力机制和内存感知调度,确保在单个NVIDIA RTX 5090桌面GPU上以720P视频流实现高效推理,具有1.2秒的动作到首帧延迟和约19GiB的峰值显存使用。
🖼️ 关键图片
📊 实验亮点
实验结果显示,ABot-World-0在WorldRoamBench上表现出色,能够在单个NVIDIA RTX 5090 GPU上以720P视频流实现高达16 FPS的性能,且具有1.2秒的动作到首帧延迟,展现出优越的可控性和连贯的长时间世界演变能力。
🎯 应用场景
ABot-World-0的研究具有广泛的潜在应用,尤其是在游戏开发、虚拟现实和机器人控制等领域。通过实现实时的长时间闭环交互,该模型能够提升用户体验和系统的智能化水平,推动相关技术的进一步发展。未来,该技术可能在自动化决策、智能助手和交互式娱乐等方面发挥重要作用。
📄 摘要(原文)
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.