ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space
作者: Mingchao Sun, Luyang Tang, Yu Liu, Xu Yan, Zhan Li, Yunwei Zhang, Fei Yu, Zengye Ge, Yumin Liu, Jiacheng Zhang, Yongchang Zhang, Jiawei Zhang, Zhicheng Liu, Zhongxu Sun, Tianjian Ouyang, Wenzheng Chen, Shixing Yang, Nianfei Fan, Guodong Sun, Huan Li, Zheng Zhou, Yongze Li, Yingliang Peng, Mengmeng Du, Yuan Liu, Haozhe Shi, Chunnuo Gong, Chengzhen Yu, Chunxue Jia, Yang Liu, Shiying Zeng, Junnan Lai, Hang Zhang, Ning Guo, Baoquan Chen, Mu Xu, Hongyu Pan
分类: cs.CV
发布日期: 2026-07-13
备注: Official Page: https://abot-world.amap.com/plaza
💡 一句话要点
提出ABot-3DWorld 0以实现多模态3D空间探索
🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture) 支柱三:空间感知与语义 (Perception & Semantics) 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 多模态融合 3D世界生成 空间生成原语 全景视频 高保真度
📋 核心要点
- 现有方法在多模态输入的3D世界生成中存在局限,难以实现高保真度和一致性。
- 论文提出的ABot-3DWorld 0通过统一的空间生成原语(SGP)有效整合多模态输入,生成可探索的3D世界。
- 实验结果显示,ABot-3DWorld 0在处理丰富多模态输入时,场景保真度优于现有的Marble方法。
📝 摘要(中文)
我们提出了ABot-3DWorld 0,这是一个通用的多模态3D世界模型,可以将文本、图像和视频输入转化为高保真、可探索的3D世界。该框架的核心是统一的空间生成原语(SGP),它是高质量全景图和空间点云的紧凑元组,能够高效描述任何3D空间。多模态输入首先被提升到这一原语;然后,3D一致的全景视频生成器沿着规划的轨迹探索该原语;最后,我们的全景视频重建引擎将生成的视频转换为干净、逼真的3D高斯点云世界。该管道覆盖了两种模式:丰富输入(多视角集、随意视频)通过几何严格的恢复提升到SGP,而单张图像或句子则生成性地完成到一个创意世界。实验表明,ABot-3DWorld 0在开源方法中设定了最新的技术水平,并在丰富的多模态输入下展现出比Marble更强的场景保真度。
🔬 方法详解
问题定义:本论文旨在解决现有多模态3D世界生成方法在场景保真度和一致性方面的不足,尤其是在处理不同类型输入时的局限性。
核心思路:ABot-3DWorld 0的核心思路是通过统一的空间生成原语(SGP)将多模态输入(如文本、图像和视频)转化为高保真的3D世界,确保生成过程的高效性和一致性。
技术框架:该框架包括三个主要模块:首先,将多模态输入提升到SGP;其次,利用3D一致的全景视频生成器沿规划轨迹探索SGP;最后,通过全景视频重建引擎将生成的视频转换为3D高斯点云世界。
关键创新:最重要的技术创新在于引入了统一的空间生成原语(SGP),这一设计使得不同类型的输入能够被有效整合并生成高保真的3D场景,显著提升了生成的灵活性和质量。
关键设计:在技术细节方面,论文强调了几何严格的恢复过程,以及在生成过程中使用的损失函数和网络结构,这些设计确保了生成的3D世界在视觉上与输入数据保持一致。
🖼️ 关键图片
📊 实验亮点
实验结果表明,ABot-3DWorld 0在处理丰富的多模态输入时,场景保真度显著优于现有的Marble方法,具体表现为在多个基准测试中设定了最新的技术水平,展示了更强的生成能力和一致性。
🎯 应用场景
ABot-3DWorld 0的潜在应用场景包括虚拟现实、游戏开发、城市规划和教育等领域。其高效的3D内容生成能力能够为用户提供沉浸式体验,并促进多模态数据的交互与探索,具有广泛的实际价值和未来影响。
📄 摘要(原文)
We present ABot-3DWorld 0, a universal multimodal 3D world model that turns text, image, and video inputs into high-fidelity, explorable 3D worlds. At the heart of our framework is a unified Spatial Generative Primitive (SGP), a compact tuple of a high-quality panorama and a spatial point cloud that delivers an efficient description of any 3D space. Multimodal inputs are first lifted into this primitive; a 3D-consistent panoramic video generator then explores the primitive along a planned trajectory; finally, our panoramic video reconstruction engine converts the generated video into a clean, photorealistic 3D Gaussian Splatting (3DGS) world. This pipeline covers two regimes: rich inputs (multi-view sets, casual video) are lifted into the SGP through a geometry-rigorous recovery that mirrors the observed scene, while a single image or sentence is completed generatively into a creative world. The result is one low-barrier engine for general 3D content creation that further anchors generated worlds to geographic points of interest, enabling map-native spatial exploration at consumer scale. Experiments show that ABot-3DWorld 0 sets the state of the art among open-source methods and demonstrates stronger scene fidelity than Marble under rich multimodal inputs.