ABot-N1: Toward a General Visual Language Navigation Foundation Model

📄 arXiv: 2607.10383 📥 PDF

作者: Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Yang Cai, Jingjing Ma, Shihui Su, Zixiao Tang, Linbo Zheng, Zedong Chu, Xiaolong Wu, Wenbin Tang, Mu Xu

分类: cs.CV, cs.AI, cs.RO

发布日期: 2026-07-20


💡 一句话要点

提出ABot-N1以解决视觉语言导航中的长尾语义和可解释性问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 视觉语言导航 深度推理 慢-快架构 长尾语义 可解释性 城市规模导航 多模态融合

📋 核心要点

  1. 现有视觉语言导航方法存在坐标漂移和长尾语义处理不佳的问题,且缺乏可解释性。
  2. ABot-N1通过慢-快架构解耦认知与控制,采用双重视觉-语言信号进行导航。
  3. 实验结果显示,ABot-N1在城市规模导航中提升了POI到达率35.0%,并在复杂场景中表现优异。

📝 摘要(中文)

视觉语言导航基础模型旨在统一深度推理与多样化的具身任务。现有方法通常通过单一策略直接将观察映射到动作,但存在坐标漂移和长尾语义处理不佳的问题。此外,这些黑箱映射缺乏可解释性,阻碍了通用性、鲁棒性和透明性的实现。本文提出的ABot-N1通过慢-快架构解耦认知与控制,利用双重视觉-语言信号来应对这些挑战。慢速视觉-语言推理器执行显式的思维链推理并生成像素目标,而快速动作专家则利用文本线索和像素指导生成连续的路径点。ABot-N1在城市规模导航中建立了新的最先进记录,特别是在复杂场景中提升了POI到达率35.0%。

🔬 方法详解

问题定义:本文旨在解决现有视觉语言导航模型在坐标漂移、长尾语义处理和可解释性方面的不足。这些问题导致模型在复杂环境中的表现不佳。

核心思路:ABot-N1的核心思路是通过慢-快架构将认知与控制解耦。慢速视觉-语言推理器负责进行显式的思维链推理,而快速动作专家则利用生成的像素目标进行高频控制。

技术框架:ABot-N1的整体架构包括两个主要模块:慢速视觉-语言推理器和快速动作专家。前者生成像素目标,后者基于这些目标和文本线索生成连续的路径点。

关键创新:ABot-N1的关键创新在于通过像素锚点与显式语言轨迹的结合,成功实现了高层意图与低层控制的桥接。这种设计显著提高了模型的鲁棒性和可解释性。

关键设计:在设计中,ABot-N1采用了特定的损失函数以优化推理和控制的协同作用,同时在网络结构上引入了多层次的特征提取,以增强模型对复杂场景的适应能力。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

ABot-N1在城市规模导航中取得了显著的实验结果,POI到达率提升了35.0%(达到77.3%),在复杂室内和室外场景中分别实现了95.4%和92.9%的成功率,展示了其在多种任务中的优越性和鲁棒性。

🎯 应用场景

ABot-N1在城市规模导航、室内外场景的任务执行中具有广泛的应用潜力。其可解释性和鲁棒性使其适用于智能机器人、自动驾驶和增强现实等领域,未来可能推动这些领域的技术进步与应用普及。

📄 摘要(原文)

Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift and poor handling of long-tail semantics. Furthermore, these black-box mappings lack interpretability, hindering the simultaneous achievement of generality, robustness, and transparency. We present ABot-N1, a step toward a general Visual Language Navigation foundation model, that addresses these challenges by decoupling cognition from control via a slow-fast architecture guided by dual visual-language signals. More specifically, a slow vision-language reasoner performs explicit Chain-of-Thought reasoning while producing a pixel goal. This compact set of image-space anchor points serves as a universal interface for diverse tasks, including point-goal, object-goal, poi-goal, instruction-following, and person-following. Subsequently, a fast action expert leverages both the textual cues and the pixel guidance to generate continuous waypoints at the native control frequency. By bridging high-level intents and low-level control through pixel-grounded anchors paired with explicit linguistic traces, our approach ensures robust, generalizable, and interpretable navigation across simulation and real-world benchmarks. ABot-N1 establishes new state-of-the-art records, delivering massive gains specifically in urban-scale navigation: boosting POI arrival by 35.0% (to 77.3%) and achieving 95.4%/92.9% SR in complex indoor and outdoor scenes. It also maintains superior robustness across object-reaching, person-following, and instruction-following tasks. New Point-Goal/POI-Goal benchmarks are released as open source to advance the field of urban-scale navigation.