Embodied GPT-5.1: Evidence of a World Model?

📄 arXiv: 2607.23899v1 📥 PDF

作者: Roberto Spinelli, Thiago C. Martins

分类: cs.RO, cs.AI

发布日期: 2026-07-27

备注: 6 pages, 16 figures. Published in the 2026 Brazilian Conference on Robotics (CROS)

期刊: R. Spinelli and T. C. Martins, "Embodied GPT-5.1: Evidence of a World Model?," 2026 Brazilian Conference on Robotics (CROS), pp. 394-399, 2026

DOI: 10.1109/CROS69211.2026.11565684


💡 一句话要点

探索GPT-5.1作为物理移动机器人高层控制器的可能性

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 多模态语言模型 物理机器人 空间推理 物理理解 自主导航 认知科学 人工智能

📋 核心要点

  1. 现有方法通常认为具身经验是发展空间推理和物理理解的必要条件,限制了对无具身模型的研究。
  2. 本研究通过使用GPT-5.1模型,仅依赖低分辨率图像和离散动作集,探索其在物理环境中的控制能力。
  3. 实验结果表明,GPT-5.1在物体定位和导航任务中展现出一定的空间推理能力,尽管存在一些感知限制。

📝 摘要(中文)

本研究探讨了大型多模态语言模型GPT-5.1在没有任何先前具身经验和模拟环境训练的情况下,能否作为物理移动机器人的高层控制器。通过低分辨率的第一人称图像和离散动作集,模型被要求进行导航和物体导向行为。实验结果显示,GPT-5.1展现出空间推理和物理理解的初步能力,如在物体离开视野后仍能保持短期记忆、推断自身运动的物理后果,以及执行连贯的动作序列。这些发现挑战了认知科学和机器人领域的传统观点,即具身经验是发展此类智能的必要前提,激发了对大型语言模型物理理解的深入研究的兴趣。

🔬 方法详解

问题定义:本研究旨在解决大型语言模型在没有具身经验的情况下,是否能够有效控制物理机器人这一问题。现有方法普遍认为具身经验是发展智能的基础,限制了对无具身模型的探索。

核心思路:论文提出通过低分辨率的第一人称图像和离散动作集,测试GPT-5.1在物理环境中的导航和物体导向行为,探索其是否具备空间推理和物理理解能力。

技术框架:研究设计了一个实验框架,GPT-5.1作为控制器,通过图像输入和动作输出进行交互。主要模块包括图像处理、动作选择和反馈机制,形成闭环控制系统。

关键创新:本研究的创新在于展示了GPT-5.1在缺乏具身训练的情况下,仍能展现出类似世界模型的行为,这一发现挑战了传统认知科学和机器人理论。

关键设计:在实验中,使用低分辨率图像作为输入,设定离散的动作集,并设计了短期记忆机制以保持物体位置的记忆。模型的训练未涉及任何具身经验,直接测试其在物理环境中的表现。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,GPT-5.1在多次试验中成功执行了物体定位和导航任务,展现出短期记忆和空间推理能力。尽管存在一些感知限制,如不精确的对齐策略和偶尔的误识别,但整体表现超出预期,表明其具备一定的世界模型特征。

🎯 应用场景

该研究的潜在应用领域包括自主机器人导航、智能家居系统以及人机交互等。通过进一步探索大型语言模型的物理理解能力,可以推动智能机器人在复杂环境中的自主决策能力,提升其在实际应用中的表现和适应性。

📄 摘要(原文)

This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level controller of a physical mobile robot despite having no prior embodiment, no training in simulated environments, and no exposure to sensorimotor experience. Using only low-resolution first-person images and a discrete action set, the model was tasked with navigation and object-directed behaviors such as locating and contacting a target toy. Across multiple trials, GPT-5.1 demonstrated emergent capabilities that suggest elements of spatial reasoning and physical understanding. These included maintaining short-term memory of object locations after they left the camera frame, inferring the physical consequences of its own movements, and executing coherent action sequences such as colliding with an object and reversing to visually verify the outcome. At the same time, the model displayed inefficiencies and perceptual limitations, including imprecise alignment strategies and occasional misidentification of distant distractors. Overall, the results indicate that GPT-5.1 exhibits signs of world-model-like behavior in an embodied setting, despite the absence of any embodiment-related training, a finding that challenges long-standing views in cognitive science and robotics which hold that a physical body is a necessary prerequisite for developing such forms of intelligence. The findings motivate deeper investigation into the emergence, limits, and robustness of physical understanding in large language models.