From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments

📄 arXiv: 2609.04894v1 📥 PDF

作者: Linsen Zhu, Mengqing Cai

分类: cs.AI, cs.LG, cs.MA

发布日期: 2026-09-04

备注: Review article. 29 pages, 1 figure, 3 tables. Literature cutoff: 31 August 2026


💡 一句话要点

提出合理授权框架以解决代理人工智能的局限性

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 代理人工智能 合理授权 多环境耦合 安全性 模型能力扩展 智能系统

📋 核心要点

  1. 现有的代理人工智能系统在安全授权和独立验证方面存在显著不足,限制了其在复杂环境中的应用。
  2. 论文提出合理授权的框架,强调在扩展代理能力时需基于证据,确保安全性和可靠性。
  3. 通过对现有研究的综合分析,论文指出行动接口的扩展效果显著,但在模型的独立性和恢复能力上仍需改进。

📝 摘要(中文)

大型语言模型在外部系统中成为重要代理,能够通过输出改变外部状态。本文对代理人工智能在数字、社交、虚拟和物理环境中的进展与局限进行了批判性回顾,整理了委托权、时间持久性和环境耦合的证据。尽管行动接口扩展的证据更为充分,但模型的独立验证和安全授权仍显不足。我们提出合理授权作为分析和规范的启发式工具,强调在证据支持的情况下扩展行动范围,以确保安全的恢复和人类控制。该框架为模型与环境的耦合评估、基于能力的权限设置和跨代理问责提供了研究议程。

🔬 方法详解

问题定义:本文旨在解决代理人工智能在多环境中的局限性,尤其是在安全授权、独立验证和持久性方面的不足。现有方法往往将模型能力与系统集成混淆,导致对自主性的误解。

核心思路:提出合理授权作为分析和规范的启发式工具,强调在扩展代理能力时需确保证据支持的安全性、失败检测和人类控制。

技术框架:整体架构包括模型、环境和授权机制三个主要模块。模型负责生成输出,环境提供交互接口,而授权机制则确保行动的安全性和可靠性。

关键创新:最重要的创新在于合理授权框架的提出,它为代理人工智能的能力扩展提供了系统化的分析工具,与现有方法相比,更加注重证据支持和安全性。

关键设计:在设计中,强调了基于能力的权限设置和跨代理问责机制,确保在多代理环境中能够有效管理和控制各个代理的行为。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果表明,行动接口的扩展在多种环境中表现出显著的效果,尤其是在任务完成率和系统集成度上有明显提升。然而,模型的独立验证和安全授权仍需进一步研究,以确保在复杂环境中的可靠性。

🎯 应用场景

该研究的潜在应用领域包括智能机器人、自动驾驶系统和复杂的多智能体系统。通过合理授权框架,可以提高这些系统在动态环境中的安全性和可靠性,推动代理人工智能的实际应用和发展。

📄 摘要(原文)

Large language models become consequential agents when surrounding systems let outputs change external state. Models now call tools, operate interfaces, delegate work, retain state, inhabit generated worlds, and control robots or laboratory equipment. Such advances are often narrated as one march toward autonomy, conflating model competence, system integration, persistence, and safe authority. This critical review synthesizes primary research and official technical specifications available by 31 August 2026. We organize the evidence along delegated authority, temporal persistence, and environmental coupling, while separating model, harness, and environment. Within the evidence examined, action-interface expansion is documented more convincingly than robust completion, recovery, authorization, or independent verification. Model Context Protocol and Agent2Agent improve interoperability but do not establish trustworthy delegation; multi-agent organization adds specialization alongside cost and correlated failure. Persistent simulations and world models support training and planning but do not themselves demonstrate agency; robotics and self-driving laboratories establish bounded feasibility rather than unattended open-world reliability. We propose justified delegation as an analytical and normative heuristic, not an observed law or certified score: expand action scope only where evidence supports provenance, bounded authority, failure detection, safe recovery, and calibrated human control. This framing yields a research agenda for coupled model-harness evaluation, capability-based permissions, durable state, cross-agent accountability, and staged physical validation.