KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

📄 arXiv: 2607.12625v1 📥 PDF

作者: Yunxin Li, Jinchao Li, Shibo Su, Zhenran Xu, Chenrui Zhao, Tongshu Bian, Xiaoman Liang, Meishan Zhang, Baotian Hu, Min Zhang

分类: cs.CL, cs.CV

发布日期: 2026-07-14

备注: 29 pages, 9 figures


💡 一句话要点

提出KnowAct-GUIClaw以解决OpenClaw的GUI交互与自我进化问题

🎯 匹配领域: 支柱一:机器人控制 (Robot Control)

关键词: 跨平台交互 自我进化机制 任务自动化 用户交互经验 智能助手

📋 核心要点

  1. 现有的OpenClaw框架在跨平台GUI交互和自我进化机制上存在不足,限制了其适应性和性能提升。
  2. 提出的KnowAct-GUIClaw框架通过积累用户交互经验和任务知识,实现任务的长远分解与分配,支持跨平台迁移。
  3. 实验结果显示,KnowAct-GUIClaw在Android、iOS等平台上表现出色,特别是在MobileWorld基准测试中达到了64.1%的最佳性能。

📝 摘要(中文)

OpenClaw作为复杂任务自动化的领先代理框架,面临跨平台GUI交互支持不足和自我进化机制不完善的问题。这些缺陷限制了其在多样设备生态系统中的适应性,并阻碍了通过执行经验的持续学习来提升性能。为此,本文提出了个人助手的Know Deeply, Act Perfectly范式,强调用户交互和任务执行经验的积累直接提升执行的准确性和效率。基于此范式,本文介绍了KnowAct-GUIClaw,一个旨在解决OpenClaw GUI操作不足的Know-Route-Act-Reflect框架,能够打破跨平台和递归自我改进的限制。实验表明,KnowAct-GUIClaw在多个平台上实现了卓越的效率和准确性,尤其在MobileWorld基准测试中表现优异。

🔬 方法详解

问题定义:本文旨在解决OpenClaw在跨平台GUI交互支持不足及自我进化机制不完善的问题。这些痛点限制了其在多设备生态系统中的适应性和性能提升。

核心思路:提出Know Deeply, Act Perfectly范式,强调通过积累用户交互和任务执行经验来提升执行的准确性和效率,统一认知理解与操作执行。

技术框架:KnowAct-GUIClaw框架包含Know-Route-Act-Reflect四个模块。首先,主代理利用积累的交互经验和任务相关知识进行任务分解与分配(Know);其次,具备经验可归属记忆系统和自我进化技能库的可插拔GUI子代理(Act),实现无缝的跨平台迁移和快速集成。

关键创新:最重要的创新在于引入了经验可归属的记忆系统和自我进化的技能库,使得框架能够持续存储用户资料和反馈,从而提高任务分解和工具调用的准确性。

关键设计:框架设计中,用户交互经验的存储和反馈机制是关键,确保了任务执行的持续改进。具体的参数设置和网络结构细节在论文中进行了详细描述。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,KnowAct-GUIClaw在多个平台上实现了卓越的效率和准确性,特别是在MobileWorld基准测试中,GUIClaw与开源Kimi-2.6模型结合,达到了64.1%的最佳性能,超越了所有现有的代理框架和闭源模型,提升幅度达到8.5%。

🎯 应用场景

该研究的潜在应用领域包括个人助手、智能家居控制、企业自动化等。通过提升跨平台的适应性和执行效率,KnowAct-GUIClaw能够在多种设备和环境中提供更为智能的用户体验,未来可能对智能助手的发展产生深远影响。

📄 摘要(原文)

OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a well-built self-evolution mechanism. These flaws limit its adaptation to diverse device ecosystems and prevent performance improvements through continuous learning from execution experience. To resolve these issues, we propose the Know Deeply, Act Perfectly paradigm for personal assistants, which holds that accumulated user interaction and task-running experience directly improve execution accuracy and efficiency, unifying cognitive comprehension and operational execution. Based on this paradigm, we introduce KnowAct-GUIClaw, a novel Know-Route-Act-Reflect framework designed to address OpenClaw's GUI manipulation deficits and break through its cross-platform and recursive self-improvement constraints. First, the host agent leverages accumulated interaction experience and task-relevant knowledge for long-horizon task decomposition and allocation (Know). Second, a pluggable GUI subagent with an experience-attributable memory system (Know) and self-evolving skill library (Act), enabling seamless cross-platform migration and fast-path integration. Especially, this framework continuously stores user profiles and feedback to improve the accuracy of task decomposition and tool calls. Extensive experiments across Android, iOS, HarmonyOS and Windows show that KnowAct-GUIClaw achieves superior efficiency, accuracy and cross-platform adaptability. Especially, the GUIClaw with open-source Kimi-2.6 models achieves the best performance (64.1%) on the long-horizon MobileWorld benchmark, beating all agentical frameworks and closed-source agentical models, e.g., Seed-2.0-Pro and GPT-5.5. Additionally, the knowledgeable memory and execution skills supported by our framework are transferable across diverse base models, improving by 8.5% with Kimi-2.6.