Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs

📄 arXiv: 2607.16193 📥 PDF

作者: Like Liu, Zhengzheng Xu, Haitao He, Hongzhe Li, Shuchang Zhang, Dian Shao

分类: cs.CV

发布日期: 2026-07-20


💡 一句话要点

提出UAV-DualCog以解决无人机双重认知能力不足问题

🎯 匹配领域: 支柱三:空间感知与语义 (Perception & Semantics) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 无人机 双重认知 多模态大语言模型 时空推理 基准评估 数据自动化构建 环境理解

📋 核心要点

  1. 现有无人机基准主要关注单一任务,缺乏对无人机双重认知能力的综合评估。
  2. 本文提出UAV-DualCog基准,联合评估无人机自身状态与外部环境的推理能力。
  3. 实验结果表明,当前多模态大语言模型在双重认知任务中表现不佳,存在明显瓶颈。

📝 摘要(中文)

多模态大语言模型在多种视觉-语言任务中表现出色,但在无人机场景中的能力仍未得到充分探索。现有的无人机基准主要集中于场景理解、事件识别或导航完成,而未能联合评估无人机代理所需的双重认知能力,即在多视角时空上下文中推理无人机自身状态和外部环境。为填补这一空白,本文提出了UAV-DualCog基准,旨在评估无人机的多视角时空推理能力。该基准包括图像和视频任务,要求在离散答案预测之外进行空间或时间的基础推理。通过构建场景级语义点云的数据自动化流程,UAV-DualCog提供了一个可扩展的基准,涵盖多样的场景、数百个地标和数千个问答样本。评估结果显示,当前的多模态大语言模型在无人机双重认知方面仍然存在显著不足。

🔬 方法详解

问题定义:本文旨在解决无人机在多视角时空推理中的双重认知能力不足问题。现有方法多集中于单一任务,未能有效评估无人机对自身状态与外部环境的综合理解能力。

核心思路:提出UAV-DualCog基准,通过图像和视频任务联合评估无人机的自我状态与环境状态推理,强调空间和时间的基础推理能力。

技术框架:UAV-DualCog基准由多个模块组成,包括数据自动化构建、任务设计和评估流程。数据构建利用场景级语义点云,确保基准的多样性和可扩展性。

关键创新:UAV-DualCog的创新在于其双重认知评估框架,区别于现有方法的单一任务评估,能够更全面地反映无人机在复杂环境中的推理能力。

关键设计:在数据构建中,采用了场景级语义点云,确保了数据的丰富性;同时,设计了针对自我状态和环境状态的多样化任务,以提升评估的有效性。实验中使用了轻量级优化探针,提供了有结构的监督信息。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,当前多模态大语言模型在UAV-DualCog基准上表现不佳,尤其在自我状态推理、视角转换、精确空间定位和时间间隔定位等方面存在显著瓶颈。与人类基线相比,现有模型的理解能力仍有较大差距,表明该基准对模型的挑战性。

🎯 应用场景

该研究的潜在应用领域包括无人机自主导航、环境监测和灾害响应等。通过提升无人机的双重认知能力,能够显著增强其在复杂环境中的决策能力和任务执行效率,推动无人机技术的实际应用和发展。

📄 摘要(原文)

Multimodal large language models have achieved strong performance across diverse vision-language tasks, yet their capabilities in UAV scenarios remain insufficiently explored. Recent UAV-oriented benchmarks have begun to evaluate MLLMs in aerial scenarios, but they typically focus on scene understanding, event recognition, or navigation completion, rather than jointly assessing the dual-cognition capability required for UAV agents: reasoning about both the UAV's own state and the external environment in multiview spatio-temporal contexts. To address this gap, we present UAV-DualCog, a benchmark for aerial multiview spatio-temporal reasoning built on this dual-cognition perspective. UAV-DualCog includes both image and video tasks to jointly evaluate self-state and environment-state reasoning, while requiring spatial or temporal grounding beyond discrete answer prediction. We also develop an automated pipeline that constructs data from scene-level semantic point clouds, yielding a scalable benchmark with diverse scenes, hundreds of landmarks, and thousands of QA samples. Extensive evaluations show that current MLLMs remain far from reliable in UAV dual cognition. Self-state reasoning, viewpoint transformation, precise spatial grounding, and temporal interval localization are persistent bottlenecks, and additional validation with thinking/frontier models and a human baseline confirms that the benchmark is understandable to humans but challenging for existing models. We further construct UAV-DualCog-Train from disjoint scenes and show through a lightweight optimization probe that it provides useful structured supervision, suggesting its value not only as an evaluation benchmark but also as a data resource for advancing MLLM-based UAV agents. Project website and supplementary materials:this https URL