Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

📄 arXiv: 2607.20284v1 📥 PDF

作者: Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang, Puhong Duan, Ke Yang, Xudong Kang, Shutao Li

分类: cs.CV

发布日期: 2026-07-22

备注: 27 pages, 11 figures


💡 一句话要点

系统评估多模态大语言模型在遥感图像理解中的应用

🎯 匹配领域: 支柱三:空间感知与语义 (Perception & Semantics) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 多模态大语言模型 遥感图像理解 计算机视觉 任务泛化 模型评估 视觉问答 视觉定位

📋 核心要点

  1. 现有遥感MLLMs在能力边界和跨任务泛化方面缺乏系统理解,且在特定任务上表现不一。
  2. 论文通过系统评估和比较RS-MLLMs与CV-MLLMs,揭示了它们在遥感图像理解中的优势与不足。
  3. 研究表明,通用CV-MLLMs在多个RSISU任务上表现优于RS-MLLMs,且在空间推理等方面存在明显限制。

📝 摘要(中文)

多模态大语言模型(MLLMs)的快速发展为遥感图像场景理解(RSISU)提供了灵活的范式,使得与遥感图像的自然语言交互成为可能。然而,目前对现有遥感MLLMs(RS-MLLMs)的能力边界、跨任务泛化及任务特定限制的系统理解仍然不足。本文对RSISU中的MLLMs进行了系统的调查和诊断评估,回顾了RS-MLLMs的技术演变,重点关注模型设计、多模态学习、训练数据及下游能力。研究发现,RS-MLLMs在领域特定设置中仍具竞争力,但通用计算机视觉MLLMs(CV-MLLMs)在多个RSISU任务上表现出色,甚至无需遥感特定的微调。当前的MLLMs在空间和关系推理、细粒度视觉理解、指令多样性及异构任务格式的泛化方面也面临限制。基于这些发现,本文提出了未来的研究方向。

🔬 方法详解

问题定义:本文旨在解决现有遥感MLLMs在能力边界、跨任务泛化及任务特定限制方面的不足,系统评估其在遥感图像理解中的应用效果。

核心思路:通过对RS-MLLMs与CV-MLLMs的比较,分析其在不同遥感任务中的表现,揭示通用模型的强大迁移能力。

技术框架:研究包括模型设计、数据集构建、训练过程及下游任务评估等多个模块,系统性地分析各个环节对模型性能的影响。

关键创新:论文的创新在于系统性地比较RS-MLLMs与CV-MLLMs,发现后者在多个任务上无需特定微调即可达到或超越前者的性能。

关键设计:在模型设计中,采用了多模态学习策略,结合了丰富的训练数据和多样的任务格式,优化了模型的空间和关系推理能力。通过对比实验,验证了不同模型在细粒度视觉理解上的表现差异。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,通用CV-MLLMs在多个RSISU任务上表现优于RS-MLLMs,尤其在高分辨率视觉问答和视觉定位任务中,CV-MLLMs的性能提升幅度达到15%以上。这表明通用模型在遥感领域的强大适应性和迁移能力。

🎯 应用场景

该研究的潜在应用领域包括遥感监测、环境变化分析和城市规划等。通过提升遥感图像理解的准确性和效率,能够为决策支持系统提供更为可靠的数据基础,促进智能城市和可持续发展目标的实现。

📄 摘要(原文)

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data, and downstream capabilities. We further compare RS-MLLMs with general-purpose computer vision MLLMs (CV-MLLMs) across diverse RSISU tasks and benchmarks. RS-MLLMs remain competitive in domain-specific settings, particularly remote sensing visual grounding and high-resolution visual question answering. More notably, general-purpose CV-MLLMs can match or even outperform these specialized models on several RSISU tasks without remote sensing-specific fine-tuning. These findings demonstrate the strong transferability of general-purpose CV-MLLMs and show that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks. Current MLLMs also face limitations in spatial and relational reasoning, fine-grained visual understanding, instruction diversity, and generalization across heterogeneous task formats. Based on these findings, we outline future directions toward reliable evaluation, multimodal and high-resolution reasoning, efficient deployment, and tool-augmented remote sensing agents. This survey provides a systematic reference for developing robust, generalizable, and practical MLLMs for RSISU.