ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
作者: Han Li, Si Liu, Zehao Huang, Dongxin Lyu, Longfei Xu, Jiahui Fu, Daxin Tian, Yuliang Xiu, Naiyan Wang
分类: cs.CV
发布日期: 2026-07-23
备注: 37 pages, 37 figures
💡 一句话要点
提出ViSTR-Bench以评估MLLMs在动态场景中的推理能力
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 多模态大型语言模型 空间-时间推理 动态场景 视觉推理 基准评估 视频问答 定性推理
📋 核心要点
- 现有的多模态大型语言模型在动态场景的空间-时间推理能力上存在明显不足,尤其是在直观推理方面。
- 本文提出了ViSTR-Bench基准,旨在通过连续视觉线索评估MLLMs的定性推理能力,填补现有评估的空白。
- 实验结果表明,尽管当前模型在视频理解上表现良好,但在复杂推理任务中仍显著低于人类水平。
📝 摘要(中文)
多模态大型语言模型(MLLMs)在多种专家级任务中取得了显著成功,但在空间感知和动态推理等基本能力上仍存在不足。现有基准主要关注静态场景或要求精确的定量预测,缺乏对时间线索的直观推理评估。本文提出了视觉空间-时间推理基准(ViSTR-Bench),旨在系统评估MLLMs在动态场景中从连续视觉线索进行定性推理的能力。ViSTR-Bench建立了一个涵盖运动感知、空间关系、结果预测和物理动态的四维评估体系,包含15个子任务和1340对高质量视频问答对。广泛评估表明,尽管当前模型在视频理解方面表现强劲,但在复杂的空间-时间推理上仍面临显著瓶颈,远低于人类表现。
🔬 方法详解
问题定义:本文旨在解决现有多模态大型语言模型在动态场景中进行空间-时间推理的不足,尤其是缺乏对时间线索的直观推理能力。
核心思路:ViSTR-Bench基准通过引入定性推理的评估框架,强调时间、推理方向和定性评估,系统性地评估MLLMs在动态场景中的表现。
技术框架:ViSTR-Bench的评估体系包括四个主要维度:运动感知、空间关系、结果预测和物理动态,涵盖15个子任务和1340对视频问答对。
关键创新:该基准的创新之处在于其专注于定性推理,而非仅依赖于定量预测,填补了现有基准的不足。
关键设计:在设计上,ViSTR-Bench采用了高质量的视频问答对,确保了评估的多样性和挑战性,同时设置了明确的评估标准以衡量模型的推理能力。
🖼️ 关键图片
📊 实验亮点
实验结果显示,尽管当前的多模态大型语言模型在视频理解方面表现优异,但在复杂的空间-时间推理任务中,其性能仍显著低于人类,表明该领域仍有广阔的改进空间。
🎯 应用场景
该研究的潜在应用领域包括智能监控、自动驾驶、机器人导航等,能够帮助模型更好地理解和推理动态场景中的复杂交互。这将推动多模态AI系统在实际应用中的性能提升,具有重要的实际价值和未来影响。
📄 摘要(原文)
Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturally develop through continuous observation of the real world, such as spatial perception and dynamic reasoning. Recent studies have recognized this gap and introduced dedicated benchmarks to evaluate the spatial-temporal capabilities of MLLMs. However, existing benchmarks mostly focus on static scenes or require exact quantitative predictions, leaving intuitive reasoning from temporal cues largely underexplored. In this paper, we introduce the Visual Spatial-Temporal Reasoning Benchmark (ViSTR-Bench), a novel evaluation suite designed to systematically assess whether MLLMs can perform qualitative reasoning from continuous visual cues in dynamic scenes. Guided by the principles of temporal emphasis, reasoning orientation, and qualitative evaluation, ViSTR-Bench establishes a comprehensive four-dimensional evaluations covering Motion Perception, Spatial Relations, Outcome Prediction, and Physical Dynamics. The benchmark comprises 15 distinct subtasks and 1,340 high-quality video question-answer pairs spanning diverse tabletop, indoor, and outdoor scenarios. Extensive evaluations of a broad spectrum of state-of-the-art proprietary, open-source, and specialized spatial MLLMs reveal that, despite their strong general video understanding capabilities, current models still face substantial bottlenecks in complex spatial-temporal reasoning and remain far below human performance.