Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
作者: Shravan Murlidaran, Miguel P. Eckstein
分类: cs.CV, cs.AI
发布日期: 2026-07-10
💡 一句话要点
提出复杂社会行为数据集以评估视觉语言模型的准确性与错误类型
🎯 匹配领域: 支柱三:空间感知与语义 (Perception & Semantics) 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 视觉语言模型 复杂社会行为 多模态大语言模型 场景描述 视觉认知错误
📋 核心要点
- 现有的视觉语言模型评估主要集中在简单场景,未能充分反映复杂人类行为的理解能力。
- 本文提出复杂社会行为(CSB)数据集,并分析了VLMs在场景描述准确性和错误类型方面的演变。
- 实验结果显示,MLLMs在复杂场景描述准确性上显著提升,几乎消除了与人类描述的差距。
📝 摘要(中文)
视觉语言模型(VLMs)在过去十年中在视觉推理方面取得了显著进展。然而,大多数评估使用简单场景(如MS-COCO),未能展示复杂的人类互动或行为。本文引入了复杂社会行为(CSB)数据集,包含100幅描绘复杂社会互动的图像。我们分析了2017至2025年间VLMs的场景描述进展,评估了模型的准确性及20个人类描述的表现。结果表明,MLLMs在场景描述准确性上几乎消除了与人类描述的差距,并且在检测、识别和幻觉错误方面的影响最大。我们的研究为视觉语言模型在过去十年的进展提供了更全面的评估。
🔬 方法详解
问题定义:本文旨在解决现有视觉语言模型在复杂场景理解中的不足,尤其是对复杂人类行为的描述能力。现有方法主要依赖简单场景,未能有效评估模型的错误类型和准确性。
核心思路:通过引入复杂社会行为(CSB)数据集,提供更具挑战性的评估基准,分析VLMs在复杂场景中的表现,特别是错误类型的识别与分析。
技术框架:研究首先构建CSB数据集,包含100幅复杂社交互动图像。接着,评估四个预多模态大语言模型(pre-MLLMs)和五个多模态大语言模型(MLLMs)的准确性,并与20个人类描述进行对比。
关键创新:最重要的创新在于引入CSB数据集,填补了现有评估中对复杂人类行为的缺失,同时系统分析了五种视觉认知错误类型。与现有方法相比,MLLMs在复杂场景描述准确性上表现出色,几乎消除了与人类描述的差距。
关键设计:在实验中,采用了与金标准进行对比的方式,分析了模型在对象检测、识别、幻觉、场景理解和空间依赖等方面的错误类型。
🖼️ 关键图片
📊 实验亮点
实验结果显示,MLLMs在CSB数据集上的场景描述准确性显著提高,达到了与顶级人类描述相似的水平,且在MS-COCO数据集上也表现出色。相比于预MLLMs,MLLMs在错误类型上几乎消除了所有问题,尤其是在检测、识别和幻觉错误方面的影响最大。
🎯 应用场景
该研究的潜在应用领域包括社交机器人、智能监控和人机交互等。通过提升视觉语言模型在复杂场景中的理解能力,可以为这些领域提供更智能的解决方案,促进人机协作的效率与准确性。
📄 摘要(原文)
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark, and have not focused on understanding the model's error types. Here, we introduce the Complex Social Behavior (CSB) dataset, containing 100 images depicting complex social interactions/behaviors. We analyze the progression of scene descriptions over a decade (2017-2025) of VLMs (four pre-Multimodal Large Language Models, MLLMs, and five MLLMs). We evaluate the accuracy of the models and 20 human descriptions relative to a gold standard on the CSB dataset and on a sample from MS-COCO. We analyzed five visual-cognitive error types: object detection, recognition, hallucination, scene understanding, and spatial dependence. The CSB dataset showed a more pronounced improvement than MS-COCO in scene description accuracy, with pre-MLLMs achieving much lower accuracy than the bottom-ranked human descriptions and MLLMs attaining accuracies similar to the top-ranked human descriptions. We show that MLLMs have eliminated the gap in scene description accuracy between simpler MS-COCO scenes and scenes depicting complex behaviors (CSB). MLLMs have almost eliminated all error types in our tested datasets, except for occasionally relying on different image regions for scene descriptions than humans do (spatial dependence error). We also show that detection, recognition, and hallucination errors have the highest impact on scene description accuracy. Together, our findings provide a more thorough evaluation of how visual language models have advanced over the last decade.