LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios

📄 arXiv: 2609.09790v1 📥 PDF

作者: Hanjing Zhou, Mingze Yin, Ying Lian, Jun Ma, Chang-Yu Hsieh, Yanbing Zhou

分类: cs.CV, cs.AI, cs.CL

发布日期: 2026-09-09


💡 一句话要点

提出LogiScope-VQA以解决工业场景中的物流危险识别问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 物流危险识别 多模态模型 工业安全 数据集构建 风险推理 人类标注 模型评估

📋 核心要点

  1. 现有方法在工业场景中缺乏足够的真实数据,导致模型在危险识别方面的能力不足。
  2. 本文提出LogiScope-VQA数据集,旨在评估主流LMMs在物流操作中的实际应用能力,并设计了39个子任务。
  3. 实验结果表明,当前强大的模型在危险识别任务上与人类表现相比仍有显著差距,显示出进一步改进的潜力。

📝 摘要(中文)

大型多模态模型(LMMs)在工业仓库环境中的大规模应用要求模型具备人类专家级的危险感知、理解和推理能力。然而,真实工业数据的稀缺性严重阻碍了进一步发展。为填补这一空白,本文构建了LogiScope-VQA,以研究主流LMMs在实际物流操作中的适用性。该数据集包含2476张图像和2918段视频,主要来源于真实物流园区,并经过人类标注者精心策划和验证的10274个VQA。基于18个核心对象和20种风险类型,设计了39个子任务,涵盖工业元素感知、仓库知识理解和潜在风险推理等主题。实验结果显示,即使是强大的专有模型,如GPT-5.5、Gemini-3.1-Pro和Claude-Opus-4.7,相较于人类表现仍存在显著差距,表明在危险识别的感知、理解和推理的联合集成方面仍有很大的提升空间。

🔬 方法详解

问题定义:本文旨在解决工业场景中物流危险识别的不足,现有方法因缺乏真实数据而无法有效评估模型性能。

核心思路:通过构建LogiScope-VQA数据集,提供丰富的图像和视频数据,结合人类标注的VQA,来评估和提升LMMs的危险识别能力。

技术框架:整体架构包括数据集构建、任务设计和模型评估三个主要模块。数据集包含图像、视频及其对应的VQA,任务设计围绕感知、理解和推理展开。

关键创新:LogiScope-VQA的独特之处在于其针对工业环境的专门设计,涵盖了多种风险类型和核心对象,填补了现有数据集的空白。

关键设计:在模型评估中,采用动态思维预算配置和双维风险偏差分析,深入探讨LMMs的性能特征。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,尽管使用了先进的模型,如GPT-5.5和Gemini-3.1-Pro,它们在危险识别任务上的表现仍显著低于人类水平,表明该领域仍有大量改进空间。

🎯 应用场景

该研究的潜在应用领域包括智能仓储、工业安全监控和自动化物流管理。通过提升模型在危险识别方面的能力,能够有效降低工业操作中的安全风险,推动智能化物流的发展。

📄 摘要(原文)

Large Multimodal Models (LMMs) large-scale deployment in industrial warehouse settings specifically necessitates that models exhibit human-expert-level hazard-oriented perception, understanding, and reasoning capabilities. However, the scarcity of real industrial data, tightly coupled to commercial terms, significantly hampers further advancement. To bridge this gap, we curate LogiScope-VQA to investigate the practical applicability of mainstream LMMs in real-world logistics operations. LogiScope-VQA comprises 2,476 images and 2,918 videos primarily sourced from real-world logistics parks, along with 10,274 VQAs meticulously curated and validated by human annotators. Grounded in 18 core objects and 20 risk types, we devise 39 subtasks aligned with three principal themes: industrial element perception, warehouse knowledge understanding, and potential risk reasoning. Furthermore, we incorporate dynamic thinking-budget configurations and dual-dimensional risk bias analyses to elucidate the properties of LMMs. Extensive experiments unveil that even powerful proprietary models, including GPT-5.5, Gemini-3.1-Pro, and Claude-Opus-4.7, exhibit a significant gap relative to human performance. The unique challenge of jointly integrating perception, understanding, and reasoning for hazard identification poses substantial headroom for further improvement on LogiScope-VQA. We additionally reveal the pervasive security bias issue that impedes LLMs' practical deployment in real-world settings. The industrial dataset is publicly available under the CC BY-NC-SA 4.0 license.