Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos

📄 arXiv: 2607.20903v1 📥 PDF

作者: Liu Liu, Freya Huying Tan, Fábio Duarte

分类: cs.CV, cs.HC

发布日期: 2026-07-23

备注: 35 pages, 18 figures. Under review at Scientific Reports


💡 一句话要点

提出多模态表示方法以优化城市步行视频的参与度评估

🎯 匹配领域: 支柱八:物理动画 (Physics-based Animation)

关键词: 城市视频分析 参与度评估 多模态表示 时间平均图像 人类判断对齐

📋 核心要点

  1. 现有方法假设更丰富的视觉表示总是能更好地与人类判断对齐,但实验结果却显示出相反的趋势。
  2. 论文提出通过时间平均图像(TAIs)作为一种新的表示方式,来优化城市步行视频的参与度评估,尤其是在构图驱动的场景中。
  3. 实验结果表明,TAIs在高低参与度分类任务中表现优于视频特征,且与人类判断的准确性相当,挑战了传统观点。

📝 摘要(中文)

本研究探讨了更丰富的视觉表示是否能提高城市参与度的评估,使用来自YouTube的61个第一人称城市步行视频,分割为超过50,000个十秒片段,并通过四种模态进行表示:时空视频特征、时间平均图像(TAIs)、音频嵌入和基于文本的语义描述。Spearman相关分析显示,视频特征在连续对齐中表现最佳,但在高低参与度的二元分类中,TAIs在大多数分类器和分位数阈值下表现相当或更好。独立的两选一强迫选择研究确认了这一结果,参与者在识别参与时刻时,TAIs与完整视频的准确性相当,而文本和音频的表现较差。差距分析揭示了功能性分离:视频特征在动态内容的活动驱动场景中占优势,而TAIs在以稳定空间结构为主的构图驱动场景中更符合人类判断。这些发现挑战了更丰富的表示总是更符合人类的假设,并建议感知基础的时间压缩可以作为全视频编码的原则性替代方案。

🔬 方法详解

问题定义:本研究旨在解决在城市步行视频中,如何有效评估参与度的问题。现有方法通常依赖于丰富的视觉表示,但其有效性在不同场景中存在不确定性。

核心思路:论文提出使用时间平均图像(TAIs)作为一种新的表示方式,认为在某些场景中,TAIs能够更好地与人类的参与度判断对齐。

技术框架:研究通过分析61个城市步行视频,提取四种模态的特征,包括时空视频特征、TAIs、音频嵌入和文本描述,并进行Spearman相关分析和二元分类实验。

关键创新:最重要的技术创新在于提出了TAIs作为一种有效的替代表示,挑战了更丰富的视觉表示总是更符合人类判断的传统观念。

关键设计:在实验中,使用了多种分类器和量化阈值来评估不同模态的表现,特别关注了动态内容和稳定空间结构场景的表现差异。实验设计中还包括了独立的两选一强迫选择研究,以验证人类判断的准确性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,时间平均图像(TAIs)在高低参与度的二元分类任务中,表现出与完整视频相当的准确性,且在大多数分类器和分位数阈值下均优于视频特征。这一发现表明,TAIs在构图驱动场景中具有显著优势,挑战了传统的视觉表示假设。

🎯 应用场景

该研究的潜在应用领域包括城市规划、智能交通系统和人机交互设计等。通过优化视频内容的参与度评估,能够提升用户体验和参与感,进而推动相关技术的发展和应用。未来,该方法可能在视频分析和内容推荐系统中发挥重要作用。

📄 摘要(原文)

We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouTube segmented into over 50,000 ten-second clips and represented across four modalities: spatiotemporal video features, temporally averaged images (TAIs), audio embeddings, and text-based semantic descriptions. Spearman correlation analysis reveals the expected ordering along the temporal-richness continuum, with video features showing the strongest continuous alignment. However, this ordering breaks down under binary classification of high- versus low-engagement moments (the paradigm most commonly used to train perceptual scoring models), where TAIs consistently match or outperform video across most classifiers and quantile thresholds. An independent two-alternative forced-choice study on Amazon Mechanical Turk confirms that this parity reflects human judgment: participants identified engaging moments with comparable accuracy from TAIs and full video clips, while text performed substantially worse and audio remained near chance. Gap analysis reveals a functional dissociation: video features are advantaged in activity-driven scenes with dynamic content, whereas TAIs better align with human judgments in composition-driven scenes dominated by stable spatial structure. These findings challenge the assumption that richer representations are inherently more human-aligned, and suggest that perceptually grounded temporal compression can be a principled alternative to full video encoding.