PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation

📄 arXiv: 2609.04867v1 📥 PDF

作者: Yuchen Sun, Qian Yang, Jun Wang, Detai Xin, Guoqiao Yu, Guanglu Wan, Qi Jia

分类: cs.MM, cs.AI, cs.SD

发布日期: 2026-09-04

备注: 19 pages, 10 figures, 4 tables. Accepted at ACM Multimedia 2026 (MM '26). This arXiv version includes supplementary appendices not included in the conference proceedings version

DOI: 10.1145/3767308.3836100


💡 一句话要点

提出PRISM-Bench以解决T2AV生成中音频评估不足的问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 音频生成 多模态评估 文本到音频视频 生成模型 评估基准

📋 核心要点

  1. 现有的T2AV生成评估方法未能充分重视音频模态,导致难以准确诊断系统的性能。
  2. PRISM-Bench通过构建一个包含900个样本的音频中心基准,因子化音频评估,提供更全面的评估框架。
  3. 实验结果显示,当前生成模型在音频生成方面存在显著性能差距,尤其在复杂的音频生成任务中表现不佳。

📝 摘要(中文)

文本到音频视频生成(T2AV)技术迅速发展,但其评估仍低估了音频模态。现有基准要么将音频视为视频质量的辅助成分,要么与视听基础无关地单独评估,难以诊断当前系统在音频生成中的真正成功与失败。为此,本文提出了PRISM-Bench,这是首个以音频为中心的T2AV生成诊断基准。该基准基于900个经过人工验证的样本构建,从两个正交轴(音频类型和声音源可见性)对音频评估进行因子化,并通过35个细化标准在四个感知维度上进行评估。通过增强的MLLM作为评判者协议,确保了评估的可靠性,显示出与人类评审者的强一致性(超过70%的平均一致性)。

🔬 方法详解

问题定义:本文旨在解决现有T2AV生成评估中音频模态被低估的问题。现有方法往往将音频视为视频质量的附属部分,难以全面评估音频生成的效果。

核心思路:PRISM-Bench通过因子化音频评估,分别从音频类型(如语音、音乐和声音)和声音源可见性(如屏幕内和屏幕外)两个维度进行评估,提供更细致的评估标准。

技术框架:PRISM-Bench的整体架构包括数据集构建、评估维度设定和评估协议。数据集由900个经过人工验证的样本组成,评估维度包括音频-视觉一致性、音频质量、音频表现力和提示遵循。

关键创新:最重要的创新在于采用了增强的MLLM作为评判者协议,通过盲测和对比真实参考,确保评估的可靠性和一致性。与现有方法相比,PRISM-Bench提供了更全面的音频评估框架。

关键设计:在评估过程中,使用了35个细化标准来评估生成内容,并确保与人类评审者的强一致性(超过70%的平均一致性)。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,当前的T2AV生成模型在音频生成方面存在显著的性能差距,尤其是在复杂的音频生成任务中表现不佳。PRISM-Bench的评估显示,前沿模型与开源模型之间存在明显的性能差异,强调了对音频生成能力的重视。

🎯 应用场景

PRISM-Bench的提出为文本到音频视频生成领域提供了一个新的评估标准,能够帮助研究人员更好地理解和改进生成模型的音频表现。其潜在应用包括影视制作、游戏音效生成以及教育领域的多媒体内容创作等,未来可能推动相关技术的进一步发展与应用。

📄 摘要(原文)

Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio modality. Existing benchmarks either treat audio as an auxiliary component of video quality or assess it in isolation from audiovisual grounding, making it difficult to diagnose where current systems truly succeed or fail in audio generation. We present PRISM-Bench, the first audio-centric diagnostic benchmark for T2AV generation. Built from a rigorously curated dataset of 900 human-verified samples, PRISM-Bench factorizes audio evaluation along two orthogonal axes: audio type (Speech, Music, and Sound) and sound-source visibility (On-screen vs. Off-screen). It evaluates generated content across four perceptual dimensions (Audio-Visual Coherence, Audio Quality, Audio Expressiveness, and Prompt Following) with 35 fine-grained criteria. To ensure reliable assessment, we adopt an enhanced MLLM-as-a-Judge protocol based on blind, side-by-side comparison against ground-truth references, demonstrating strong alignment (over 70% mean agreement) with human raters. Our evaluation of recent T2AV systems highlights a significant performance gap between frontier and open-source models. Furthermore, we demonstrate that current generation paradigms overfit to perceptual fidelity while struggling with complex grounding and control tasks, particularly in generating music and synchronized On-screen audio.