MultiHuSE: A Multimodal Dataset for Humour Styles and Emotions
作者: Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat
分类: cs.CL, cs.CV, cs.MM
发布日期: 2026-09-10
备注: 7 pages, 3 figures, 5 tables. Accepted at IEEE CBMI 2025 (International Conference on Content-Based Multimedia Indexing), Dublin, Ireland
期刊: 2025 International Conference on Content-Based Multimedia Indexing (CBMI), Dublin, Ireland, 2025, pp. 1-7
DOI: 10.1109/CBMI66578.2025.11339313
💡 一句话要点
提出MultiHuSE数据集以解决幽默风格与情感识别问题
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 多模态数据集 幽默识别 情感分析 心理学 人机交互 视频理解 文本分析
📋 核心要点
- 现有方法多集中于二元分类,缺乏能够捕捉幽默心理维度及表达变化的数据集,导致幽默识别的准确性不足。
- 论文提出MultiHuSE数据集,包含多模态视频和文本样本,旨在系统分析幽默风格与情感之间的关系。
- 实验结果显示,多模态融合方法在幽默风格分类上准确率达到80.1%,相比单一模态方法提升了2.7%。
📝 摘要(中文)
计算机对语言幽默的识别仍然是一项挑战,需理解语言、表达风格、情感和文化背景。现有方法多集中于二元分类,缺乏捕捉幽默心理维度及表达变化的数据集。我们提出MultiHuSE,这是一个多模态数据集,包含2407个高质量视频,50位不同背景的演员表演1463个文本样本,涵盖四种幽默风格及中性内容。数据集独特地捕捉了多个演员对同一文本的不同解读,便于系统分析表现多样性。基线实验表明,多模态融合在幽默风格分类上优于单一模态方法,特别是在亲和幽默上提升显著。我们希望MultiHuSE为幽默与情感的心理理论提供实证支持,并为人类沟通、幸福感及AI驱动的互动研究开辟新途径。
🔬 方法详解
问题定义:本论文旨在解决计算机对语言幽默的识别问题,现有方法在幽默的心理维度和表达多样性方面存在不足,无法有效捕捉幽默的复杂性。
核心思路:提出MultiHuSE数据集,通过多模态视频和文本样本的结合,捕捉幽默的多样性和情感维度,从而提升幽默风格的分类准确性。
技术框架:数据集包含2407个高质量视频,50位演员表演1463个文本样本,涵盖四种幽默风格。实验中采用多模态融合技术,结合文本和视频信息进行分类。
关键创新:MultiHuSE数据集的最大创新在于其多模态特性,能够系统性地分析不同演员对同一文本的幽默解读,填补了现有数据集在心理维度上的空白。
关键设计:在实验中,采用了多模态融合模型,设置了适当的损失函数以优化分类效果,确保文本信息与视频信息的有效结合。
🖼️ 关键图片
📊 实验亮点
实验结果表明,多模态融合方法在幽默风格分类上准确率达到80.1%,相比单一模态方法的77.4%提升了2.7%。特别是在亲和幽默的分类中,准确率从66%提升至74%,显示出显著的效果提升。
🎯 应用场景
该研究的潜在应用领域包括人机交互、情感计算和心理健康等。通过更好地理解幽默与情感的关系,MultiHuSE数据集可为AI系统提供更自然的互动方式,提升用户体验,并促进心理健康相关研究的发展。
📄 摘要(原文)
Computational recognition of verbal humour remains a challenging task, requiring an understanding of language, delivery style, emotions, and cultural context. Most existing approaches focus on binary classification and lack datasets that capture psychological dimensions of humour alongside variations in expression. We introduce MultiHuSE, a multimodal dataset comprising 2,407 high-definition videos of 50 demographically diverse actors performing 1,463 text samples across four psychological humour styles (affiliative, aggressive, self-enhancing, and self-deprecating), as well as neutral content. A subset is additionally annotated for underlying emotions. The dataset uniquely captures multiple actor interpretations of the same texts, enabling systematic analysis of expressive diversity. Baseline experiments show that multimodal fusion outperforms unimodal approaches (80.1% vs. 77.4% accuracy) in humour style classification, with particularly strong gains for affiliative humour (66% to 74%). While text provides the strongest individual signal, fusion models deliver meaningful improvements. We hope that MultiHuSE provides empirical support for psychological theories linking humour and emotion, while also opening new avenues for research in human communication, well-being, and AI-driven interaction. The dataset is available for academic use under an End-User Licence Agreement.