Investigating Large Language Models' Perception of Emotion Using Appraisal Theory
作者: Nutchanon Yongsatianchot, Parisa Ghanad Torshizi, Stacy Marsella
分类: cs.CL, cs.AI
发布日期: 2023-10-03
期刊: 11th International Conference on Affective Computing and Intelligent Interaction Workshop and Demo (ACIIW) 2023 1-8
💡 一句话要点
通过评估理论研究大型语言模型的情感感知
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 情感感知 大型语言模型 评估理论 心理学 人机交互
📋 核心要点
- 现有大型语言模型在理解人类情感方面存在局限,尤其是在心理评估维度的反应上。
- 本研究通过应用压力与应对过程问卷(SCPQ)来评估LLM对情感的感知,探索其与人类反应的相似性与差异性。
- 实验结果表明,LLM在评估和应对的动态上与人类相似,但在关键维度上反应幅度存在显著差异,且对提问方式敏感。
📝 摘要(中文)
大型语言模型(LLM)如ChatGPT近年来取得了显著进展,广泛应用于公众交互中。理解这些黑箱模型,尤其是它们对人类心理方面的理解至关重要。本研究通过压力与应对过程问卷(SCPQ)探讨LLM的情感感知。SCPQ是一个经过验证的临床工具,包含多个随时间演变的故事,涉及可控性和可变性等关键评估变量。我们将SCPQ应用于OpenAI的三种最新LLM(davinci-003、ChatGPT和GPT-4),并将结果与评估理论和人类数据进行比较。结果显示,LLM的反应在评估和应对的动态方面与人类相似,但在关键评估维度上未能如理论和数据预测的那样有所不同。LLM的反应幅度在多个变量上也与人类存在显著差异。研究还发现,GPT对指令和提问方式非常敏感。这项工作丰富了对LLM心理学方面的评估文献,增进了对当前模型的理解。
🔬 方法详解
问题定义:本研究旨在解决大型语言模型在情感感知方面的不足,尤其是它们在评估理论中的表现与人类的差异。现有方法未能充分揭示LLM在情感理解中的复杂性。
核心思路:通过使用压力与应对过程问卷(SCPQ),本研究探讨LLM在情感评估中的表现,比较其与人类反应的相似性和差异性,以揭示LLM的情感理解机制。
技术框架:研究采用SCPQ作为评估工具,设计了多个故事情境,涵盖不同的评估变量。将三种LLM(davinci-003、ChatGPT、GPT-4)应用于这些情境,并与人类数据进行对比分析。
关键创新:本研究的创新点在于首次将评估理论应用于LLM的情感感知研究,揭示了LLM在情感理解中的动态反应与人类的相似性和差异性。
关键设计:研究中使用的SCPQ包含多个情境故事,重点关注可控性和可变性等变量。实验设计考虑了指令的敏感性,确保LLM在不同提问方式下的反应被准确记录和分析。
🖼️ 关键图片
📊 实验亮点
实验结果显示,LLM在情感评估的动态反应上与人类相似,但在关键评估维度上未能如理论预测的那样表现出差异。LLM的反应幅度在多个变量上显著不同,且对提问方式表现出高度敏感性。
🎯 应用场景
该研究的潜在应用领域包括心理健康评估、情感计算和人机交互等。通过深入理解LLM的情感感知能力,可以为开发更具人性化的AI系统提供理论基础,提升用户体验和交互质量。
📄 摘要(原文)
Large Language Models (LLM) like ChatGPT have significantly advanced in recent years and are now being used by the general public. As more people interact with these systems, improving our understanding of these black box models is crucial, especially regarding their understanding of human psychological aspects. In this work, we investigate their emotion perception through the lens of appraisal and coping theory using the Stress and Coping Process Questionaire (SCPQ). SCPQ is a validated clinical instrument consisting of multiple stories that evolve over time and differ in key appraisal variables such as controllability and changeability. We applied SCPQ to three recent LLMs from OpenAI, davinci-003, ChatGPT, and GPT-4 and compared the results with predictions from the appraisal theory and human data. The results show that LLMs' responses are similar to humans in terms of dynamics of appraisal and coping, but their responses did not differ along key appraisal dimensions as predicted by the theory and data. The magnitude of their responses is also quite different from humans in several variables. We also found that GPTs can be quite sensitive to instruction and how questions are asked. This work adds to the growing literature evaluating the psychological aspects of LLMs and helps enrich our understanding of the current models.