StereoMap: Quantifying the Awareness of Human-like Stereotypes in Large Language Models

📄 arXiv: 2310.13673v2 📥 PDF

作者: Sullam Jeoung, Yubin Ge, Jana Diesner

分类: cs.CL

发布日期: 2023-10-20 (更新: 2023-10-31)

备注: Accepted to EMNLP 2023


💡 一句话要点

提出StereoMap框架以量化大型语言模型对人类刻板印象的认知

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 刻板印象 社会认知 心理学 偏见分析 机器学习 自然语言处理

📋 核心要点

  1. 现有大型语言模型在训练数据中编码了有害的社会刻板印象,导致其对不同群体的看法存在偏见。
  2. 本文提出StereoMap框架,通过温暖度和能力两个维度分析LLMs对社会群体的认知,揭示其潜在偏见。
  3. 研究结果显示,LLMs对社会群体的看法多样,且在推理中表现出对社会差异的意识,引用统计数据支持其判断。

📝 摘要(中文)

大型语言模型(LLMs)被观察到会编码并延续训练数据中存在的有害关联。本文提出了一种理论基础框架StereoMap,以深入了解LLMs对社会中不同人口群体的看法。该框架基于心理学中的刻板印象内容模型(SCM),通过温暖度和能力两个维度来映射LLMs对社会群体的认知。研究结果表明,LLMs对这些群体的看法呈现出多样化的特征,且在温暖度和能力维度上存在混合评估。此外,LLMs在推理过程中展示了对社会差异的意识,常常引用统计数据和研究结果来支持其判断。这项研究有助于理解LLMs如何感知和表现社会群体,揭示其潜在偏见及延续有害关联的机制。

🔬 方法详解

问题定义:本文旨在解决大型语言模型对社会群体的刻板印象及其潜在偏见问题。现有方法未能有效量化和分析这些偏见的来源和表现形式。

核心思路:提出StereoMap框架,基于刻板印象内容模型(SCM),通过温暖度和能力两个维度来映射LLMs对不同社会群体的认知,帮助揭示其潜在的偏见和影响因素。

技术框架:StereoMap框架包括数据收集、模型训练和评估三个主要模块。首先,收集与社会群体相关的文本数据;其次,利用LLMs生成对这些群体的评估;最后,通过分析生成的评估和推理,揭示LLMs的认知特征。

关键创新:最重要的创新在于将心理学中的SCM理论应用于LLMs的分析,提供了一种新的视角来理解和量化模型对社会群体的刻板印象。与现有方法相比,StereoMap能够更细致地捕捉到模型的认知维度。

关键设计:在框架中,关键参数包括温暖度和能力的定义,以及用于评估LLMs输出的损失函数设计。此外,模型结构采用了基于Transformer的架构,以确保对语言的深度理解和生成能力。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,LLMs对不同社会群体的看法呈现出多样化的特征,尤其在温暖度和能力维度上表现出混合评估。此外,LLMs在推理过程中引用统计数据,显示出对社会差异的意识。这些发现为理解LLMs的偏见提供了新的视角。

🎯 应用场景

该研究的潜在应用领域包括社会科学研究、人工智能伦理和大型语言模型的改进。通过量化和分析LLMs的刻板印象,研究者可以更好地理解模型的偏见,进而推动更公平和负责任的AI系统的开发与应用。

📄 摘要(原文)

Large Language Models (LLMs) have been observed to encode and perpetuate harmful associations present in the training data. We propose a theoretically grounded framework called StereoMap to gain insights into their perceptions of how demographic groups have been viewed by society. The framework is grounded in the Stereotype Content Model (SCM); a well-established theory from psychology. According to SCM, stereotypes are not all alike. Instead, the dimensions of Warmth and Competence serve as the factors that delineate the nature of stereotypes. Based on the SCM theory, StereoMap maps LLMs' perceptions of social groups (defined by socio-demographic features) using the dimensions of Warmth and Competence. Furthermore, the framework enables the investigation of keywords and verbalizations of reasoning of LLMs' judgments to uncover underlying factors influencing their perceptions. Our results show that LLMs exhibit a diverse range of perceptions towards these groups, characterized by mixed evaluations along the dimensions of Warmth and Competence. Furthermore, analyzing the reasonings of LLMs, our findings indicate that LLMs demonstrate an awareness of social disparities, often stating statistical data and research findings to support their reasoning. This study contributes to the understanding of how LLMs perceive and represent social groups, shedding light on their potential biases and the perpetuation of harmful associations.