DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs

📄 arXiv: 2609.10253v1 📥 PDF

作者: Bhuvan Arora, Devesh Saraogi, Sravya Varada, Dhruv Kumar

分类: cs.CL, cs.AI

发布日期: 2026-09-09


💡 一句话要点

提出DiSCo框架以评估大型语言模型的文化偏好偏差

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 文化偏好 偏见评估 上下文适应 分布优先 多文化评估 人工智能伦理

📋 核心要点

  1. 现有文化基准仅通过单一正确答案评估,难以准确表征LLM的文化偏好,且未能区分默认偏好与上下文驱动的适应。
  2. 本文提出DiSCo框架,通过分布优先的强制选择评估,测试LLM在不同文化背景下的默认文化先验和可引导性。
  3. 实验结果表明,默认文化偏好高度集中,且提示引导无法有效缩小高低资源文化之间的选择差距。

📝 摘要(中文)

大型语言模型(LLMs)在全球助手中的应用日益广泛,但其在文化背景下的默认选择可能系统性地偏向某些文化,影响本地化、用户信任和公平行为。现有文化基准通过单一“正确”答案评估准确性,难以表征LLM的文化偏好。本文提出DiSCo,一个以分布为先的强制选择评估框架,旨在隔离默认文化先验并通过四级上下文梯度进行测试。使用DiSCo-Bench(304个项目)评估六种多样的指令调优LLM,发现默认偏好高度集中,英美文化占据约35%的选择,且提示引导无法有效解决文化偏好偏差。

🔬 方法详解

问题定义:本文旨在解决大型语言模型在文化背景下的偏好评估问题,现有方法无法有效区分文化偏好与上下文适应,导致评估结果不准确。

核心思路:提出DiSCo框架,以分布为先的强制选择评估方法,能够隔离默认文化先验并测试其可引导性,提供更全面的文化偏好评估。

技术框架:DiSCo框架包括四级上下文梯度(C0-C3),通过对304个项目的评估,分析不同文化背景下的选择分布,评估LLM的文化偏好。

关键创新:DiSCo框架的创新在于其分布优先的评估方法,能够有效识别和隔离LLM的默认文化偏好,与传统的单一答案评估方法有本质区别。

关键设计:在实验中,使用了DiSCo-Bench数据集,涵盖12种文化,评估六种不同的指令调优LLM,重点关注选择分布和文化偏好差异。实验设计中还考虑了提示引导的影响。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,默认文化偏好高度集中,英美文化的选择占比达到约35%,而提示引导未能有效缩小高低资源文化之间的选择差距,表明文化偏好偏差不能仅通过提示个性化解决。

🎯 应用场景

该研究的潜在应用领域包括全球化产品的本地化设计、文化敏感的人工智能助手开发以及用户信任的提升。通过更准确地评估和调整LLM的文化偏好,可以促进更公平和包容的技术应用,减少文化偏见。未来,该框架还可扩展至其他领域的偏见评估与调整。

📄 摘要(原文)

Large language models (LLMs) are increasingly deployed in globally used assistants, yet their default choices in culturally grounded everyday situations can systematically favour some cultures over others, affecting localisation, user trust, and equitable behaviour. Existing cultural benchmarks evaluate accuracy against a single "correct" answer, making it difficult to characterise an LLM's cultural preference prior when multiple culturally grounded responses are all valid; they also conflate default preferences with context-driven adaptation. We propose DiSCo, a distribution-first forced-choice evaluation framework that isolates default cultural priors and tests steerability via a four-level context gradient (C0--C3). Using DiSCo-Bench (304 items) derived from BLEnD spanning 12 cultures, we evaluate six diverse instruction-tuned LLMs. Default priors are heavily concentrated, with UK and US together absorbing approximately 35\% of all selections despite representing only 2 of 12 cultures. Most critically, prompt-based steering consistently widens the selection gap between high- and low-resource cultures, and injecting explicit cultural facts produces negligible distributional disruption, confirming that cultural preference bias cannot be resolved through prompt-based personalisation alone.