Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models

📄 arXiv: 2609.08637v1 📥 PDF

作者: Luka Debevc, Nishan Chatterjee, Antoine Doucet, Senja Pollak, Matej Martinc

分类: cs.CL, cs.CY

发布日期: 2026-09-08


💡 一句话要点

提出政治坐标测试框架以评估大语言模型的政治偏见与公平性

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 政治偏见 大语言模型 评估框架 多维扰动 跨语言分析 仇恨言论检测 模型公平性

📋 核心要点

  1. 现有方法在评估大语言模型的政治偏见时存在测量伪影和反应引导偏差,导致结果不可靠。
  2. 本文提出了一种新的政治坐标测试框架,通过多维扰动空间的采样来增强评估的稳健性。
  3. 实验结果显示,大多数模型倾向于自由主义左派,且指令措辞和答案格式对坐标恢复有显著影响。

📝 摘要(中文)

大语言模型作为信息中介的使用日益普遍,但其政治行为的测量仍然脆弱,因为问卷结果混合了模型倾向、测量伪影和反应引导偏差。本文提出了一种稳健的政治坐标测试评估框架,采样300种配置,涵盖语言、框架、指令、答案格式、选项顺序和角色措辞等八维扰动空间。对八种Gemma 3和Qwen 3模型在14种语言和三种量化水平下进行评估,获得了设计平均的政治坐标及其不确定性。大多数模型平均倾向于自由主义左派,但指令措辞、语言和答案格式显著影响恢复的坐标。跨语言差异主要反映坐标漂移,而非文化推理的差异。

🔬 方法详解

问题定义:本文旨在解决大语言模型在政治偏见评估中的脆弱性,现有方法常因测量伪影和反应引导偏差而导致结果不准确。

核心思路:提出一种稳健的政治坐标测试评估框架,通过在八维扰动空间中采样300种配置,增强评估的可靠性和准确性。

技术框架:整体流程包括模型选择、配置采样、数据收集和坐标计算等主要模块,确保多样性和全面性。

关键创新:引入了多维扰动空间的概念,能够有效捕捉模型在不同条件下的政治倾向,克服了传统方法的局限性。

关键设计:在实验中,采用了多种语言和量化水平,设计了不同的指令措辞和答案格式,以评估其对政治坐标的影响。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,大多数模型在政治坐标上倾向于自由主义左派,指令措辞和答案格式对坐标恢复的影响显著。此外,较大模型在角色分离方面表现更为明显,尤其是在检测仇恨言论时,基础和中立提示的主题级情感一致性最高。

🎯 应用场景

该研究的潜在应用领域包括社交媒体内容审核、政治广告监测和信息传播分析。通过量化大语言模型的政治偏见,可以为政策制定者和技术开发者提供重要参考,促进信息的公平性与透明性。

📄 摘要(原文)

Large Language Models are increasingly deployed as information intermediaries, yet measuring their political behavior remains fragile because questionnaire results mix model dispositions with measurement artifacts and response-elicitation biases. We introduce a robust Political Compass Test evaluation framework that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording. We evaluate eight Gemma 3 and Qwen 3 models across 14 languages and three quantization levels, obtaining design-averaged political coordinates with quantified uncertainty. Most models lean Libertarian-Left on average, but instruction phrasing, language, and answer format significantly affect recovered coordinates. Cross-lingual differences primarily reflect coordinate drift rather than distinct cultural reasoning. Reverse-engineering the test also exposes axis-weighting imbalances and the collapse of degenerate responses toward the center, so near-origin estimates for the smallest models can reflect weak signal rather than centrism. Free-text reasoning and chat-then-classify elicitation alter recovered coordinates, and larger models show clearer persona separation, with a specific failure of the Authoritarian-Left persona to move most models in the intended social direction. In downstream tasks, persona effects are modest relative to model size and target group for hate-speech detection, while base and centrist prompts give the highest agreement for topic-level sentiment. Political role prompting therefore has measurable but task- and dataset-specific downstream effects.