Geopolitical alignment: Endorsement effects in large language models
作者: Maxim Chupilkin
分类: cs.CY, cs.AI
发布日期: 2026-07-10
💡 一句话要点
研究大语言模型中的地缘政治影响及其评估偏差
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大语言模型 地缘政治 政策评估 endorsement实验 模型偏差 国际关系 可信度
📋 核心要点
- 核心问题:现有大语言模型在政策评估中可能受到地缘政治线索的影响,但这一现象尚未被系统研究。
- 方法要点:通过endorsement实验,研究不同国家支持下的政策评估差异,揭示模型对地缘政治的敏感性。
- 实验或效果:发现不同国家的支持影响模型评分,尤其是西方国家的支持被视为可信度的标志。
📝 摘要(中文)
大语言模型(LLMs)在总结和评估政策相关信息方面的应用日益增加,但其判断是否受到地缘政治线索的影响仍不明确。本文通过一项endorsement实验研究了这一问题,四个LLMs在随机描述为美国、欧盟、中国或俄罗斯支持的情况下评估相同的国际经济和安全政策。在仅数字评分的条件下,GPT-5、Claude Sonnet和Gemini对中国和俄罗斯支持的政策评分显著低于美国或欧盟支持的相同政策,而DeepSeek是主要例外。在要求模型提供简短理由的第二个条件下,GPT-5和Claude Sonnet的西方与非西方评分差距保持不变,Gemini的惩罚减弱,而DeepSeek对中国和俄罗斯的惩罚显著增强。这些发现表明,即使政策内容固定,LLM的政策评估也可能依赖于外国支持者的身份。
🔬 方法详解
问题定义:本文旨在探讨大语言模型在政策评估中是否受到地缘政治线索的影响。现有方法未能系统性地分析这一问题,导致对模型判断的理解不足。
核心思路:通过设计endorsement实验,随机将政策描述为不同国家支持,观察模型评分的变化,以此揭示地缘政治对模型评估的影响。
技术框架:实验分为两个主要条件:第一是仅数字评分,第二是要求模型提供评分理由。每个条件下,四个模型(GPT-5、Claude Sonnet、Gemini和DeepSeek)对相同政策进行评估。
关键创新:本研究的创新在于通过endorsement实验揭示了大语言模型在政策评估中对地缘政治支持的敏感性,尤其是西方与非西方国家支持的显著差异。
关键设计:实验中,模型的评分机制和理由生成模块被设计为能够捕捉到不同国家支持下的评估变化,尤其关注模型对西方支持的可信度和对非西方支持的风险感知。
🖼️ 关键图片
📊 实验亮点
实验结果显示,GPT-5、Claude Sonnet和Gemini对中国和俄罗斯支持的政策评分显著低于西方支持的政策,DeepSeek则表现出不同的评分模式。在提供理由的条件下,西方支持被普遍视为可信度的标志,而非西方支持则引发对数据安全和地缘风险的担忧。
🎯 应用场景
该研究的潜在应用领域包括国际关系分析、政策制定和舆论监测等。通过理解大语言模型如何受到地缘政治影响,可以更好地利用这些模型进行政策评估和决策支持,提升其在复杂国际环境中的应用价值。
📄 摘要(原文)
Large language models (LLMs) are increasingly used to summarize and evaluate policy-relevant information, but it remains unclear whether their judgments are implicitly shaped by geopolitical cues. I study this question with an endorsement experiment in which four LLMs evaluate the same international economic and security policies after each policy is randomly described as supported by the United States, the European Union, China, or Russia. In the numeric-only condition, GPT-5, Claude Sonnet, and Gemini rate China- and Russia-endorsed policies substantially lower than identical policies endorsed by the United States or the European Union; DeepSeek is the main exception. A second condition asks models to provide a short justification with the score. This request leaves the broad Western/non-Western gap intact for GPT-5 and Claude Sonnet, attenuates Gemini's penalties, and sharply activates China and Russia penalties in DeepSeek. The justifications indicate that Western endorsement is often treated as a credibility cue, whereas Chinese and Russian endorsement is treated as a cue for data security, sovereignty, surveillance, or geopolitical risk. These findings show that LLM policy evaluations can depend on the identity of a foreign endorser even when policy content is held fixed.