Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

📄 arXiv: 2607.12739v1 📥 PDF

作者: Binwen Liu, Yilin Ren

分类: cs.CL

发布日期: 2026-07-14


💡 一句话要点

提出ESFP基准以评估语言模型的认知立场灵活性

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 语言模型 认知立场 基准评估 灵活性 内容密度 对话系统 自我归因 模型评估

📋 核心要点

  1. 现有基准未能有效评估语言模型在不同归因条件下的认知立场变化,导致模型表现的全面性不足。
  2. 提出ESFP基准,通过对比外部归因与自我归因的提示,系统性地评估模型的认知立场灵活性。
  3. 实验评估了八个前沿模型,发现认知灵活性与模型能力并不完全相关,内容密度提供了最强的评估信号。

📝 摘要(中文)

本文探讨了语言模型在处理专家观点与自身观点时的认知立场差异。为此,提出了ESFP基准,旨在评估模型在不同归因条件下的反应灵活性。ESFP包含104个经过精心控制的项目,涵盖六个认知类别和五种措辞模板,评估模型在词汇自我归因、角色框架响应、句子内容密度及跨条件一致性等四个维度的表现。实验结果显示,认知灵活性与模型的整体能力大致正交,且内容密度是最强的信号。

🔬 方法详解

问题定义:本文旨在解决现有基准无法有效评估语言模型在不同归因条件下的认知立场变化的问题。现有方法主要关注准确性和安全性,未能深入探讨模型的认知灵活性。

核心思路:论文提出ESFP基准,强调外部归因与自我归因提示的对比,作为评估模型认知立场灵活性的基本单位。这种设计能够更好地捕捉模型在不同情境下的反应差异。

技术框架:ESFP基准由104个项目组成,涵盖六个认知类别和五种措辞模板。评估模型反应的四个维度包括词汇自我归因、角色框架响应、句子内容密度和跨条件一致性。

关键创新:ESFP基准的最大创新在于将认知立场的灵活性作为评估模型能力的新维度,强调了模型在不同归因条件下的适应性,而非单一的能力测量。

关键设计:在设计中,采用了多维度评估方法,包括内容密度的量化和词汇标记的灵活性分析。通过引入LLM评审小组,确保了评估的客观性和准确性。实验还提供了项目级的自助置信区间和敏感性分析。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,认知灵活性与模型的整体能力大致正交。例如,一个27B参数的开放模型在认知灵活性上与一些最强的专有系统相匹配,而某些旗舰模型的表现却低于其轻量级对应物。内容密度被发现是最强的信号,表明模型在表达立场时的内容丰富性至关重要。

🎯 应用场景

该研究的潜在应用领域包括智能对话系统、信息检索和内容生成等。通过提高语言模型在不同情境下的认知立场适应能力,可以增强其在复杂对话中的表现,提升用户体验和信任度。未来,ESFP基准可能成为评估语言模型认知能力的重要工具。

📄 摘要(原文)

A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversational agent should distinguish these two requests and respond in different epistemic registers: neutral attribution in the first case and stance expression in the second. Whether such a shift occurs-and whether it occurs coherently-is not directly assessed by existing benchmarks for accuracy, instruction following, or safety. We introduce ESFP, a behavioral benchmark that treats the contrast between externally attributed and self-attributed prompts as the fundamental unit of measurement. ESFP consists of 104 carefully controlled items spanning six epistemic categories and five phrasing templates, and evaluates model responses along four complementary dimensions: lexical self-attribution, representation-level responsiveness to role framing, sentence-level stance content density assessed by an LLM judge panel, and cross-condition stance consistency. Evaluating eight frontier models from five vendors, we find that epistemic flexibility is largely orthogonal to general model capability: a 27B open-weight model matches the strongest proprietary systems, the flagship model of one family underperforms its lightweight counterpart, and reasoning-optimized models do not consistently exhibit higher flexibility. Stance content density provides the strongest signal, while surface-level lexical markers such as 'I think' can change substantially without corresponding changes in expressed stance. We provide item-level bootstrap confidence intervals, weight-sensitivity analyses, and an explicit discussion of the interpretation limits of the composite score. ESFP measures a model's propensity to adapt its epistemic stance under changing attribution conditions, rather than a general competence measure.