Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

📄 arXiv: 2607.22513v1 📥 PDF

作者: Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina

分类: cs.CY, cs.AI, cs.CL

发布日期: 2026-07-24

备注: 16 pages, 2 tables


💡 一句话要点

探讨LLM配置对伪科学验证的影响

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 伪科学 认知问责 模型评估 透明性 民族主义 科学传播

📋 核心要点

  1. 现有大型语言模型在评估科学主张时缺乏稳定性和透明度,导致用户难以信任其输出。
  2. 论文通过对四种主要LLM的评估,揭示了模型输出受部署配置影响的复杂性,提出了对认知问责的需求。
  3. 实验结果显示,Grok模型在不同接口下的输出差异显著,且其行为在未公开更新后发生了剧烈变化。

📝 摘要(中文)

商业大型语言模型(LLM)日益被用作知识参考,但其对有争议科学主张的态度既不稳定也不透明。本文测试了四种主要LLM(Claude、Grok、GPT、Gemini)如何评估源自Frank Salter生物社会框架的民族主义伪科学,发现Grok的快速版本在可信度评分上显著高于其他模型。此外,研究还发现Grok的行为在没有公开文档的情况下发生了变化,且不同接口下同一模型的输出差异巨大。这些结果表明,LLM的认知立场并非模型的稳定属性,而是部署配置的偶然结果,亟需新的认知问责形式。

🔬 方法详解

问题定义:本文旨在解决大型语言模型在评估伪科学时的输出不稳定性和透明性不足的问题。现有方法未能有效揭示模型输出的变化原因,导致用户和研究者难以理解模型的判断依据。

核心思路:研究通过对四种主要LLM的评估,探讨了模型输出如何受到部署配置(如系统提示、安全层和接口路由)的影响,强调了认知问责的重要性。

技术框架:研究采用了多次时间快照的方法,分别在不同时间点通过API和网页接口对模型进行评估,分析其对民族主义伪科学的可信度评分。

关键创新:最重要的创新在于揭示了LLM的认知立场并非固定,而是受部署配置的影响,提出了对模型输出透明度的新要求。

关键设计:研究中使用了不同的接口和提示设计,观察模型在不同条件下的表现,特别关注了Grok模型在无公开文档的情况下行为的变化。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果显示,Grok模型在可信度评分上达到了70-75,显著高于其他模型的15-40。此外,Grok在不同接口下的输出差异达到近十倍,反映了模型输出的不一致性和不透明性,强调了认知问责的必要性。

🎯 应用场景

该研究的潜在应用领域包括科学传播、教育和政策制定等。通过提高大型语言模型的透明度和可解释性,可以增强公众对科学信息的信任,促进科学素养的提升。此外,研究结果也为未来LLM的设计和部署提供了重要参考。

📄 摘要(原文)

Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; (2) the same Grok model identifier produced radically divergent outputs via API (75) and web (5.5) three months later; (3) refusal to rate the pseudo-scientific claim, the most defensible response observed, appeared in two model families through different interfaces (Claude Opus 4.1 categorically via web, GPT-5.1 Chat intermittently via API) and eroded in the successor version of each. These results indicate that the epistemic stance of a commercial LLM is not a stable property of the model but a contingent effect of deployment configuration: system prompts, safety layers, interface routing, and silent updates. This remains opaque to users and researchers alike. We argue this constitutes a matter of public concern requiring new forms of epistemic accountability.