Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

📄 arXiv: 2607.20058v1 📥 PDF

作者: Markus J. Buehler

分类: cs.AI, cond-mat.mes-hall, cond-mat.mtrl-sci, cs.CL

发布日期: 2026-07-22


💡 一句话要点

提出材料科学机制的开放权重语言模型解析方法

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 材料科学 语言模型 因果推理 状态变换 机制识别 科学计算 机器学习

📋 核心要点

  1. 现有的大型语言模型在科学问题回答上存在局限,无法明确其是否有效利用物理规律。
  2. 论文提出了一种新的方法,通过分析隐藏状态和状态变换来解析材料科学机制信息。
  3. 实验结果表明,该方法在识别机制家族和控制答案方面表现优异,成功识别了90%的机制家族。

📝 摘要(中文)

大型语言模型能够回答科学问题,但正确的输出并不一定表明模型是否有效地表示或利用了物理规律。本文展示了开放权重的google/gemma-4-E4B-it模型中材料科学机制信息的三种可分离形式:概念在单个隐藏状态中可读,构成方向通过状态间的受控变换传递,选定的内部表示因果控制工程答案。通过匹配的直接和雅可比词汇读出、无选项状态几何、60条反事实基准和因果干预,研究在50个保留的材料描述中实现了概念排名的再现,并通过目标无关的词集实现了10个机制家族的盲识别。

🔬 方法详解

问题定义:本文旨在解决大型语言模型在科学问题回答中缺乏物理规律表示的不足,现有方法无法有效区分模型输出与物理机制的关系。

核心思路:通过分析模型的隐藏状态和状态间的变换,提出了一种新的解析材料科学机制的方法,强调了因果关系的控制。

技术框架:整体架构包括直接和雅可比词汇读出、状态几何分析和因果干预,结合反事实基准进行验证。

关键创新:最重要的创新在于将材料科学机制信息分为可读概念、构成方向和因果控制,提供了新的视角来理解模型的内部工作机制。

关键设计:采用了60条反事实基准和匹配的状态变换,确保了实验的严谨性,并通过盲识别验证了机制家族的准确性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,使用新方法在50个材料描述中成功识别了90%的机制家族,且在方向性法律的正确排序中达到了39/40的高准确率,显著优于传统方法。

🎯 应用场景

该研究可广泛应用于材料科学、工程设计和智能系统等领域,帮助科学家和工程师更好地理解和利用材料特性,推动新材料的开发和应用。未来,该方法可能在其他科学领域的模型解析中发挥重要作用。

📄 摘要(原文)

Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families. A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. We therefore compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations ordered direct, physically neutral and inverse laws across 60 frozen relations and correctly oriented 39 of 40 directional laws, whereas lexical controls were near chance. Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases, while counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. Physical relationships were therefore more visible in controlled state changes than in absolute states alone.