Do Reviewers Still Reward Lexical Complexity? A Frozen-Rater Study of Preference Drift in 124K ICLR Reviews

📄 arXiv: 2609.08475v1 📥 PDF

作者: Jiabin Zheng

分类: cs.CL, cs.DL, cs.LG

发布日期: 2026-09-08

备注: 23 pages, 8 figures, 11 tables. Code and the machine-readable records behind every number: https://github.com/Biajin-PKU/frozen-rater-drift


💡 一句话要点

提出冷冻评审者模型以分析ICLR评审偏好变化

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 词汇复杂性 同行评审 冷冻评审者 大型语言模型 偏好变化 文本生成 评审标准

📋 核心要点

  1. 核心问题:评审者对词汇复杂性的重视是否因大型语言模型的普及而发生变化,现有研究无法明确区分评审者与文本的影响。
  2. 方法要点:通过冷冻评审者模型,使用同一模型和提示生成评审,以隔离评审者偏好的变化与提交内容的变化。
  3. 实验或效果:分析结果显示,人类评审者对非领域词汇复杂性的重视从+0.142降至-0.015,而冷冻评审者的评估保持稳定。

📝 摘要(中文)

随着大型语言模型降低了生成复杂文本的成本,同行评审者是否仍然重视词汇复杂性成为一个关键问题。本文通过冷冻评审者方法,分析了2018至2025年间ICLR提交的81,850篇机器生成评审的变化,发现人类评审者对非领域词汇复杂性的重视程度显著下降,而冷冻评审者的评估保持稳定。研究表明,评审者对生产成本降低的信号产生了折扣,且人类与冷冻评审者的偏好逐渐不一致。

🔬 方法详解

问题定义:本文旨在探讨评审者对词汇复杂性的偏好是否因大型语言模型的普及而发生变化。现有方法无法有效区分评审者的偏好变化与提交内容的变化,导致结果的不确定性。

核心思路:论文提出使用冷冻评审者模型,通过在同一时间窗口内生成评审,确保评审者的偏好不变,从而准确分析评审者对文本特征的反应。

技术框架:整体流程包括生成81,850篇机器评审,使用相同的模型和提示,分析2018至2025年间的评审数据,比较人类评审者与冷冻评审者的偏好变化。

关键创新:最重要的技术创新在于冷冻评审者的使用,使得研究能够独立于评审者的变化,准确识别文本特征对评审结果的影响。

关键设计:在实验中,采用了双重假发现控制和区间排除的设计,确保结果的可靠性,并对未通过对抗性重测的发现进行了报告。

🖼️ 关键图片

fig_0
img_1
img_2

📊 实验亮点

实验结果显示,人类评审者对非领域词汇复杂性的重视显著下降,从+0.142降至-0.015,而冷冻评审者的评估保持在+0.080至+0.082之间,表明评审者的偏好发生了显著变化。

🎯 应用场景

该研究的潜在应用领域包括学术评审、教育评估和文本生成领域。通过理解评审者的偏好变化,研究者和教育者可以更好地调整评审标准和文本生成策略,以适应新的评审环境。

📄 摘要(原文)

Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about the evaluator, not about the text. When the association between a writing cue and review scores moves across years, the reviewers may have changed, the submissions may have changed, or both, and a regression of scores on text cannot say which. We separate the two with a frozen rater: 81,850 machine reviews of ICLR submissions from 2018 to 2025, all generated in one February-April 2025 window with one model family and one prompt, so that its year-to-year coefficients track submission composition alone and the human-minus-frozen trend difference identifies reviewer preference drift. On 32,638 submissions with 124,615 human reviews, the human coefficient on non-domain lexical complexity falls from +0.142 to -0.015 while the frozen rater moves from +0.080 to +0.082; the three-way difference-in-differences is -0.0100 (q=0.013), and forty random-wordlist placebos through the same specification centre on zero. Humans still reward sentence-length variability, which the frozen rater never registers, while the frozen rater still pays for lexical complexity at its earlier rate. Every claim is held to a double gate of false-discovery control and interval exclusion, and the findings that failed adversarial re-testing are reported. Reviewers discounted a cue whose production cost collapsed, as models of manipulable signals prescribe; an LLM judge calibrated to historical human preferences inherits the earlier schedule and drifts out of alignment while its agreement with humans on totals stays ordinary.