Can We Trust Item Response Theory for AI Evaluation?

📄 arXiv: 2607.15190 📥 PDF

作者: Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao, Susu Zhang

分类: cs.AI

发布日期: 2026-07-20


💡 一句话要点

探讨IRT在AI评估中的可靠性挑战与解决方案

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 项目反应理论 AI评估 基准测试 统计模型 模型排名 能力分布 估计工具

📋 核心要点

  1. 现有的IRT方法在AI基准测试中面临数据模式不匹配的问题,影响了其可靠性。
  2. 论文通过模拟响应矩阵和比较多种估计工具,探讨IRT在AI评估中的适用性。
  3. 实验结果显示,经典估计器在大规模基准中不可行,而可扩展估计器在特定条件下可能导致不可靠的推断。

📝 摘要(中文)

随着AI基准测试越来越多地利用项目级统计模型,特别是项目反应理论(IRT),来评估模型能力、排名系统、选择信息性示例和诊断基准质量,然而AI基准数据往往与人类测试的数据模式存在差异。本文研究了这些差异如何影响IRT建模在AI评估中的可靠性。通过对六个广泛使用的LLM基准的项目参数和能力分布进行分析,模拟了三种常见IRT模型下的响应矩阵,并比较了四种在近期基准研究中使用的估计工具。结果表明,经典估计器在大型基准设置中可能变得不可行,而可扩展的估计器在小型或非正态分布的模型集上可能产生不可靠的项目级和排名推断。

🔬 方法详解

问题定义:本文旨在解决IRT在AI评估中由于数据模式不匹配而导致的可靠性问题。现有方法在处理少量模型和大量项目时表现不佳,且能力分布可能偏斜或多模态。

核心思路:通过模拟不同IRT模型下的响应矩阵,比较多种估计工具的性能,评估其在AI基准测试中的适用性和可靠性。

技术框架:研究采用了六个广泛使用的LLM基准,分析其项目参数和能力分布,使用三种IRT模型进行模拟,并比较了边际最大似然、马尔可夫链蒙特卡洛、变分推断和神经伪Siamese估计器。

关键创新:本研究的创新点在于系统性地评估了IRT在AI基准测试中的适用性,揭示了经典估计器在大规模设置中的局限性,以及可扩展估计器在特定条件下的潜在不可靠性。

关键设计:在实验中,设置了18,000种模拟条件,重点考察了计算可行性、可扩展性及IRT推断的可靠性,特别是在模型排名和项目特征推断方面的表现。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,经典IRT估计器在大型基准测试中不可行,而可扩展估计器在小型或非正态分布的模型集上可能导致不可靠的推断。这一发现强调了在AI评估中选择合适估计工具的重要性。

🎯 应用场景

该研究的潜在应用领域包括AI模型评估、基准测试设计和教育测评等。通过提供对IRT在AI评估中的适用性分析,研究为未来的基准测试提供了重要的理论支持和实践指导,促进了AI领域的标准化和可靠性提升。

📄 摘要(原文)

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or non-normally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.