Can We Trust Item Response Theory for AI Evaluation?

📄 arXiv: 2607.15190v1 📥 PDF

作者: Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao, Susu Zhang

分类: cs.AI

发布日期: 2026-07-16


💡 一句话要点

评估AI时提出IRT模型的局限性与改进建议

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 项目反应理论 AI评估 基准测试 统计模型 模型排名 性能评估

📋 核心要点

  1. 现有的IRT模型在AI评估中面临数据模式不匹配的问题,导致推断结果的可靠性受到质疑。
  2. 本文通过模拟响应矩阵和比较不同的IRT估计工具,探讨了在AI基准测试中使用IRT的有效性。
  3. 实验结果表明,经典估计器在大规模基准中不可行,而可扩展估计器在特定条件下可能导致不可靠的推断。

📝 摘要(中文)

随着AI基准测试越来越多地利用项目级统计模型,特别是项目反应理论(IRT),来估计模型能力、排名系统、选择信息性示例和诊断基准质量。然而,AI基准数据通常与人类测试的数据模式存在差异,这对IRT建模的可靠性提出了挑战。本文通过模拟响应矩阵并比较四种估计工具,系统评估了IRT推断的可靠性和可行性,发现经典估计器在大规模基准设置中可能变得不可行,而可扩展的估计器在小型或非正态分布模型集上可能产生不可靠的推断。研究指出了何时潜在特征模型能够可靠支持AI基准声明,以及需要什么样的样本量和诊断以确保可信使用。

🔬 方法详解

问题定义:本文旨在解决AI评估中使用IRT模型的可靠性问题,现有方法在数据模式不匹配时表现不佳,导致推断结果不可靠。

核心思路:通过模拟响应矩阵并比较不同IRT估计工具,评估其在AI基准测试中的适用性,旨在识别何时IRT模型能够可靠支持AI评估。

技术框架:研究采用了六个广泛使用的LLM基准数据,模拟了三种常见的IRT模型下的响应矩阵,并比较了四种估计工具的性能。

关键创新:本研究的创新在于系统性地评估了IRT模型在AI基准测试中的适用性,揭示了经典与可扩展估计器在不同条件下的表现差异。

关键设计:在实验中,设置了18,000种模拟条件,关注计算可行性、可扩展性及IRT推断的可靠性,特别是模型排名和项目特征的推断。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,经典估计器在大规模基准设置中不可行,而可扩展估计器在小型或非正态分布模型集上可能产生不可靠的推断。这一发现强调了在AI评估中选择合适IRT工具的重要性。

🎯 应用场景

该研究的潜在应用领域包括AI模型的性能评估、基准测试设计以及教育测评等。通过改进IRT模型的应用,可以为AI系统的开发和优化提供更可靠的评估工具,进而提升AI技术的实际应用价值。

📄 摘要(原文)

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or nonnormally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.