From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models

📄 arXiv: 2607.22182v1 📥 PDF

作者: Shixin Fang, Jiachen Wo, Wenjuan Qin, Sihang Jiang, Yanghua Xiao

分类: cs.CL, cs.AI

发布日期: 2026-07-24

备注: 34 pages, 5 figures, 20 tables


💡 一句话要点

提出多层次分类法以解决大语言模型能力评估碎片化问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大语言模型 能力评估 多层次分类 认知科学 研究组织 覆盖审核 推理能力 语言能力

📋 核心要点

  1. 现有大语言模型评估主要集中在任务上,导致能力分析碎片化,难以进行有效比较和识别覆盖缺口。
  2. 本文提出了一种基于人类认知科学的多层次能力分类法,涵盖14个能力领域和91个子技能,旨在系统化能力评估。
  3. 通过对31,505篇论文的分析,发现语言-语义能力和推理是研究关注的主要领域,支持了能力的结构化分析。

📝 摘要(中文)

大语言模型(LLM)的评估涵盖多种任务和基准,但现有证据主要围绕任务而非能力,导致跨研究比较困难,能力任务的招募情况不清晰,覆盖缺口难以识别。为此,本文提出了一种多层次的分类法,涵盖14个能力领域和91个子技能,分为原始层、构建层和整合层。能力的定义和组织受到人类认知科学的指导,而非LLM架构。通过对31,505篇相关论文的筛选,本文展示了该分类法的实际应用,支持研究组织、覆盖审核和可测试假设的制定。

🔬 方法详解

问题定义:本文旨在解决大语言模型能力评估的碎片化问题,现有方法多围绕具体任务,缺乏系统化的能力分析框架。

核心思路:提出一种多层次的能力分类法,依据人类认知科学定义和组织能力,强调能力的结构化而非单一任务。

技术框架:该框架分为原始层、构建层和整合层,涵盖14个能力领域和91个子技能,利用多模型注释、共识和仲裁对文献进行标注和映射。

关键创新:最重要的创新在于将分析单位从孤立任务转向结构化能力,提供了更系统的能力评估方法,便于研究组织和覆盖审核。

关键设计:在能力分类中,层级分配基于发展优先级和假设的功能支持,确保能力定义与可观察的模型行为相适应。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,语言-语义能力和推理领域的研究集中度最高,分别占22.3%和21.3%。在14个能力领域中,最常见的子技能在90%的论文中出现,表明能力分类法的有效性和实用性。

🎯 应用场景

该研究的潜在应用领域包括大语言模型的评估、训练和转移学习。通过系统化的能力分类,研究人员可以更有效地识别模型的强项和弱项,从而优化模型设计和应用,推动人工智能领域的进一步发展。

📄 摘要(原文)

Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation limits cross-study comparison, obscures capabilities tasks recruit, and makes coverage gaps difficult to identify. We introduce a multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers. Human cognitive science guides capability definition and organization, not LLM architecture. Layer assignments draw on developmental precedence and hypothesized functional support, while human-origin constructs are adapted to observable model behavior. To demonstrate operational utility, we screened 31,505 papers from ACL, AAAI, ICML, and NeurIPS between 2023 and 2025 and mapped 15,934 LLM-focused papers through multi-model annotation, consensus, and arbitration. Direct research attention concentrated on Language-Semantic Competence (3,551; 22.3%), Reasoning (3,388; 21.3%), Planning and Decision-Making (2,149; 13.5%), and Perception (1,954; 12.3%), whereas six domains appeared in fewer than 2% of papers. Within domains, the most frequent subskill had a median prevalence of 97.9% and appeared in at least 90% of papers in 10 of 14 domains. Language-Semantic Competence and Reasoning formed the highest-volume pair (n = 1,864; 11.7%; lift = 2.47), whereas Theory of Mind and Social Reasoning and Interaction showed the highest lift among pairs with at least 20 co-occurrences (n = 62; lift = 30.84). By shifting the unit of analysis from isolated tasks to structured capabilities, the taxonomy supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.