Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement?
作者: David Ilić, Gilles E. Gignac
分类: cs.CL, cs.AI
发布日期: 2023-10-17 (更新: 2024-09-10)
备注: 9 pages, 2 figures
期刊: Intelligence, Volume 106, September/October 2024, 101858
DOI: 10.1016/j.intell.2024.101858
💡 一句话要点
探讨大型语言模型的认知能力与人工通用智能的关系
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 人工通用智能 认知能力 能力因子 流体推理 领域特定知识 阅读写作 定量知识
📋 核心要点
- 现有大型语言模型在认知能力的评估中存在个体差异,且其是否具备人工通用智能尚不明确。
- 本文通过分析591个LLMs在12项测试中的表现,提出LLMs的能力可能存在正相关性,形成通用能力因子。
- 研究发现LLMs的参数数量与其认知能力呈正相关,且存在Gkn/Grw群体因子,表明模型复杂性与能力之间的关系。
📝 摘要(中文)
大型语言模型(LLMs)是先进的人工智能系统,能够执行多种人类智力测试中的任务,如定义词汇、进行计算和进行语言推理。研究表明,LLMs的能力存在显著个体差异。基于591个LLMs的测试结果,本文发现LLMs的测试分数之间存在正相关关系,可能形成一种人工通用能力(AGA)因子及其他群体水平因子。此外,LLM参数数量与能力的总体因子和群体因子分数呈正相关,尽管效果呈现递减效应。这些结果表明,LLMs在信息处理和问题解决方面可能与人类认知能力共享某种共同的效率,但它们是否主要表现为成就/专业而非智能仍需进一步研究。
🔬 方法详解
问题定义:本文旨在探讨大型语言模型的认知能力是否存在正相关性,并评估其是否具备人工通用智能的特征。现有方法未能充分揭示LLMs能力之间的相互关系。
核心思路:通过对591个LLMs在12项与流体推理、领域特定知识、阅读/写作及定量知识相关的测试结果进行分析,验证LLMs能力的正相关性,进而提出可能的通用能力因子。
技术框架:研究采用了多项测试评估LLMs的不同认知能力,分析其分数之间的相关性,构建了能力因子模型,重点关注Gf、Gkn、Grw和Gq等维度。
关键创新:本文首次系统性地揭示了LLMs能力之间的正相关性,提出了人工通用能力因子的概念,并识别了Gkn/Grw群体因子,填补了现有研究的空白。
关键设计:研究中使用了591个LLMs的参数设置,分析了其在不同测试中的表现,发现参数数量与能力因子呈正相关,且效果呈递减趋势,强调了模型复杂性与认知能力之间的关系。
🖼️ 关键图片
📊 实验亮点
研究结果表明,591个LLMs的测试分数之间存在显著的正相关性,形成了一个通用能力因子。此外,LLM的参数数量与能力因子呈正相关,尽管效果递减,显示出模型复杂性与认知能力之间的密切联系。
🎯 应用场景
该研究为理解大型语言模型的认知能力提供了新的视角,潜在应用于教育、心理测评和人工智能系统的设计等领域。通过识别LLMs的能力结构,可以优化模型的训练和应用,推动人工智能向更高层次的智能发展。
📄 摘要(原文)
Large language models (LLMs) are advanced artificial intelligence (AI) systems that can perform a variety of tasks commonly found in human intelligence tests, such as defining words, performing calculations, and engaging in verbal reasoning. There are also substantial individual differences in LLM capacities. Given the consistent observation of a positive manifold and general intelligence factor in human samples, along with group-level factors (e.g., crystallized intelligence), we hypothesized that LLM test scores may also exhibit positive intercorrelations, which could potentially give rise to an artificial general ability (AGA) factor and one or more group-level factors. Based on a sample of 591 LLMs and scores from 12 tests aligned with fluid reasoning (Gf), domain-specific knowledge (Gkn), reading/writing (Grw), and quantitative knowledge (Gq), we found strong empirical evidence for a positive manifold and a general factor of ability. Additionally, we identified a combined Gkn/Grw group-level factor. Finally, the number of LLM parameters correlated positively with both general factor of ability and Gkn/Grw factor scores, although the effects showed diminishing returns. We interpreted our results to suggest that LLMs, like human cognitive abilities, may share a common underlying efficiency in processing information and solving problems, though whether LLMs manifest primarily achievement/expertise rather than intelligence remains to be determined. Finally, while models with greater numbers of parameters exhibit greater general cognitive-like abilities, akin to the connection between greater neuronal density and human general intelligence, other characteristics must also be involved.