Theory of Mind in Large Language Models: Examining Performance of 11 State-of-the-Art models vs. Children Aged 7-10 on Advanced Tests

📄 arXiv: 2310.20320v1 📥 PDF

作者: Max J. van Duijn, Bram M. A. van Dijk, Tom Kouwenhoven, Werner de Valk, Marco R. Spruit, Peter van der Putten

分类: cs.CL, cs.AI

发布日期: 2023-10-31

备注: 14 pages, 4 figures, Forthcoming in Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL)


💡 一句话要点

通过对比儿童与LLMs,探讨大语言模型的心智理论能力

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 心智理论 大语言模型 推理能力 非字面语言 递归意图 教育心理学 人机交互

📋 核心要点

  1. 现有研究主要集中在大语言模型的基本能力,缺乏对其心智理论能力的深入探讨。
  2. 本文通过测试多种LLMs在非字面语言和递归意图等复杂任务上的表现,提出了新的评估方法。
  3. 实验结果表明,指令调优的GPT模型在ToM任务上表现优异,超越了其他模型和儿童的表现。

📝 摘要(中文)

本文探讨了大语言模型(LLMs)在心智理论(ToM)能力方面的表现,特别是推理意图和信念的能力。研究测试了11种基础和指令调优的LLMs,超越了传统的错误信念范式,涵盖了非字面语言使用和递归意图等能力。通过重新编写的标准化测试评估LLMs的稳健性,并与7-10岁儿童的表现进行对比。结果显示,GPT系列的指令调优模型在许多任务中超越了其他模型,甚至在某些情况下超越了儿童,而基础LLMs在ToM任务上表现不佳。研究认为语言与ToM的相互发展可能解释了指令调优的优势,强调了考虑对话者和上下文的合作沟通的重要性。

🔬 方法详解

问题定义:本文旨在探讨大语言模型在心智理论(ToM)方面的能力,现有方法主要集中在简单的错误信念任务,未能全面评估模型的复杂推理能力。

核心思路:研究通过设计新的标准化测试,评估LLMs在非字面语言使用和递归意图等方面的表现,提供更全面的ToM能力评估。

技术框架:整体研究框架包括四个主要模块:1) 测试设计,涵盖多种ToM相关任务;2) 模型选择,包含11种基础和指令调优的LLMs;3) 评估方法,采用开放式与封闭式问题的评分;4) 性能对比,基于儿童的表现进行基准测试。

关键创新:最重要的创新在于将ToM能力的评估扩展到非字面语言和递归意图,突破了传统的错误信念范式,提供了更丰富的评估视角。

关键设计:在测试设计中,采用了重新编写的标准化测试,确保任务的适应性和有效性;在模型评估中,特别关注指令调优的效果,强调了对话者和上下文的影响。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,指令调优的GPT模型在多项ToM任务中表现优于其他模型,且在某些任务上超越了7-10岁儿童的表现。这一发现表明,指令调优显著提升了模型的推理能力,尤其是在复杂的非字面语言和意图理解方面。

🎯 应用场景

该研究为理解大语言模型的认知能力提供了新的视角,潜在应用于教育、心理学和人机交互等领域。通过更深入的ToM能力评估,未来可以优化模型在复杂对话和情感理解中的表现,提升人机交互的自然性和有效性。

📄 摘要(原文)

To what degree should we ascribe cognitive capacities to Large Language Models (LLMs), such as the ability to reason about intentions and beliefs known as Theory of Mind (ToM)? Here we add to this emerging debate by (i) testing 11 base- and instruction-tuned LLMs on capabilities relevant to ToM beyond the dominant false-belief paradigm, including non-literal language usage and recursive intentionality; (ii) using newly rewritten versions of standardized tests to gauge LLMs' robustness; (iii) prompting and scoring for open besides closed questions; and (iv) benchmarking LLM performance against that of children aged 7-10 on the same tasks. We find that instruction-tuned LLMs from the GPT family outperform other models, and often also children. Base-LLMs are mostly unable to solve ToM tasks, even with specialized prompting. We suggest that the interlinked evolution and development of language and ToM may help explain what instruction-tuning adds: rewarding cooperative communication that takes into account interlocutor and context. We conclude by arguing for a nuanced perspective on ToM in LLMs.