Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model

📄 arXiv: 2310.15113v2 📥 PDF

作者: Leonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, Anna Cai, Ritam Dutt, Amey Hengle, Anubha Kabra, Atharva Kulkarni, Abhishek Vijayakumar, Haofei Yu, Hinrich Schütze, Kemal Oflazer, David R. Mortensen

分类: cs.CL

发布日期: 2023-10-23 (更新: 2023-10-26)

备注: EMNLP 2023


💡 一句话要点

对ChatGPT的形态学能力进行多语言分析以揭示其局限性

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 形态学能力 多语言分析 Berko测试 语言理解 模型评估

📋 核心要点

  1. 现有研究对大型语言模型的语言能力缺乏系统性分析,尤其是在形态学方面。
  2. 本文通过对ChatGPT进行Berko的wug测试,分析其在四种语言中的形态学能力,填补了这一空白。
  3. 实验结果显示,ChatGPT在形态学任务上的表现显著低于专门设计的系统,尤其是在英语中表现不佳。

📝 摘要(中文)

大型语言模型(LLMs)最近在语言能力上取得了显著进展,但对其语言能力的系统性研究仍然较少。现有研究往往忽视人类的概括能力,仅关注英语,并且主要研究句法或语义,而忽略了形态学等人类语言的核心能力。本文通过对ChatGPT在英语、德语、泰米尔语和土耳其语四种不同类型语言的形态学能力进行首次严格分析,发现其表现远低于专门构建的系统,尤其是在英语方面。我们的结果表明,关于ChatGPT具有人类语言能力的说法是过早且具有误导性的。

🔬 方法详解

问题定义:本文旨在分析ChatGPT在形态学能力上的表现,现有方法未能全面评估大型语言模型的语言能力,尤其是形态学方面的不足。

核心思路:通过应用Berko的wug测试,使用未污染的数据集,对ChatGPT进行多语言的形态学能力评估,旨在揭示其在不同语言中的表现差异。

技术框架:研究首先定义了形态学能力的评估标准,然后设计了实验流程,包括数据集的构建、模型的选择和测试的实施,最后对结果进行分析和比较。

关键创新:本研究首次系统性地评估了ChatGPT在多种语言中的形态学能力,提供了与现有研究不同的视角,强调了形态学在语言能力中的重要性。

关键设计:实验中使用了四种语言的专用数据集,确保数据的纯净性和代表性,采用了标准化的测试方法,以便于结果的比较和分析。实验设计还考虑了不同语言的形态特征,以确保评估的全面性。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果显示,ChatGPT在形态学任务中的表现远低于专门构建的系统,尤其是在英语中,其性能显著低于基线,表明其在语言能力上的局限性。具体数据未提供,但整体表现被认为是令人失望的。

🎯 应用场景

该研究的结果对大型语言模型的开发和评估具有重要的指导意义,尤其是在多语言处理和形态学能力的提升方面。未来,研究者可以基于此研究进一步探索如何改进语言模型的形态学理解能力,从而增强其在实际应用中的表现。

📄 摘要(原文)

Large language models (LLMs) have recently reached an impressive level of linguistic capability, prompting comparisons with human language skills. However, there have been relatively few systematic inquiries into the linguistic capabilities of the latest generation of LLMs, and those studies that do exist (i) ignore the remarkable ability of humans to generalize, (ii) focus only on English, and (iii) investigate syntax or semantics and overlook other capabilities that lie at the heart of human language, like morphology. Here, we close these gaps by conducting the first rigorous analysis of the morphological capabilities of ChatGPT in four typologically varied languages (specifically, English, German, Tamil, and Turkish). We apply a version of Berko's (1958) wug test to ChatGPT, using novel, uncontaminated datasets for the four examined languages. We find that ChatGPT massively underperforms purpose-built systems, particularly in English. Overall, our results -- through the lens of morphology -- cast a new light on the linguistic capabilities of ChatGPT, suggesting that claims of human-like language skills are premature and misleading.