Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity
作者: Pengzhao Lyu, Yeun Joon Kim, Hanlin Xiao, Yingyue Luna Luan
分类: cs.CL, cs.AI
发布日期: 2026-07-24
💡 一句话要点
探讨大型语言模型与人类在创造力评估中的一致性与差异性
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 创造力评估 人类评估 新颖性 上下文信息 模型特定标准 评估一致性
📋 核心要点
- 现有研究表明,LLMs在创造力评估中与人类评估的一致性存在混乱,缺乏系统性分析。
- 本文通过三项研究,识别LLMs创造力评估的标准,并探讨其对评估结果的影响。
- 研究发现,LLMs与人类评估的相关性适中,且在新颖性维度上表现出更强的一致性。
📝 摘要(中文)
尽管大型语言模型(LLMs)在创造力评估中的应用日益增加,但其与人类评估的一致性证据仍然不一,本文探讨了何时及为何其判断与人类判断趋同或背离。通过三项研究和六种广泛使用的LLMs,研究发现LLMs通常依赖于较窄的人类创造力评估标准。在新颖性维度上与人类标准的趋同最强,而在上下文维度上则表现出明显的背离。每种LLM展现出独特的模型特定标准,这些标准在广度上存在显著差异。研究结果有助于解释LLM与人类评估之间的混合证据,表明一致性取决于判断所需的证据和每个模型应用的标准。
🔬 方法详解
问题定义:本文旨在解决大型语言模型在创造力评估中与人类评估标准不一致的问题,现有方法未能充分揭示其评估标准的多样性和局限性。
核心思路:通过对六种LLMs的评估标准进行系统分析,探讨其与人类评估的一致性和差异性,特别是在新颖性和上下文维度上的表现。
技术框架:研究分为三项主要实验:第一项分析LLMs的评估标准,第二项评估LLMs与人类评估的相关性,第三项考察上下文信息对评估结果的影响。
关键创新:本文的创新在于系统性地比较不同LLMs在创造力评估中的标准,揭示了模型特定的评估标准如何影响创造力判断,尤其是在新颖性和上下文维度上的差异。
关键设计:研究中使用了多种评估标准,设计了针对性的问题集,并通过统计分析方法评估LLMs与人类评估之间的相关性和差异性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,LLMs的评估与人类评估的相关性为中等水平,且在新颖性维度上表现出更强的一致性。具体而言,具有更广泛标准的LLMs能够更有效地区分人类认为更具创造性的想法与较少创造性的想法。
🎯 应用场景
该研究的潜在应用领域包括教育、创意产业和人机协作等。通过理解LLMs在创造力评估中的表现,可以更好地选择合适的模型进行创意生成和评估,从而提升创意工作的效率与质量。
📄 摘要(原文)
Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from human judgments. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications. Study 1 showed that LLMs generally relied on a narrower subset of human creativity evaluation standards. Convergence with human standards was strongest in the novelty dimension, whereas divergence was clearest in the contextual dimension, which captures social, market, and reputational information. Moreover, each LLM exhibited distinct, model-specific standards that varied substantially in breadth. These differences in evaluation standards were reflected in actual creativity judgments. Study 2 (N = 1,103 ideas) showed that LLM evaluations were moderately correlated with human evaluations, and individual LLMs with broader standards better distinguished ideas humans judged as more versus less creative. Study 3 (N = 1,195) showed that LLMs were less sensitive to contextual information: such information significantly altered human creativity ratings but left LLM ratings largely unchanged. Together, our findings help explain the mixed evidence on LLM-human alignment, showing that alignment depends on the evidence a judgment demands and the standards each model applies. LLMs may resemble humans when evaluations emphasize intrinsic qualities such as novelty, yet diverge when judgments require contextual information. Selecting an LLM evaluator is therefore a consequential decision: different models, applying different standards, recognize different ideas as creative.