Experimental Narratives: A Comparison of Human Crowdsourced Storytelling and AI Storytelling
作者: Nina Begus
分类: cs.CL, cs.AI
发布日期: 2023-10-19 (更新: 2024-11-04)
期刊: Humanities and Social Sciences Communications 11: 1392 (2024)
DOI: 10.1057/s41599-024-03868-8
💡 一句话要点
提出一种框架以比较人类与AI的叙事能力
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 叙事学 人工智能 生成模型 社会偏见 文化研究 人机交互 故事创作
📋 核心要点
- 现有的叙事研究多集中于人类创作,缺乏对AI生成叙事的系统比较,难以揭示两者的异同。
- 本文提出的框架结合了行为实验与计算实验,通过相同的虚构提示,系统比较人类与AI的叙事能力。
- 实验结果显示,AI生成的故事在性别角色和性取向方面更为前卫,而人类故事在创意和修辞上更具优势。
📝 摘要(中文)
本文提出了一种结合行为实验和计算实验的框架,利用虚构提示作为研究人类和生成AI叙事中的文化产物和社会偏见的新工具。研究分析了2019年由众包工作者创作的250个故事和2023年由GPT-3.5及GPT-4生成的80个故事。通过叙事学和推论统计的方法,比较了人类与大型语言模型在相同提示下的叙事反应。结果表明,AI生成的叙事在性别角色和性取向方面更具进步性,而人类叙事则在想象力和修辞上更为丰富。该框架为理解人类与AI的集体想象和社会维度提供了新的视角。
🔬 方法详解
问题定义:本文旨在解决人类与AI在叙事能力上的比较研究不足的问题。现有方法未能有效揭示两者在叙事内容和风格上的差异。
核心思路:通过设计相同的虚构提示,结合行为实验和计算实验,直接比较人类和大型语言模型(LLM)在叙事创作中的表现,以揭示文化和社会偏见。
技术框架:研究分为几个主要阶段:首先,收集人类创作的故事和AI生成的故事;其次,应用叙事学和推论统计方法进行分析;最后,比较两者在叙事内容、性别角色和创新性方面的差异。
关键创新:该研究的创新点在于提出了一种新的实验范式,能够直接且可控地比较人类与AI生成的叙事,填补了现有研究的空白。
关键设计:在实验中,使用了相同的虚构提示,确保了对比的公平性;同时,采用了叙事学和推论统计的结合方法,以增强分析的深度和广度。具体的参数设置和模型选择在研究中进行了详细描述。
🖼️ 关键图片
📊 实验亮点
实验结果表明,GPT-3.5和GPT-4生成的故事在性别角色和性取向方面表现出更高的进步性,尤其是GPT-4的表现更为突出。同时,AI叙事在创新情节方面虽有优势,但在想象力和修辞上仍不及人类创作的故事。
🎯 应用场景
该研究的潜在应用领域包括教育、文学创作和人机交互等。通过理解人类与AI在叙事上的异同,可以为未来的AI创作工具提供设计依据,促进更具人性化的AI应用发展。
📄 摘要(原文)
The paper proposes a framework that combines behavioral and computational experiments employing fictional prompts as a novel tool for investigating cultural artifacts and social biases in storytelling both by humans and generative AI. The study analyzes 250 stories authored by crowdworkers in June 2019 and 80 stories generated by GPT-3.5 and GPT-4 in March 2023 by merging methods from narratology and inferential statistics. Both crowdworkers and large language models responded to identical prompts about creating and falling in love with an artificial human. The proposed experimental paradigm allows a direct and controlled comparison between human and LLM-generated storytelling. Responses to the Pygmalionesque prompts confirm the pervasive presence of the Pygmalion myth in the collective imaginary of both humans and large language models. All solicited narratives present a scientific or technological pursuit. The analysis reveals that narratives from GPT-3.5 and particularly GPT-4 are more progressive in terms of gender roles and sexuality than those written by humans. While AI narratives with default settings and no additional prompting can occasionally provide innovative plot twists, they offer less imaginative scenarios and rhetoric than human-authored texts. The proposed framework argues that fiction can be used as a window into human and AI-based collective imaginary and social dimensions.