Prevalence and prevention of large language model use in crowd work

📄 arXiv: 2310.15683v1 📥 PDF

作者: Veniamin Veselovsky, Manoel Horta Ribeiro, Philip Cozzolino, Andrew Gordon, David Rothschild, Robert West

分类: cs.CL

发布日期: 2023-10-24

备注: VV and MHR equal contribution. 14 pages, 1 figure, 1 table


💡 一句话要点

研究LLM在众包工作中的使用及其预防策略

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 众包工作 数据质量 使用频率 干预策略

📋 核心要点

  1. 当前众包工作中,LLM的使用普遍存在,且对研究结果的影响尚未得到充分重视。
  2. 研究提出通过提高使用成本和明确禁止使用LLM的方式来减少其使用频率。
  3. 实验结果显示,LLM使用率从约30%降低至15%,但同时也影响了摘要中关键信息的呈现。

📝 摘要(中文)

本研究表明,大型语言模型(LLMs)在众包工作者中广泛使用,且针对性的缓解策略可以显著减少但无法完全消除LLM的使用。在未对工作者进行任何指导的文本摘要任务中,LLM使用的估计普遍率约为30%,通过要求工作者不使用LLM并提高使用成本(例如禁用复制粘贴),这一比例减少了约一半。进一步分析揭示了LLM使用的高质量但同质化的响应可能对关注人类行为的研究造成危害,同时阻止LLM的使用可能与获取高质量响应的目标相悖。理解LLM工具与用户的共同演化对于维持众包研究的有效性至关重要。

🔬 方法详解

问题定义:本研究旨在解决大型语言模型(LLMs)在众包工作中普遍使用的问题,现有方法未能有效控制LLM的影响,可能导致研究结果的偏差。

核心思路:通过实施针对性的干预措施,如要求工作者不使用LLM和提高使用成本,来减少LLM的使用频率,从而提高数据的多样性和质量。

技术框架:研究设计了一个文本摘要任务,工作者在未被指导的情况下完成任务,随后实施干预措施并进行效果评估。主要模块包括任务设计、干预实施和效果分析。

关键创新:本研究的创新点在于系统性地评估了LLM在众包工作中的使用情况及其对研究结果的影响,提出了有效的干预策略,并揭示了LLM使用与数据质量之间的权衡。

关键设计:在实验中,通过禁用复制粘贴等方式提高LLM使用的成本,同时对工作者进行明确的使用指导,以此来观察其对LLM使用率和摘要质量的影响。

🖼️ 关键图片

fig_0
img_1

📊 实验亮点

实验结果显示,未干预的情况下LLM使用率约为30%,而在实施干预后,该比例降低至15%。然而,干预措施也导致摘要中关键信息的减少,表明在提高数据质量与获取高质量响应之间存在权衡。

🎯 应用场景

该研究的潜在应用领域包括社会科学研究、市场调研和在线内容生成等。通过理解LLM的使用情况及其影响,研究者可以更好地设计众包任务,确保数据质量,从而提升研究的有效性和可靠性。未来,随着LLM技术的不断发展,该研究的发现将为众包平台的设计和管理提供重要参考。

📄 摘要(原文)

We show that the use of large language models (LLMs) is prevalent among crowd workers, and that targeted mitigation strategies can significantly reduce, but not eliminate, LLM use. On a text summarization task where workers were not directed in any way regarding their LLM use, the estimated prevalence of LLM use was around 30%, but was reduced by about half by asking workers to not use LLMs and by raising the cost of using them, e.g., by disabling copy-pasting. Secondary analyses give further insight into LLM use and its prevention: LLM use yields high-quality but homogeneous responses, which may harm research concerned with human (rather than model) behavior and degrade future models trained with crowdsourced data. At the same time, preventing LLM use may be at odds with obtaining high-quality responses; e.g., when requesting workers not to use LLMs, summaries contained fewer keywords carrying essential information. Our estimates will likely change as LLMs increase in popularity or capabilities, and as norms around their usage change. Yet, understanding the co-evolution of LLM-based tools and users is key to maintaining the validity of research done using crowdsourcing, and we provide a critical baseline before widespread adoption ensues.