Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages

📄 arXiv: 2310.17526v2 📥 PDF

作者: Qusai Khraisha, Sophie Put, Johanna Kappenberg, Azza Warraitch, Kristin Hadfield

分类: cs.CL, cs.AI, cs.LG

发布日期: 2023-10-26 (更新: 2023-10-27)

备注: 9 pages, 2 figures, 1 table

DOI: 10.1002/jrsm.1715


💡 一句话要点

评估GPT-4在系统评价过程中的应用潜力

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 系统评价 GPT-4 数据提取 文献筛选 多语言处理 人工智能

📋 核心要点

  1. 系统评价过程通常耗时且劳动密集,现有方法效率低下,难以满足快速决策的需求。
  2. 本研究采用GPT-4评估其在系统评价中的表现,探索其在多语言文献筛选和数据提取中的应用潜力。
  3. 实验结果显示,GPT-4在特定任务中表现出与人类相当的准确性,但在数据提取和筛选阶段存在一定的局限性。

📝 摘要(中文)

系统评价对于指导实践、研究和政策至关重要,但通常耗时且劳动密集。大型语言模型(LLMs)可能加速和自动化系统评价,但其在此类任务中的表现尚未得到全面评估。本研究评估了GPT-4在标题/摘要筛选、全文审查和数据提取中的能力,采用“人类不参与”的方法。尽管GPT-4在大多数任务中的准确性与人类相当,但结果受到偶然一致性和数据集不平衡的影响。经过调整后,数据提取的表现中等,而在使用高度可靠提示的情况下,全文筛选的表现几乎完美。研究结果表明,目前在进行系统评价时应谨慎使用LLMs,但在某些任务中,LLMs的表现可以与人类相媲美。

🔬 方法详解

问题定义:本研究旨在解决大型语言模型在系统评价过程中的有效性问题,现有方法在效率和准确性上存在不足,尤其是在处理多语言文献时的挑战。

核心思路:通过评估GPT-4在标题/摘要筛选、全文审查和数据提取中的表现,探索其在系统评价中的应用潜力,采用“人类不参与”的方法进行评估。

技术框架:研究分为三个主要阶段:1) 标题/摘要筛选;2) 全文审查;3) 数据提取。每个阶段均使用不同的提示策略进行测试,以评估GPT-4的表现。

关键创新:本研究首次全面评估了GPT-4在系统评价中的表现,特别是在多语言文献处理方面,揭示了其在特定任务中与人类相当的能力。

关键设计:在实验中,使用了高度可靠的提示来优化GPT-4的表现,并通过调整数据集不平衡和偶然一致性来提高结果的可靠性。

🖼️ 关键图片

fig_0
fig_1

📊 实验亮点

实验结果显示,GPT-4在数据提取任务中的表现为中等水平,而在使用高度可靠提示进行全文筛选时,其准确性几乎达到了完美。通过对漏检关键研究的惩罚,GPT-4的表现得到了进一步提升,表明其在特定条件下能够与人类表现相媲美。

🎯 应用场景

该研究的潜在应用领域包括医学、社会科学和政策研究等需要进行系统评价的领域。通过提高系统评价的效率,LLMs有望加速文献综述的过程,帮助研究人员和决策者更快地获取关键信息,进而推动科学研究和政策制定的进展。

📄 摘要(原文)

Systematic reviews are vital for guiding practice, research, and policy, yet they are often slow and labour-intensive. Large language models (LLMs) could offer a way to speed up and automate systematic reviews, but their performance in such tasks has not been comprehensively evaluated against humans, and no study has tested GPT-4, the biggest LLM so far. This pre-registered study evaluates GPT-4's capability in title/abstract screening, full-text review, and data extraction across various literature types and languages using a 'human-out-of-the-loop' approach. Although GPT-4 had accuracy on par with human performance in most tasks, results were skewed by chance agreement and dataset imbalance. After adjusting for these, there was a moderate level of performance for data extraction, and - barring studies that used highly reliable prompts - screening performance levelled at none to moderate for different stages and languages. When screening full-text literature using highly reliable prompts, GPT-4's performance was 'almost perfect.' Penalising GPT-4 for missing key studies using highly reliable prompts improved its performance even more. Our findings indicate that, currently, substantial caution should be used if LLMs are being used to conduct systematic reviews, but suggest that, for certain systematic review tasks delivered under reliable prompts, LLMs can rival human performance.