Large Language Model Prediction Capabilities: Evidence from a Real-World Forecasting Tournament
作者: Philipp Schoenegger, Peter S. Park
分类: cs.CY, cs.AI, cs.CL, cs.LG
发布日期: 2023-10-17
备注: 13 pages, six visualizations (four figures, two tables)
💡 一句话要点
评估大型语言模型在真实世界预测中的能力
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 概率预测 真实世界应用 预测比赛 人工智能评估
📋 核心要点
- 现有研究对大型语言模型在真实世界预测中的能力了解不足,尤其是其概率预测的准确性。
- 本研究通过将GPT-4参与真实的预测比赛,实证检验其在多样化主题下的预测能力。
- 实验结果表明,GPT-4的预测准确性显著低于人类群体中位数,且与随机预测策略相似。
📝 摘要(中文)
准确预测未来是人工智能能力的重要里程碑。然而,关于大型语言模型提供未来事件概率预测的研究仍处于初级阶段。为实证测试这一能力,研究者将OpenAI的GPT-4模型参与了为期三个月的预测比赛。比赛涵盖了多个主题,包括科技、美国政治、病毒爆发和乌克兰冲突。结果显示,GPT-4的概率预测显著低于人群中位数预测,且其预测与随机猜测的结果相似。这表明GPT-4在真实世界预测任务中的表现显著不如人类群体预测,提示未来在此领域的研究方向。
🔬 方法详解
问题定义:本研究旨在评估大型语言模型(如GPT-4)在真实世界预测中的能力,尤其是其概率预测的准确性。现有方法在此领域的研究较少,缺乏实证数据支持。
核心思路:通过将GPT-4参与为期三个月的预测比赛,研究者可以直接比较其预测结果与人类参与者的表现,从而评估其在复杂和未知环境中的预测能力。
技术框架:研究采用Metaculus平台进行比赛,涵盖多个主题,参与者包括843名人类预测者。比赛主要集中在二元预测上,评估模型的概率预测能力。
关键创新:本研究的创新在于将大型语言模型置于真实的预测环境中进行评估,而非传统的基准测试。这种方法能够更真实地反映模型在面对未知时的表现。
关键设计:研究中,GPT-4的预测结果与人类中位数预测进行比较,发现其预测结果与50%的随机猜测相似,未能显示出有效的概率预测能力。
🖼️ 关键图片
📊 实验亮点
实验结果显示,GPT-4的概率预测准确性显著低于人类群体中位数,且其预测与随机猜测的结果相似。这表明GPT-4在真实世界预测任务中的表现不佳,提示未来研究需关注模型的推理和预测能力。
🎯 应用场景
该研究为大型语言模型在真实世界预测中的应用提供了重要的实证数据,揭示了其在复杂和动态环境中的局限性。这一发现对未来模型的改进和应用具有指导意义,尤其是在需要高准确性预测的领域,如金融市场、公共卫生和政策制定等。
📄 摘要(原文)
Accurately predicting the future would be an important milestone in the capabilities of artificial intelligence. However, research on the ability of large language models to provide probabilistic predictions about future events remains nascent. To empirically test this ability, we enrolled OpenAI's state-of-the-art large language model, GPT-4, in a three-month forecasting tournament hosted on the Metaculus platform. The tournament, running from July to October 2023, attracted 843 participants and covered diverse topics including Big Tech, U.S. politics, viral outbreaks, and the Ukraine conflict. Focusing on binary forecasts, we show that GPT-4's probabilistic forecasts are significantly less accurate than the median human-crowd forecasts. We find that GPT-4's forecasts did not significantly differ from the no-information forecasting strategy of assigning a 50% probability to every question. We explore a potential explanation, that GPT-4 might be predisposed to predict probabilities close to the midpoint of the scale, but our data do not support this hypothesis. Overall, we find that GPT-4 significantly underperforms in real-world predictive tasks compared to median human-crowd forecasts. A potential explanation for this underperformance is that in real-world forecasting tournaments, the true answers are genuinely unknown at the time of prediction; unlike in other benchmark tasks like professional exams or time series forecasting, where strong performance may at least partly be due to the answers being memorized from the training data. This makes real-world forecasting tournaments an ideal environment for testing the generalized reasoning and prediction capabilities of artificial intelligence going forward.