(Whose defaults?) Is artificial intelligence reorienting archaeological methods?
作者: Lorenzo Cardarelli, Roberto Ragno
分类: cs.CY, cs.AI, cs.CL, cs.HC
发布日期: 2026-09-10
💡 一句话要点
评估大型语言模型对考古方法多样性的影响
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 考古学 大型语言模型 方法多样性 生成性人工智能 计算研究 贝叶斯模型 推荐系统
📋 核心要点
- 考古学方法的多样性受到大型语言模型影响,但具体影响尚不明确。
- 通过分析考古学文献,识别出方法变化并进行分类,探讨LLMs的推荐效果。
- 实验结果显示,LLMs推荐的方法多样性低于文献,且更倾向于使用2023年前的广泛方法。
📝 摘要(中文)
生成性人工智能和“氛围编码”的实践正在改变考古学家进行计算研究的方式,但其对学科方法多样性的影响仍然缺乏深入研究。本文分析了2010至2025年间约119,000篇考古学摘要,识别出25个广泛类别和241个更细致的计算方法。结果显示,自2023年后方法使用发生了小幅变化,但总体方法多样性反而有所增加。此外,通过控制实验发现,LLMs推荐的方法集相较于考古学家实际使用的更为狭窄,尤其是在缺乏方法指导的情况下。这些结果表明,LLMs可能推动方法选择趋向收敛,但尚无法确定因果关系。
🔬 方法详解
问题定义:本文旨在探讨大型语言模型(LLMs)是否导致考古学方法的多样性下降,现有研究对这一影响尚缺乏系统分析。
核心思路:通过分析大量考古学文献,识别出使用的计算方法,并评估LLMs在推荐方法时的多样性,以此了解其对考古学研究方法的潜在影响。
技术框架:研究首先对119,000篇考古学摘要进行分析,使用本地运行的LLM识别方法,分类为25个广泛类别和241个细分集群。随后进行控制实验,比较LLMs推荐的研究方法与实际文献中的方法多样性。
关键创新:本研究首次系统性地评估了LLMs对考古学方法多样性的影响,揭示了LLMs在推荐方法时的趋同现象,提供了对考古学研究方法未来发展的重要见解。
关键设计:使用贝叶斯Dirichlet-多项式模型分析方法组成,设计了三种不同的提示级别(初学者、中级、专家)来测试LLMs的推荐效果,确保实验的全面性和准确性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,LLMs推荐的方法多样性显著低于考古学文献,尤其在缺乏方法指导的情况下。推荐方法更倾向于2023年前的技术,表明LLMs可能导致方法选择的收敛。
🎯 应用场景
该研究为考古学领域提供了对大型语言模型影响的深入理解,帮助考古学家在日益依赖AI工具的背景下,保持研究方法的多样性。未来,研究成果可用于优化考古学的计算研究方法,促进学科的创新发展。
📄 摘要(原文)
Generative AI and the practice of "vibe coding" are changing how archaeologists carry out computational research, but their effects on the discipline's range of methods is still understudied. In this paper, we evaluate whether large language models (LLMs) are narrowing the variety of methods archaeologists use. We first analysed approximately 119,000 archaeology abstracts from Scopus, covering publications from 2010 to 2025. Using a locally run LLM, we identified the computational methods reported in each abstract and organised them into 25 broad categories (L2) and 241 finer clusters (L3). A Bayesian Dirichlet-multinomial model of method composition within sub-disciplines found a small but credible shift in method use after 2023. However, this shift was smaller than the variation already present across the full study period. No individual technique showed a significant change, and overall methodological diversity increased rather than declined. We then ran a controlled experiment to see whether LLMs recommend a narrower set of methods than archaeologists have used in practice. Two different open-weight models were asked to suggest methods for 28 archaeological research problems, with prompts providing three levels of methodological guidance: novice, intermediate, and expert. Recommendation diversity was much lower than in the published literature, particularly without methodological guidance. The models also tended to favour methods that were widely used before 2023, and their recommendations more closely resembled the post-2023 literature. Taken together, these results are consistent with LLMs pushing methodological choice towards convergence, although our study cannot establish a causal effect. They raise a broader question: how can archaeology retain methodological diversity as LLMs become more involved in research?