Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI
作者: Deyu Jing
分类: cs.AI, cs.DL
发布日期: 2026-07-23
备注: 33 pages, 8 tables. This paper proposes a normative and infrastructural framework for traceable, AI-assisted humanistic research and presents an auditable Kant case study
💡 一句话要点
提出可追溯学术研究以解决生成性AI带来的信任问题
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 可追溯性 生成性AI 人文学科 学术研究 知识基础设施
📋 核心要点
- 现有的生成性AI技术在学术研究中面临信任危机,缺乏明确的来源和证据支持。
- 论文提出可追溯学术研究的概念,通过引入页锚和其他工具,确保学术解释的透明性和可靠性。
- 案例研究表明,可追溯性在检索纠正和证据评估中具有显著的支持作用,提升了研究的可信度。
📝 摘要(中文)
生成性AI使大型语言模型能够在几秒钟内生成看似学术的文本,但流畅性并不等同于有效解释。最深层的风险不仅在于事实错误,还在于没有明确来源、页码、版本或证据的情况下,解释似乎已经确立。本文提出可追溯学术研究作为AI辅助人文学科研究的最低规范条件,涵盖印刷、数字和生成性AI三次知识基础设施革命。我们引入了页锚、双页码、引用优先生成、无证据、人类验证、四级合规和范围合同,并展示了AIH-Infra作为三层参考实现的结构。通过对29卷《康德学术版》的案例研究,说明了可追溯性如何支持检索纠正、证据评分和判断降级。可追溯性不是软件特性,而是在生成性AI时代人文学科研究保持公开和可反驳的条件。
🔬 方法详解
问题定义:论文要解决的问题是生成性AI生成的文本缺乏可追溯性,导致学术研究的信任危机。现有方法未能提供清晰的来源和证据,影响了研究的有效性和可靠性。
核心思路:论文的核心解决思路是提出可追溯学术研究,利用页锚和双页码等工具,确保学术解释能够追溯到原始来源,从而增强研究的透明度和可信度。
技术框架:整体架构包括三个主要模块:Contexture(文档结构化)、Open WebUI AIH-Infra(可追溯知识库)、AIH-Infra MCP Server(代理网关)。这些模块共同构成了一个支持可追溯性的知识基础设施。
关键创新:最重要的技术创新点在于引入了页锚和双页码的概念,使得学术研究能够在生成性AI的环境中保持可追溯性。这与现有方法的本质区别在于强调了来源的透明性,而不仅仅是生成文本的流畅性。
关键设计:在设计中,采用了四级合规标准和范围合同,确保生成的内容符合学术规范。此外,实施了人类验证机制,以进一步提升生成内容的可信度。具体的参数设置和损失函数设计尚未详细披露。
🖼️ 关键图片
📊 实验亮点
案例研究表明,通过引入可追溯性机制,检索纠正和证据评分的准确性显著提升,具体的性能数据尚未披露,但研究强调了可追溯性在学术研究中的重要性,提升了研究的整体可信度。
🎯 应用场景
该研究的潜在应用领域包括人文学科的学术研究、教育和知识管理等。通过确保学术解释的可追溯性,研究能够提升学术交流的透明度和可信度,促进更为严谨的学术环境。未来,该方法可能会影响生成性AI在学术领域的广泛应用,推动学术研究的规范化。
📄 摘要(原文)
Generative AI lets large language models produce scholarly-looking text within seconds, yet fluency does not equal valid explanation. The deepest risk is not factual error alone but the appearance that an explanation is already established without clear sources, page numbers, editions, or evidence. We liken the page anchor to Ariadne's thread: within the labyrinth of generative fluency, it is the thread that leads the scholar back to the source. This paper proposes Traceable Scholarship as the minimum normative condition for AI-assisted humanistic research, situating it across the three revolutions of knowledge infrastructure: print, digital, and generative AI. We introduce page anchors, dual page numbers, citation-first generation, NO_EVIDENCE, human verification, four-level compliance, and Scope Contract, and present AIH-Infra as a three-layer reference implementation: Contexture (document structuring), Open WebUI AIH-Infra (traceable knowledge base), and AIH-Infra MCP Server (agent gateway). A case study on a 29-volume Kant Akademie-Ausgabe knowledge base illustrates how traceability supports retrieval correction, evidence grading, and judgment downgrading. Traceability is not a software feature; it is the condition under which humanistic research can remain public and refutable in the age of generative AI.