A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment
作者: María Eugenia Curi, Germán Capdehourat, Isabel Amigo, Magdalena Romano, Rosana Serra, Adrián Silveira, Andrés Peri
分类: cs.CL, cs.AI
发布日期: 2026-09-04
备注: 25 pages, 8 figures
💡 一句话要点
提出人机协作框架以提升大规模写作评估的评分效率
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 人工智能 写作评估 人机协作 大型语言模型 教育测评 评分系统 效率提升
📋 核心要点
- 现有的写作评估方法在效率和可扩展性方面存在挑战,人工评分工作量大且容易受到主观因素影响。
- 论文提出了一种AI辅助评分框架,结合人机协作策略,以提高评分效率并保持评估质量。
- 实验结果显示,AI评分与人工评分在大多数维度上具有中到高的一致性,支持AI辅助评分的可行性。
📝 摘要(中文)
本研究探讨了将人工智能(AI),特别是大型语言模型(LLMs),融入教育评估的可能性,以提高评分过程的效率和可扩展性。研究设计并验证了一种AI辅助评分框架,专注于约150-200字的短文本,采用人机协作策略以保持评估质量并减少人工工作量。通过分析来自两次全国性测试的约5000份学生回答的数据,研究发现AI生成的评分与人工评分在多个维度上具有中到高的一致性,表明在此环境中AI辅助评分的可行性。研究结果还指出,AI辅助评分需与精心设计的人类监督相结合,以确保安全整合到大规模评估过程中。
🔬 方法详解
问题定义:本研究旨在解决传统写作评估中人工评分效率低、主观性强的问题。现有方法在大规模评估中难以处理大量学生的写作回答,导致评分过程缓慢且不一致。
核心思路:论文的核心思路是设计一个AI辅助评分框架,结合人机协作,利用AI模型生成初步评分,再由人工评审进行校正,以确保评分的准确性和一致性。
技术框架:整体架构包括数据预处理、AI评分模型、人工校正模块和决策流。首先对学生的写作文本进行预处理,然后使用训练好的AI模型生成评分,最后通过人工评审来确认和调整评分。
关键创新:最重要的技术创新在于引入人机协作的评分流程,AI模型不仅提供初步评分,还通过决策流识别需要人工审查的案例,从而优化专家的工作分配。
关键设计:在模型训练中,采用了多维度评分标准,损失函数设计为考虑评分一致性和准确性,同时在网络结构上使用了适合短文本处理的Transformer架构。
🖼️ 关键图片
📊 实验亮点
实验结果显示,AI生成的评分与人工评分在多个维度上具有中到高的一致性,具体而言,在大多数维度上,AI评分与人工评分的相关系数达到0.7以上,表明AI辅助评分的有效性和可靠性。此外,研究还发现AI辅助评分能够有效识别需要人工审查的案例,从而优化专家的工作分配。
🎯 应用场景
该研究的潜在应用领域包括国家级写作评估、教育测评机构以及在线学习平台。通过引入AI辅助评分,能够显著提高评估效率,减轻教师的工作负担,并在大规模评估中保持评分质量。未来,随着技术的不断发展,该框架可能会扩展到其他类型的评估和反馈系统中。
📄 摘要(原文)
The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational assessment has opened new opportunities to enhance the efficiency and scalability of grading processes. This study presents the design and validation of an AI-assisted scoring framework for written responses in a large-scale national assessment. The proposed approach focuses on short written texts of approximately 150-200 words and incorporates a human-in-the-loop strategy to preserve assessment quality while reducing manual workload. The study is grounded in a real operational context, using data from two recent editions of a nationwide test, each comprising approximately 5,000 student responses. We analyze the alignment between AI-generated scores and human raters across multiple rubric dimensions, as well as the impact of the proposed decision flow on pass/fail outcomes. Results show moderate to high agreement between the model and human evaluations in most dimensions, supporting the feasibility of AI assistance in this setting. Moreover, the proposed correction workflow identifies cases where human review is most valuable, enabling a more efficient allocation of expert effort. The findings suggest that AI-assisted scoring can be safely integrated into large-scale assessment processes only when combined with carefully designed human oversight. The paper concludes by discussing practical implications for deployment in national assessment systems and outlining future research directions, including longitudinal monitoring of model-human alignment and the analysis of potential cognitive bias introduced by AI-supported review workflows.