LAiW: A Chinese Legal Large Language Models Benchmark

📄 arXiv: 2310.05620v2 📥 PDF

作者: Yongfu Dai, Duanyu Feng, Jimin Huang, Haochen Jia, Qianqian Xie, Yifang Zhang, Weiguang Han, Wei Tian, Hao Wang

分类: cs.CL

发布日期: 2023-10-09 (更新: 2024-02-18)


💡 一句话要点

构建LAiW基准以评估中文法律大语言模型的实用能力

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 法律大语言模型 评估基准 法律人工智能 信息检索 法律推理 复杂应用 人工评估

📋 核心要点

  1. 现有法律大语言模型的评估缺乏与法律实践逻辑的一致性,难以反映其实际能力。
  2. 本文提出LAiW基准,依据法律实践逻辑将LLMs的能力分为三个层次,以全面评估其性能。
  3. 通过自动评估和法律专家的人工评估,发现LLMs在复杂法律应用能力上表现良好,但在基础任务上存在不足。

📝 摘要(中文)

一般和法律领域的大语言模型(LLMs)在法律人工智能的多项任务中表现出色。然而,目前对这些LLMs的评估缺乏与法律实践逻辑的一致性,导致难以判断其实际能力。为了解决这一挑战,本文首次构建了基于法律实践逻辑的中文法律LLMs基准LAiW。我们将LLMs的法律能力分为三个层次:基本信息检索、法律基础推理和复杂法律应用。通过对当前LLMs的自动评估,结果表明这些模型在某些基本任务上表现不佳,可能影响其在法律领域的实际应用和接受度。为进一步确认LLMs在法律应用场景中的复杂法律应用能力,我们还结合了法律专家的人工评估,结果显示LLMs在表现上仍需加强法律逻辑。

🔬 方法详解

问题定义:本文旨在解决当前法律大语言模型评估缺乏与法律实践逻辑一致性的问题,现有方法无法有效判断模型的实际能力。

核心思路:通过构建LAiW基准,依据法律实践的思维过程,将法律能力分为三个层次,确保评估的全面性和准确性。

技术框架:整体架构包括三个层次的任务:基本信息检索、法律基础推理和复杂法律应用,每个层次下又包含多个具体任务,以便进行系统评估。

关键创新:LAiW基准是首个基于法律实践逻辑构建的中文法律LLMs评估工具,强调了法律逻辑在模型评估中的重要性。

关键设计:在评估过程中,设置了多样化的任务和标准,确保模型在不同层次的表现能够被准确衡量,同时结合了法律专家的人工评估以增强结果的可信度。

🖼️ 关键图片

fig_0
fig_1

📊 实验亮点

实验结果表明,当前的法律大语言模型在复杂法律应用任务中表现良好,但在基本信息检索等基础任务上却存在明显不足。通过与法律专家的评估结合,进一步确认了模型在法律逻辑强化方面的需求。

🎯 应用场景

该研究的潜在应用领域包括法律咨询、合同审查、法律文书生成等,能够为法律从业者提供更为精准和高效的工具。随着法律人工智能技术的发展,LAiW基准将推动法律大语言模型的实际应用和行业接受度,促进法律服务的智能化转型。

📄 摘要(原文)

General and legal domain LLMs have demonstrated strong performance in various tasks of LegalAI. However, the current evaluations of these LLMs in LegalAI are defined by the experts of computer science, lacking consistency with the logic of legal practice, making it difficult to judge their practical capabilities. To address this challenge, we are the first to build the Chinese legal LLMs benchmark LAiW, based on the logic of legal practice. To align with the thinking process of legal experts and legal practice (syllogism), we divide the legal capabilities of LLMs from easy to difficult into three levels: basic information retrieval, legal foundation inference, and complex legal application. Each level contains multiple tasks to ensure a comprehensive evaluation. Through automated evaluation of current general and legal domain LLMs on our benchmark, we indicate that these LLMs may not align with the logic of legal practice. LLMs seem to be able to directly acquire complex legal application capabilities but perform poorly in some basic tasks, which may pose obstacles to their practical application and acceptance by legal experts. To further confirm the complex legal application capabilities of current LLMs in legal application scenarios, we also incorporate human evaluation with legal experts. The results indicate that while LLMs may demonstrate strong performance, they still require reinforcement of legal logic.