Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

📄 arXiv: 2607.21340v1 📥 PDF

作者: Prerit Ahuja

分类: cs.CL

发布日期: 2026-07-23

备注: 23 pages. Resubmission of submit/7557765, which expired due to an arXiv system bug (confirmed in support ticket AH-199019); the overfull-box correction requested by moderators has been applied. Original submission held in moderation since 14 May 2026. Therefore, request priority review/approval


💡 一句话要点

提出CM-LRS以解决资本市场文档可银行性问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 资本市场 大型语言模型 文档生成 可靠性评分 合规性 信息检索 数值一致性

📋 核心要点

  1. 现有方法未能全面评估大型语言模型在资本市场文档生成中的可银行性,主要集中在表面准确性。
  2. CM-LRS通过七个维度评估工作流输出,提供了一种新的可靠性评分方法,旨在提升文档的可辩护性。
  3. 实验结果显示,封闭源模型在CM-LRS评分上优于开放权重模型,尤其在信息检索和综合方面表现突出。

📝 摘要(中文)

在资本市场工作流程中,关键问题不在于大型语言模型能否生成流畅的草稿,而在于该草稿是否具备可银行性,即在对方或监管机构面前是否具有可辩护性。现有方法仅部分解决了这一问题,CM-LRS(资本市场LLM可靠性评分)通过七个维度评估工作流输出,包括事实准确性、证据可追溯性、数值一致性等。本文展示了CM-LRS在五个工作流上的应用,结果表明,封闭源模型在CM-LRS评分上表现优于开放权重基线,尤其在检索和综合方面存在显著差距。

🔬 方法详解

问题定义:本文旨在解决资本市场工作流中大型语言模型生成文档的可银行性问题。现有方法主要关注表面准确性,未能有效评估文档在实际应用中的可靠性和可辩护性。

核心思路:CM-LRS通过七个维度(如事实准确性、证据可追溯性等)对工作流输出进行评估,提供了一种更全面的可靠性评分方法,以满足资本市场的实际需求。

技术框架:CM-LRS的整体架构包括七个评估维度,每个维度根据特定的评分标准进行打分,最终得出一个综合评分。该评分可以根据具体工作流进行调节。

关键创新:CM-LRS的主要创新在于其评估维度的设计,特别是将评估重点放在工作流输出层而非单一的问答层,这使得评分更符合实际应用场景的需求。

关键设计:每个评估维度的评分范围为0-5,依据监管环境中评审者使用的信号进行评分,确保评分的科学性和实用性。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果显示,封闭源模型在CM-LRS评分上表现优于开放权重基线,具体而言,Sonnet 4.6和Opus 4.7的平均评分分别为4.31和4.30,而Llama 3.3 70B仅为3.15。检索和综合的评分差距尤为明显,分别为2.23和2.15。

🎯 应用场景

CM-LRS可广泛应用于资本市场的文档生成和审核流程,帮助金融机构提高文档的合规性和可辩护性。未来,该方法可能在其他行业的文档生成和审核中发挥重要作用,提升整体工作效率和合规性。

📄 摘要(原文)

In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft, but whether the draft is bankable: defensible in front of a counter-party or a regulator, with the documents in hand. Existing methods address parts of that gap: open-domain QA benchmarks reward surface accuracy, and finance benchmarks (FinanceBench, FinQA, ConvFinQA) advance document-grounded and numerical QA but evaluate at the question-answer layer rather than the workflow outputs practitioners defend. We introduce CM-LRS, a Capital Markets LLM Reliability Score, evaluating outputs at the workflow-output layer across seven dimensions: factual accuracy, evidence traceability, numerical consistency, workflow completeness, source discipline, decision usefulness, and reviewability/auditability. Each is scored 0-5 against a rubric anchored on signals reviewers in regulated settings use; the aggregate is tunable to the workflow. We demonstrate CM-LRS on five workflows (DCM transaction-terms extraction, precedent retrieval, issuer profile synthesis, M&A transaction-comparable reasoning, ECM transaction-terms extraction) over public SEC EDGAR filings, a public UK takeover release, and fictional synthetic supplements, scoring four models against four independent LLM judges spanning three model families. Three findings. First, the frontier closed-source models cluster within 0.22 points on four-judge averaged CM-LRS (Sonnet 4.6 = 4.31, Opus 4.7 = 4.30, GPT-5.5 = 4.09); all four judges place the open-weights baseline (Llama 3.3 70B = 3.15) last. Second, that gap concentrates on retrieval (2.23) and synthesis (2.15), not extraction (0.84). Third, Decision Usefulness shows the widest cross-model dispersion of any dimension (4.0 points on issuer profiling) and top-tier inter-judge agreement (mean r = 0.52). Plausibility is cheap. Bankability is the bar.