LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings

📄 arXiv: 2607.24435v1 📥 PDF

作者: Brittany Harbison, Ashok K. Goel

分类: cs.CL, cs.AI

发布日期: 2026-07-27


💡 一句话要点

提出LEX-EC框架以解决大语言模型个性分类的可解释性问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 个性分类 可解释性 黑箱模型 文本分析 词汇消融 心理学 社交媒体

📋 核心要点

  1. 现有的大语言模型在个性分类中缺乏可解释性,导致难以理解模型的决策过程。
  2. LEX-EC框架通过结合流行度、一致性诊断和词汇消融,提供了一种新的黑箱审计方法。
  3. 实验结果显示,不同文本类型在个性特征上表现出显著差异,且某些文本特征的掩蔽影响了模型的分类能力。

📝 摘要(中文)

大型语言模型能够从文本中轻松分配个性标签,但模型的可解释性仍然是一个未解决的问题。为此,本文提出了LEX-EC,一个可重用的黑箱审计框架,结合了流行度和一致性诊断与受控的词汇消融,以区分边际分布效应与在限制证据下可恢复的特征相关信号。通过该框架,我们展示了不同文本类型可能表现出显著不同的个性特征。研究表明,自由形式的论文文本包含最广泛但仍然较弱的信号,而研究生自我介绍中的外向性关联在掩蔽后减弱,单条Facebook状态则几乎没有稳定证据,显示出内容或长度的可能下限。掩蔽主题和人口内容削弱了一些关联,但仍能从功能词、情感词和认知风格词汇中检测到其他关联。

🔬 方法详解

问题定义:本文旨在解决大型语言模型在个性分类中的可解释性问题,现有方法在理解模型决策过程时存在不足,尤其是在黑箱设置下。

核心思路:论文提出的LEX-EC框架通过结合流行度和一致性诊断与受控的词汇消融,能够有效区分边际分布效应与特征相关信号,从而提高模型的可解释性。

技术框架:LEX-EC框架包括多个模块:首先进行文本的流行度和一致性分析,然后实施词汇消融以评估特征信号,最后通过模型生成的解释进行综合评估。

关键创新:LEX-EC的创新在于其将词汇方法应用于黑箱可解释性,能够系统性地评估个性标签的分类效果与文本证据的关系,填补了现有方法的空白。

关键设计:框架中采用了多种文本类型进行实验,包括自由形式的论文、研究生自我介绍和社交媒体状态,设计了不同的掩蔽策略以评估其对个性特征关联的影响。

🖼️ 关键图片

fig_0
img_1
img_2

📊 实验亮点

实验结果表明,LEX-EC框架能够有效区分不同文本类型的个性特征,尤其是在研究生自我介绍中,外向性关联在掩蔽后显著减弱。此外,单条Facebook状态几乎没有稳定证据,显示出内容和长度的下限。

🎯 应用场景

该研究的潜在应用领域包括心理学研究、社交媒体分析和人机交互等。通过提高模型的可解释性,LEX-EC能够帮助研究人员和开发者更好地理解和优化个性分类系统,进而提升用户体验和信任度。

📄 摘要(原文)

Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining prevalence and agreement diagnostics with controlled lexical ablation to distinguish marginal-distribution effects from trait-associated signal recoverable under restricted evidence. Using this framework, we illustrate how various text genres may exhibit sharply different profiles: free-form essay text contains the broadest, but still weak, signal; in graduate student introductions, an observable Extraversion association weakened after masking; and single Facebook statuses yield little stable evidence even in a trait-balanced sample, indicating a possible lower bound of content or length. Masking topical and demographic content weakened some associations while leaving others detectable from function words, affective terms, and cognitive-style vocabulary. Linguistic prompting shifted model self-explanations but did not eliminate topical content. LEX-EC jointly evaluates classification prevalence, item-level association, chance-corrected agreement, persistence under lexical restriction, and prompt sensitivity in model-generated explanations. Across datasets, models, and prompts, LEX-EC characterizes how trait associations may vary with available lexical evidence, introducing a novel application of lexical methods to black-box interpretability in personality labeling.