Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns
作者: Yong Yang, Xiang Guan, Sophie Arheix-Parras, Saeed Ahmadi, Roger Newman-Norlund, Leonardo Bonilha, Christopher Rorden, Julius Fridriksson, Rutvik H. Desai, Srihari Nelakuditi
分类: cs.AI, cs.CL, cs.LG
发布日期: 2026-07-13
备注: 15 pages, 8 figures; supplementary materials (18 pages, 6 sections) included
💡 一句话要点
提出多模态语言模型以重现失语症患者的命名错误模式
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 失语症 多模态语言模型 命名错误 数字双胞胎 神经分类器 临床模拟 个体化治疗
📋 核心要点
- 现有语言模型未能有效模拟失语症患者的命名错误模式,缺乏针对性研究。
- 本文通过对多模态语言模型进行扰动,探索其在命名任务中重现失语症错误的能力。
- 实验结果显示,模型能够在大多数情况下重现失语症患者的错误特征,具有较高的准确性。
📝 摘要(中文)
中风后失语症常导致系统性的命名错误,具有特征性模式,但尚未验证通用语言模型是否能重现这些模式。本文研究了多模态语言模型在不同损伤或控制扰动下是否能重现命名错误类型,并能否再现个体失语症患者的完整错误特征。通过对278名失语症患者的评估,发现六种错误类型在不同参数空间区域以临床可比的比例出现,结果为97.8%的患者在六种错误类别中匹配成功,79.5%的患者在七种类别中匹配成功。这些结果为重现个体失语症错误模式建立了定量框架,表明语言模型有潜力成为中风后失语症患者的数字双胞胎。
🔬 方法详解
问题定义:本文旨在解决通用语言模型在重现失语症患者命名错误模式方面的不足,现有方法未能有效模拟个体患者的错误特征。
核心思路:通过对多模态语言模型LLaVA 1.6进行不同层次、比例和噪声的扰动,探索其对命名错误的重现能力,以期建立个体化的错误模式框架。
技术框架:研究采用了扰动配置评估的方式,首先对278名失语症患者进行命名测试,然后使用经过验证的神经分类器将响应分类为七种错误类型,最后分析不同扰动配置对错误类型的影响。
关键创新:本研究的创新在于首次将多模态语言模型应用于失语症错误模式的重现,且能够在高比例的患者中成功匹配个体错误特征,超越了传统方法的局限。
关键设计:在扰动配置中,研究者调整了模型的层数、噪声比例和施加的噪声量,确保在不同参数空间中探索到能够重现临床可比错误模式的配置。
🖼️ 关键图片
📊 实验亮点
实验结果显示,97.8%的失语症患者在六种错误类别中成功匹配,79.5%的患者在所有七种类别中匹配成功,表明该模型在重现个体错误特征方面具有显著的有效性,且与Monte Carlo基线对比,反映了类别间的结构关系,而非边际重叠。
🎯 应用场景
该研究为失语症的临床评估和治疗提供了新的工具,能够通过数字双胞胎技术帮助医生更好地理解和预测患者的语言能力变化,进而制定个性化的康复方案。未来,该方法有望在其他语言障碍的研究中得到应用。
📄 摘要(原文)
Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language models not designed for clinical simulation can reproduce these patterns remains untested. We investigated (1) whether lesions or controlled perturbations to a multimodal language model can reproduce different types of errors in picture naming, and (2) whether the framework can reproduce the complete error profile of individual persons with aphasia (PWAs). Using LLaVA 1.6, we evaluated perturbation configurations that varied the layer, proportion, and amount of noise applied to model units. We examined 278 PWAs on the Philadelphia Naming Test, classifying responses into seven categories using a validated neural classifier. Six of seven response categories (correct, semantic, mixed, unrelated, neologism, no response errors) emerged at clinically-comparable proportions across distinct parameter space regions, with formal paraphasia being the exception. Searching the perturbation space revealed configurations that reproduced the individual error profile in at least six of seven categories for 97.8% of PWAs and in all seven categories for 79.5% of PWAs. Monte Carlo baselines confirmed that this matching reflects joint inter-category structure rather than marginal overlap. These results establish a quantitative framework for reproducing individual aphasic error patterns in picture naming. They suggest the potential for language models to serve as digital twins of individuals with post-stroke aphasia.