FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment

📄 arXiv: 2604.23786 📥 PDF

作者: Sophie Chiang, Tom Brennan, Fethiye Irmak Dogan, Jiaee Cheong, Hatice Gunes

分类: cs.AI, cs.LG

发布日期: 2026-07-20


💡 一句话要点

提出FAIR_XAI以解决多模态模型公平性问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 多模态学习 公平性 可解释人工智能 心理健康评估 视觉-语言模型 偏见检测 临床应用

📋 核心要点

  1. 现有多模态模型在心理健康评估中的应用面临透明性不足和偏见问题,影响其临床有效性。
  2. 本研究提出了一种XAI干预框架,旨在提高VLMs在心理健康评估中的公平性和可解释性。
  3. 实验结果显示,尽管干预措施在某些方面提高了公平性,但也暴露了程序透明性与结果公平性之间的差距。

📝 摘要(中文)

近年来,多模态机器学习在心理健康监测中的应用潜力巨大。然而,随着视觉-语言模型(VLMs)的快速发展,其在临床环境中的应用引发了透明性不足和潜在偏见的担忧。尽管已有研究探讨了公平性与可解释人工智能(XAI)的交集,但其在VLMs用于心理健康评估和抑郁预测中的应用仍未得到充分探索。本研究评估了VLM在实验室和自然环境数据集上的表现,重点关注诊断可靠性和人口公平性。实验结果显示,模型在不同环境和架构下表现差异显著,且存在性别和种族偏见。我们的XAI干预框架取得了混合效果,强调未来的公平性干预需同时优化预测准确性和人口平等。

🔬 方法详解

问题定义:本论文旨在解决多模态视觉-语言模型在心理健康评估中的公平性和透明性不足的问题。现有方法在不同人群和环境下表现不均,且存在性别和种族偏见。

核心思路:提出了一种基于可解释人工智能的干预框架,通过分析模型决策过程来提高公平性,同时关注预测准确性和人口平等的平衡。

技术框架:研究分为几个主要阶段,包括数据集选择(AFAR-BSFT和E-DAIC)、模型评估(Phi3.5-Vision和Qwen2-VL)、以及XAI干预实施与效果评估。

关键创新:最重要的创新在于将XAI与公平性干预结合,探索其在VLMs中的应用,尤其是在心理健康领域的实际效果与挑战。

关键设计:在模型训练中,采用了特定的损失函数和公平性约束,针对不同模型架构进行了参数调优,以提高模型在不同数据集上的表现。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,Phi3.5-Vision在E-DAIC数据集上达到了80.4%的准确率,而Qwen2-VL仅为33.9%。尽管XAI干预措施在某些情况下提高了程序一致性,但在AFAR-BSFT上并未保证结果公平性,反而加剧了种族偏见。

🎯 应用场景

该研究的潜在应用领域包括心理健康监测、临床决策支持和社会服务等。通过提高多模态模型的公平性和可解释性,可以更好地服务于不同人群的心理健康需求,促进心理健康技术的普及与应用。

📄 摘要(原文)

In recent years, the integration of multimodal machine learning in wellbeing assessment has offered transformative potential for monitoring mental health. However, with the rapid advancement of Vision-Language Models (VLMs), their deployment in clinical settings has raised concerns due to their lack of transparency and potential for bias. While previous research has explored the intersection of fairness and Explainable AI (XAI), its application to VLMs for wellbeing assessment and depression prediction remains under-explored. This work investigates VLM performance across laboratory (AFAR-BSFT) and naturalistic (E-DAIC) datasets, focusing on diagnostic reliability and demographic fairness. Performance varied substantially across environments and architectures; Phi3.5-Vision achieved 80.4% accuracy on E-DAIC, while Qwen2-VL struggled at 33.9%. Additionally, both models demonstrated a tendency to over-predict depression on AFAR-BSFT. Although bias existed across both architectures, Qwen2-VL showed higher gender disparities, while Phi-3.5-Vision exhibited more racial bias. Our XAI intervention framework yielded mixed results; fairness prompting achieved perfect equal opportunity for Qwen2-VL at a severe accuracy cost on E-DAIC. On AFAR-BSFT, explainability-based interventions improved procedural consistency but did not guarantee outcome fairness, sometimes amplifying racial bias. These results highlight a persistent gap between procedural transparency and equitable outcomes. We analyse these findings and consolidate concrete recommendations for addressing them, emphasising that future fairness interventions must jointly optimise predictive accuracy, demographic parity, and cross-domain generalisation.