Explainable Multimodal Deep Learning Integrating Imaging and Clinical Data for Oral Potentially Malignant Disorder Detection

📄 arXiv: 2609.04512v1 📥 PDF

作者: Ruilin You, Yihan Wang, Jiabin Chen, Cherie Wink, Petra Wilder-Smith, Rongguang Liang, Bofan Song

分类: eess.IV, cs.CV

发布日期: 2026-09-03


💡 一句话要点

提出M2-OPMDNet以解决口腔潜在恶性疾病检测难题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 多模态学习 深度学习 口腔癌检测 临床数据 可解释性 图像处理 风险评估

📋 核心要点

  1. 现有的口腔潜在恶性疾病检测方法面临表型异质性和良性病变重叠的挑战,导致临床检测困难。
  2. 本文提出的M2-OPMDNet框架通过整合图像和结构化临床数据,提升了OPMD的检测准确性和可解释性。
  3. 实验结果显示,M2-OPMDNet在AUC上达到0.952,显著优于单模态方法,尤其在检测微妙病变时表现突出。

📝 摘要(中文)

口腔潜在恶性疾病(OPMDs)是口腔癌的重要前驱,但由于表型异质性和与良性病变的重叠,临床检测仍然具有挑战性。尽管基于图像的深度学习在自动筛查中显示出前景,但仅依赖视觉信息在实际应用中可能不足。本文开发了M2-OPMDNet,一个多模态深度学习框架,整合了共注册的白光和自发荧光口腔图像与结构化临床信息,以实现OPMD的检测。通过SHAP分析,结果表明结构化临床变量对风险评估有显著贡献,M2-OPMDNet在AUC上达到0.952,超越了单模态方法,尤其在视觉上微妙的病变检测中表现更佳。

🔬 方法详解

问题定义:本文旨在解决口腔潜在恶性疾病(OPMDs)检测中的准确性和可解释性问题。现有方法往往仅依赖图像信息,无法充分考虑患者特定的风险因素,导致检测效果不佳。

核心思路:M2-OPMDNet通过整合白光和自发荧光图像与结构化临床信息,提供了一种多模态的深度学习解决方案。这种设计旨在利用不同数据源的互补性,提高检测的准确性和可靠性。

技术框架:该框架包括多个模块:首先,通过定制问卷收集结构化临床信息;其次,使用多种图像编码器(如卷积神经网络和基础模型架构)处理图像数据;最后,利用SHAP进行模型可解释性分析,量化特征和模态的贡献。

关键创新:M2-OPMDNet的主要创新在于其多模态学习能力,能够有效结合图像和临床数据,显著提升了对OPMD的检测性能。这与传统单模态方法形成鲜明对比,后者往往忽视了临床信息的作用。

关键设计:在模型设计中,采用了多种图像编码器以适应不同类型的输入数据,并通过SHAP分析评估各特征的贡献。此外,定制的问卷设计确保了临床信息的标准化和可重复性,为模型提供了丰富的上下文信息。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

M2-OPMDNet在AUC上达到0.952,显著优于传统单模态方法,尤其在检测视觉上微妙的病变时表现出色。SHAP分析表明,结构化临床变量对风险评估的贡献显著,进一步验证了多模态学习的有效性。

🎯 应用场景

M2-OPMDNet可广泛应用于口腔癌的早期筛查和临床决策支持,尤其适合需要结合影像和临床信息的复杂病例。该框架的可扩展性使其能够适应不同的临床环境,提升口腔健康管理的整体效率和准确性。

📄 摘要(原文)

Oral potentially malignant disorders (OPMDs) are critical precursors to oral cancer, yet clinical detection remains challenging because of substantial phenotypic heterogeneity and overlap with benign conditions. Although image-based deep learning shows promise for automated screening, visual information alone may be insufficient in real-world settings, where diagnostic decisions also rely on patient-specific risk factors. We developed M2-OPMDNet, a multimodal deep learning framework that integrates co-registered white-light and autofluorescence intraoral images with structured clinical information for OPMD detection. A customized questionnaire was designed to capture clinically relevant risk factors and symptoms in a standardized, reproducible format for integration with image-derived features. Multiple image encoders, including conventional convolutional neural networks and foundation model-based architectures, were evaluated using a prospectively collected dataset reflecting real-world screening conditions. Model interpretability was assessed using SHapley Additive exPlanations (SHAP) to quantify feature- and modality-level contributions. M2-OPMDNet achieved an AUC of 0.952, outperforming unimodal approaches and showing improved performance for visually subtle lesions. SHAP analysis demonstrated that structured clinical variables contributed substantially to risk estimation and complemented imaging features. These results demonstrate that explainable multimodal learning combining white-light and autofluorescence imaging with structured clinical data can provide accurate, transparent, and clinically grounded OPMD detection. M2-OPMDNet offers a scalable framework for real-world oral cancer screening and decision support.