ChatRadio-Valuer: A Chat Large Language Model for Generalizable Radiology Report Generation Based on Multi-institution and Multi-system Data

📄 arXiv: 2310.05242v2 📥 PDF

作者: Tianyang Zhong, Wei Zhao, Yutong Zhang, Yi Pan, Peixin Dong, Zuowei Jiang, Xiaoyan Kui, Youlan Shang, Li Yang, Yaonai Wei, Longtao Yang, Hao Chen, Huan Zhao, Yuxiao Liu, Ning Zhu, Yiwei Li, Yisong Wang, Jiaqi Yao, Jiaqi Wang, Ying Zeng, Lei He, Chao Zheng, Zhixue Zhang, Ming Li, Zhengliang Liu, Haixing Dai, Zihao Wu, Lu Zhang, Shu Zhang, Xiaoyan Cai, Xintao Hu, Shijie Zhao, Xi Jiang, Xin Zhang, Xiang Li, Dajiang Zhu, Lei Guo, Dinggang Shen, Junwei Han, Tianming Liu, Jun Liu, Tuo Zhang

分类: cs.CL, cs.AI

发布日期: 2023-10-08 (更新: 2023-10-10)


💡 一句话要点

提出ChatRadio-Valuer以解决放射科报告生成的通用性挑战

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 放射科报告生成 大型语言模型 模型微调 多机构适应性 临床AI应用 疾病诊断 通用性挑战

📋 核心要点

  1. 现有放射科报告生成方法面临跨机构和跨系统的风格差异,导致通用性不足。
  2. ChatRadio-Valuer基于大型语言模型,通过单一机构的报告进行微调,并适应多机构的疾病诊断任务。
  3. 实验结果显示,ChatRadio-Valuer在疾病诊断准确性和临床应用效率上优于现有模型,尤其是ChatGPT系列。

📝 摘要(中文)

放射科报告生成是医学图像分析中的关键步骤,对临床决策具有重要意义。然而,来自不同机构和体检区域的报告风格差异使得现有方法面临通用性挑战。为此,本文提出了基于大型语言模型的ChatRadio-Valuer,旨在通过学习通用表示来提升模型的适应性。该模型在单一机构的放射科报告上进行监督微调,并适应于来自六个不同机构的多系统疾病诊断任务。实验结果表明,ChatRadio-Valuer在疾病诊断方面显著优于现有的最先进模型,尤其在临床应用中展现出更高的有效性和更低的部署成本。

🔬 方法详解

问题定义:本文旨在解决放射科报告生成中的通用性挑战,现有方法在处理来自不同机构和体检区域的报告时,面临风格和规范性差异的问题,导致模型的适应性不足。

核心思路:提出ChatRadio-Valuer模型,利用大型语言模型的强大能力,通过在单一机构的放射科报告上进行监督微调,学习通用的表示,从而提高模型在多机构、多系统的适应能力。

技术框架:ChatRadio-Valuer的整体架构包括两个主要阶段:首先在单一机构的放射科报告上进行训练,然后将模型适应于来自六个不同机构的多系统疾病诊断任务。该框架确保了模型能够在不同的临床环境中有效应用。

关键创新:ChatRadio-Valuer的主要创新在于其基于大型语言模型的设计,能够有效学习不同机构间的报告风格差异,从而实现更高的通用性和适应性。这一设计与传统方法相比,显著提升了模型的灵活性。

关键设计:在模型训练过程中,采用了特定的损失函数以优化报告生成的准确性,并设计了适应性强的网络结构,以便在不同的临床任务中保持高效性能。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,ChatRadio-Valuer在疾病诊断任务中表现优异,尤其在与ChatGPT(GPT-3.5-Turbo和GPT-4等)对比时,准确性提升显著,具体性能指标显示其在332,673个观察样本中均优于现有最先进模型。

🎯 应用场景

ChatRadio-Valuer的研究成果在放射科领域具有广泛的应用潜力,能够有效提升放射科报告的生成效率和准确性,减轻专家的标注负担。这将促进临床AI应用的发展,推动智能医疗的进步,尤其是在多机构合作的环境中。

📄 摘要(原文)

Radiology report generation, as a key step in medical image analysis, is critical to the quantitative analysis of clinically informed decision-making levels. However, complex and diverse radiology reports with cross-source heterogeneity pose a huge generalizability challenge to the current methods under massive data volume, mainly because the style and normativity of radiology reports are obviously distinctive among institutions, body regions inspected and radiologists. Recently, the advent of large language models (LLM) offers great potential for recognizing signs of health conditions. To resolve the above problem, we collaborate with the Second Xiangya Hospital in China and propose ChatRadio-Valuer based on the LLM, a tailored model for automatic radiology report generation that learns generalizable representations and provides a basis pattern for model adaptation in sophisticated analysts' cases. Specifically, ChatRadio-Valuer is trained based on the radiology reports from a single institution by means of supervised fine-tuning, and then adapted to disease diagnosis tasks for human multi-system evaluation (i.e., chest, abdomen, muscle-skeleton, head, and maxillofacial $\&$ neck) from six different institutions in clinical-level events. The clinical dataset utilized in this study encompasses a remarkable total of \textbf{332,673} observations. From the comprehensive results on engineering indicators, clinical efficacy and deployment cost metrics, it can be shown that ChatRadio-Valuer consistently outperforms state-of-the-art models, especially ChatGPT (GPT-3.5-Turbo) and GPT-4 et al., in terms of the diseases diagnosis from radiology reports. ChatRadio-Valuer provides an effective avenue to boost model generalization performance and alleviate the annotation workload of experts to enable the promotion of clinical AI applications in radiology reports.