Benchmarking a foundation LLM on its ability to re-label structure names in accordance with the AAPM TG-263 report

📄 arXiv: 2310.03874v1 📥 PDF

作者: Jason Holmes, Lian Zhang, Yuzhen Ding, Hongying Feng, Zhengliang Liu, Tianming Liu, William W. Wong, Sujay A. Vora, Jonathan B. Ashman, Wei Liu

分类: physics.med-ph, cs.CL

发布日期: 2023-10-05

备注: 20 pages, 5 figures, 1 table


💡 一句话要点

利用大型语言模型标准化放射肿瘤学结构名称

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 结构名称标注 放射肿瘤学 医学影像 标准化 DICOM GPT-4

📋 核心要点

  1. 现有方法在结构名称标注上存在不一致性,影响放射治疗的标准化与效果。
  2. 本研究利用GPT-4 API对DICOM文件中的结构名称进行重新标注,以符合AAPM TG-263标准。
  3. 实验结果显示,GPT-4在不同疾病类型的结构名称重新标注准确率高达98.5%,展现出LLMs的潜力。

📝 摘要(中文)

本研究旨在引入大型语言模型(LLMs)用于根据美国医学物理学会(AAPM)TG-263标准重新标注结构名称,并建立基准供未来研究参考。通过实施GPT-4 API作为数字成像与通信医学(DICOM)存储服务器,研究对前列腺、头颈和胸部三种疾病的150名患者进行评估,结果显示结构名称重新标注的准确率分别为96.0%、98.5%和96.9%。研究表明,LLMs在放射肿瘤学中标准化结构名称方面具有广阔前景,尤其是在LLMs能力快速发展的背景下。

🔬 方法详解

问题定义:本研究旨在解决放射肿瘤学中结构名称标注不一致的问题,现有方法缺乏标准化,导致治疗效果不佳。

核心思路:通过利用GPT-4这一大型语言模型,自动化地将结构名称重新标注为符合AAPM TG-263标准,以提高标注的一致性和准确性。

技术框架:整体架构包括DICOM存储服务器,接收结构集DICOM文件后,调用GPT-4进行名称重新标注。研究选择了前列腺、头颈和胸部三种疾病进行评估。

关键创新:本研究的创新在于将大型语言模型应用于医学影像领域,尤其是结构名称的标准化,突破了传统手动标注的局限性。

关键设计:在实验中,针对每种疾病类型随机选择150名患者进行指令提示的手动调优,并从中选取50名患者进行评估,确保了实验的严谨性和结果的可靠性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,GPT-4在前列腺、头颈和胸部疾病的结构名称重新标注准确率分别为96.0%、98.5%和96.9%。其中,前列腺的目标体积标注准确率达到100%,显示出LLMs在医学标注中的卓越性能。

🎯 应用场景

该研究的潜在应用领域包括放射肿瘤学、医学影像处理和临床决策支持系统。通过标准化结构名称,可以提高放射治疗的准确性和一致性,进而提升患者的治疗效果和安全性。未来,随着LLMs技术的进一步发展,其在医学领域的应用将更加广泛。

📄 摘要(原文)

Purpose: To introduce the concept of using large language models (LLMs) to re-label structure names in accordance with the American Association of Physicists in Medicine (AAPM) Task Group (TG)-263 standard, and to establish a benchmark for future studies to reference. Methods and Materials: The Generative Pre-trained Transformer (GPT)-4 application programming interface (API) was implemented as a Digital Imaging and Communications in Medicine (DICOM) storage server, which upon receiving a structure set DICOM file, prompts GPT-4 to re-label the structure names of both target volumes and normal tissues according to the AAPM TG-263. Three disease sites, prostate, head and neck, and thorax were selected for evaluation. For each disease site category, 150 patients were randomly selected for manually tuning the instructions prompt (in batches of 50) and 50 patients were randomly selected for evaluation. Structure names that were considered were those that were most likely to be relevant for studies utilizing structure contours for many patients. Results: The overall re-labeling accuracy of both target volumes and normal tissues for prostate, head and neck, and thorax cases was 96.0%, 98.5%, and 96.9% respectively. Re-labeling of target volumes was less accurate on average except for prostate - 100%, 93.1%, and 91.1% respectively. Conclusions: Given the accuracy of GPT-4 in re-labeling structure names of both target volumes and normal tissues as presented in this work, LLMs are poised to be the preferred method for standardizing structure names in radiation oncology, especially considering the rapid advancements in LLM capabilities that are likely to continue.