A Unified Framework for Comprehensive Cardiac CT Segmentation and Phenotyping: Human-in-the-Loop Data Annotation, Vision Foundation Model Development, Multicenter Evaluation and Clinical Validation

📄 arXiv: 2607.11287v1 📥 PDF

作者: Pooya Mohammadi Kazaj, Leo Fridolin Weber, Wen Xie, Seyed Amir Ahmad Safavi-Naini, Anselm Stark, Giovanni Baj, Ali Mokhtari, Toshiya Yoshida, Christoph Ryffel, Taishi Okuno, Yoshihiro Akashi, Ronny R. Buechel, Thomas Pilgrim, Waldo Valenzuela, George C. M. Siontis, Xiaowei Xu, Moritz Hundertmark, Stephan Windecker, Christoph Grani, Isaac Shiri

分类: cs.CV, cs.AI

发布日期: 2026-07-13


💡 一句话要点

提出统一框架以解决心脏CT分割与表型分析问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 心脏CT分割 自监督学习 数据增强 人机协作 临床应用 表型分析 深度学习

📋 核心要点

  1. 现有心脏CT分割方法在可扩展性和常规使用上存在显著不足,限制了其临床应用。
  2. 论文提出的框架结合了人机协作标注、CT增强技术和自监督预训练,旨在提高分割精度和效率。
  3. 实验结果表明,该框架在多个外部数据集上超越了现有工具,尤其在低数据环境下表现出色。

📝 摘要(中文)

全面量化心脏结构的计算机断层扫描(CT)仍然受到测量可扩展性的限制,使得常规使用变得不切实际。本文提出了一种统一的心脏CT分割与表型分析框架,结合了人机协作的标注流程、心脏CT增强技术以及在60,000个未标记心脏CT扫描上进行自监督预训练的基础模型。通过这种方法,组建了迄今为止最大的专家标注心脏CT分割数据集,包含1598个病例和14种不同的心脏结构。该框架在五个外部数据集上比现有开源工具更准确、全面地分割所有结构。自监督预训练提高了标注效率,尤其在低数据环境下的外部评估中表现出显著提升。该框架还扩展到人群级别的表型分析,提供了与心室功能和疾病严重性相关的功能信息。

🔬 方法详解

问题定义:本文旨在解决心脏CT分割的可扩展性问题,现有方法在数据标注和处理效率上存在显著挑战,限制了其在临床中的应用。

核心思路:提出的框架通过人机协作的标注流程与自监督学习相结合,旨在提高数据标注的效率和分割的准确性,从而实现全面的心脏CT表型分析。

技术框架:整体架构包括三个主要模块:人机协作的标注流程、心脏CT数据增强技术和自监督基础模型的预训练。首先,通过人机协作提高标注质量;其次,利用数据增强技术扩展训练数据集;最后,使用自监督学习提升模型的泛化能力。

关键创新:最重要的技术创新在于结合了人机协作与自监督学习,形成了一个高效的标注和分割流程。这一方法与传统的完全依赖人工标注或监督学习的方法有本质区别。

关键设计:在模型设计上,采用了多种网络架构(卷积、变换器和状态空间架构),并通过自监督预训练显著提高了标注效率。损失函数的选择和参数设置经过精心调整,以确保模型在不同数据集上的鲁棒性和准确性。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果显示,该框架在五个外部数据集上实现了比现有开源工具更高的分割准确性,尤其在低数据环境下,自监督预训练带来了显著的效率提升。具体而言,框架在1598个病例中成功分割14种心脏结构,标志着心脏CT分析的重大进展。

🎯 应用场景

该研究的潜在应用领域包括心脏病学、影像学和临床诊断,能够为医生提供更准确的心脏结构分析和疾病评估。通过开放发布数据集和模型,促进了相关领域的研究和应用,未来可能推动个性化医疗的发展。

📄 摘要(原文)

Comprehensive quantification of cardiac structures from computed tomography (CT) remains limited not by data availability but by the scalability of measurements, which makes routine use impractical. Here we present a unified framework for comprehensive cardiac CT segmentation and phenotyping that combines a human-in-the-loop annotation pipeline, a cardiac CT augmentation technique, and a self-supervised foundation model pre-trained on 60,000 unlabeled cardiac CT scans. Using this approach, we assembled the largest and most comprehensive expert-annotated cardiac CT segmentation dataset to date, comprising 1598 cases and 14 distinct cardiac structures (1000 for training, 598 for the external test set). Across five external datasets, the framework segmented all structures more accurately and comprehensively than existing open-source tools. Self-supervised pre-training improved labeling efficiency, with the most significant gains observed during external evaluation in the low-data regime. Benchmarking across convolutional, transformer, and state-space architectures showed comparable performance, indicating that data quality and pre-training, rather than architecture, drove accuracy. The framework was scaled to population-level phenotyping, with segmented anatomy that carries functionally relevant information about ventricular function and disease severity beyond demographic variables. By openly releasing the largest dataset with human labels, code, model weights, a CT augmentation library, and software, this work provides a reproducible foundation for opportunistic cardiac phenotyping from routinely acquired CT scans.