Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context
作者: Suneeta Mall, Vladimir Nekrasov, Ashnil Kumar, Sajith Karunasena, Aiden Nibali, Alix Bird, Mateo Diaz Shine, Jarrel Seah
分类: cs.CV, cs.AI
发布日期: 2026-07-07
💡 一句话要点
提出Harrison.Rad 1.5以解决放射科报告生成效率问题
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 放射科 多模态模型 自动报告生成 深度学习 临床应用 视觉问答 模型评估
📋 核心要点
- 现有放射科报告生成方法无法有效应对日益增长的影像需求,导致医师工作负担加重。
- Harrison.Rad 1.5通过多模态大语言模型整合影像、临床历史和先前研究,自动生成放射科报告。
- HR1.5在多个评估基准上表现突出,满足FRCR通过标准,并在临床问题准确率上领先于其他系统。
📝 摘要(中文)
随着影像需求的增长,放射科医师的工作负担加重,报告积压问题无法仅通过培训和招聘解决。Harrison.Rad 1.5(HR1.5)是一种放射科特定的多模态大语言模型,能够处理文本和视觉输入,生成结构化和非结构化文本。HR1.5通过三阶段训练流程进行训练,包括对基础语言模型的领域适应、对约600万图像-报告实例的对比视觉编码器训练,以及多轮对话的视觉问答微调。HR1.5在多个评估框架中表现优异,满足模拟FRCR通过标准,并在闭合格式临床问题上取得最高准确率。
🔬 方法详解
问题定义:论文旨在解决放射科报告生成的效率低下问题,现有方法无法满足快速增长的影像需求,导致医师工作压力加大。
核心思路:Harrison.Rad 1.5通过整合文本和视觉输入,利用多模态大语言模型自动生成报告,减少医师的手动工作量。
技术框架:HR1.5的训练流程分为三个阶段:首先对基础语言模型进行领域适应,然后进行对比视觉编码器训练,最后进行视觉问答微调。
关键创新:HR1.5是首个能够处理多种影像类型并生成结构化和非结构化文本的放射科特定模型,显著提高了报告生成的准确性和效率。
关键设计:模型训练使用了约600万图像-报告实例,采用了课程学习策略进行对比训练,并通过多轮对话进行微调,确保模型在临床环境中的有效性。
🖼️ 关键图片
📊 实验亮点
Harrison.Rad 1.5在模拟FRCR考试中表现优异,满足通过标准,并在闭合格式临床问题上实现了最高准确率,显示出其在多模态报告生成中的领先地位。
🎯 应用场景
Harrison.Rad 1.5可广泛应用于放射科领域,帮助医师快速生成影像报告,减轻工作负担,提高工作效率。该模型的成功应用将推动放射科自动化进程,提升医疗服务质量,具有重要的临床价值和社会影响。
📄 摘要(原文)
Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and recruitment alone. The most direct opportunity is reducing the time and effort radiologists spend producing reports, a task that requires interpreting images, integrating clinical history and prior studies, and drafting structured findings. We present Harrison.Rad 1.5 (HR1.5), a radiology-specific multimodal large language model that accepts interleaved text and visual inputs and generates structured and unstructured text across plain-film radiology, spanning computed radiography, chest, musculoskeletal, abdominal, spine, and pelvic x-rays, and mammography. HR1.5 is trained through a three-stage pipeline: domain adaptation of a base language model on radiology reports, contrastive vision-encoder training with curriculum-based hard negatives on ~6 million image-report instances, and visual-question-answering fine-tuning on multi-turn conversations. We evaluate it with a Findings-Diagnosis scoring framework that extends RadGraph-XL entity extraction with ontology-based synonym matching and polarity-contradiction detection, benchmarked on RadBench, a simulated FRCR 2B Short Case examination scored against Angoff-method thresholds, ReXGradient, and internal multi-modality datasets. HR1.5 is the only system evaluated to meet the simulated FRCR passing standard and achieves the highest accuracy on closed-format clinical questions, across anatomical regions, on internal multi-body-part and mammography reporting, and on the primary clinically-aligned score for public chest reporting. We further examine explainability and model behaviour, including question-sensitive Grad-CAM heatmaps, attention analysis, and confidence estimation, to support responsible future evaluation toward clinical use, and a framework for clinically grounded assessment of report quality.