Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

📄 arXiv: 2607.21694v1 📥 PDF

作者: Yong Liu, Xiaolong Fu, Zihang Xu, Wen Xue, Xueheng Li, Lin Song, Yuan Zhang, Chuyang Zhao, Haoyang Huang, Nan Duan, Yipeng Sun, Yan Li, Simiu Gu

分类: cs.CV

发布日期: 2026-07-23


💡 一句话要点

提出Oxygen-TryOn以解决多样化虚拟试穿问题

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 虚拟试穿 时尚原生模型 多物品合成 数据引擎 强化学习 图像生成 计算机视觉

📋 核心要点

  1. 现有方法通常仅支持单一服装类别的试穿,且多参考方法仍然以服装为中心,缺乏对多样化物品的支持。
  2. Oxygen-TryOn通过专用的数据引擎和多参考生成任务的重新定义,能够处理多种物品和场景,合成高质量的试穿图像。
  3. 在公开基准和内部测试中,Oxygen-TryOn在单物品和多物品试穿上均实现了领先的连贯性和真实感,超越了多种现有系统。

📝 摘要(中文)

我们提出了Oxygen-TryOn,一个统一的基础模型,专为任何物品的虚拟试穿而设计。与传统的通用图像编辑器不同,Oxygen-TryOn是时尚原生的,专注于通过专用的数据引擎和特定的训练方法实现试穿。该模型能够处理多种参考物品和单一目标主体图像,合成出主体穿着各种时尚类别物品的逼真图像。与以往系统仅支持单一服装类别不同,Oxygen-TryOn支持多样化的物品和场景,能够在保持主体身份和物品外观的同时,实现自由的多物品组合。通过构建高质量的试穿数据引擎和设计三阶段的训练流程,Oxygen-TryOn在多个基准测试中展现出卓越的连贯性和真实感。

🔬 方法详解

问题定义:本论文旨在解决现有虚拟试穿系统仅支持单一服装类别的问题,导致在多样化物品和场景下的应用受限。现有方法在处理多参考物品时仍然以服装为中心,缺乏灵活性和多样性。

核心思路:Oxygen-TryOn的核心思路是将试穿任务重新定义为一个多参考、理解驱动的生成任务,利用专用的数据引擎和训练流程来提升生成图像的质量和多样性。

技术框架:整体架构包括三个主要阶段:持续预训练(CPT)、监督微调(SFT)和强化学习(RL)。数据引擎负责收集、制造、注释和过滤高质量的试穿数据。

关键创新:Oxygen-TryOn的最大创新在于其时尚原生的设计和多样化的支持能力,能够在不同场景下合成多物品试穿图像,且保持主体身份和物品外观的一致性。

关键设计:在训练过程中,采用混合奖励机制,结合内部试穿奖励模型与通用模型,确保生成图像的细粒度一致性和指令级质量。此外,模型还能够在同一过程中遵循一般编辑指令(如姿势变化)。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

在多个公开基准和内部Oxygen-TryOn基准测试中,Oxygen-TryOn在单物品试穿上实现了最先进的连贯性和真实感,并在多物品试穿上超越了领先的专有系统(如Nano Banana Pro、GPT-Image-2、Seedream5 Lite)和开源模型(如FLUX.2),展现出显著的性能提升。

🎯 应用场景

Oxygen-TryOn的潜在应用领域包括在线时尚零售、虚拟试衣间和社交媒体平台,能够为用户提供个性化的购物体验和增强的互动性。未来,该技术可能会推动虚拟时尚行业的发展,提升消费者的购买决策效率。

📄 摘要(原文)

We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).