LPM: Industrial-Scale Generative Video Restoration
作者: Bichuan Zhu, Fulin Li, Jiachao Gong, Jinhua Hao, Kai Zhao, Kun Yuan, Pengcheng Xu, Qiang Wang, Qiao Mo, Yanlong Yuan, Yizhen Shao, Yuxiao Hu, Zixi Tuo, Ming Sun, Chao Zhou, Bin Chen, Bin Yu
分类: cs.CV
发布日期: 2026-07-15
备注: 21 pages, 7 figures
💡 一句话要点
提出LPM以解决工业规模视频恢复问题
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 视频恢复 生成模型 扩散模型 用户生成内容 工业应用 带宽优化 时间一致性 高保真
📋 核心要点
- 现有的视频恢复方法在处理复杂退化时面临挑战,尤其是在用户生成内容中多样化的质量问题。
- LPM通过扩散生成框架,结合大规模数据工程和高效推理,提供了一种统一的解决方案以恢复视频质量。
- 在快手的实际应用中,LPM处理的视频占总观看时间的45%,并在比特率和用户体验上显著提升。
📝 摘要(中文)
本文提出了大型处理模型(LPM),这是一个基于扩散的生成框架,旨在应对复杂的自然环境下的视频恢复问题。LPM是首个在工业规模上部署的生成视频恢复模型,能够通过统一系统处理用户生成内容中的多样化退化。其增强的架构、渐进式训练策略和时间金字塔推理机制共同实现了对任意长度视频的高保真、时间一致性恢复。LPM已在快手投入生产,处理的视频占总观看时间的约45%,在关键的用户体验质量指标上持续改善。此外,LPM在保持感知质量的同时,减少了20%的比特率,带来了数亿的年度带宽成本节省,显示出其在大规模视频处理中的实用性和经济性。
🔬 方法详解
问题定义:本文旨在解决在复杂环境下的视频恢复问题,现有方法在处理用户生成内容时常常无法有效应对多样化的退化情况。
核心思路:LPM采用基于扩散的生成模型,通过大规模数据工程和高效推理机制,提供了一种统一的框架来恢复视频质量,确保高保真和时间一致性。
技术框架:LPM的整体架构包括数据预处理、基础模型训练和推理阶段,采用渐进式训练策略和时间金字塔推理机制,以适应任意长度的视频恢复需求。
关键创新:LPM的主要创新在于其工业规模的部署能力和高效的比特率优化,与现有方法相比,能够在保持感知质量的同时显著降低带宽需求。
关键设计:LPM设计了特定的损失函数以优化视频质量,并在网络结构中引入了时间金字塔模块,以增强时间一致性和细节恢复能力。
🖼️ 关键图片
📊 实验亮点
LPM在快手的应用中,处理的视频占总观看时间的45%,在保持感知质量的同时,成功将比特率降低了20%。这些改进不仅提升了用户体验,还为快手带来了数亿的年度带宽成本节省,展示了其在实际应用中的显著效果。
🎯 应用场景
LPM在视频处理领域具有广泛的应用潜力,尤其适用于社交媒体平台和视频分享网站。其高效的恢复能力和显著的成本节省使其在大规模视频处理上具备实际价值,未来可能推动更多生成模型在工业应用中的落地。
📄 摘要(原文)
We present the Large Processing Model (LPM), a diffusion-based generative framework for photorealistic video restoration under complex, in-the-wild degradations. To our knowledge, LPM is the first generative video restoration model deployed at industrial scale. LPM addresses the diverse degradations in user-generated content (UGC) through a unified system encompassing large-scale data engineering, foundation-model training, and efficient inference. Its enhanced architecture, progressive training strategy, and temporal-pyramid inference mechanism jointly enable high-fidelity, temporally consistent restoration of arbitrarily long videos across the broad content distribution encountered on UGC platforms. LPM has been deployed in production at Kuaishou, where videos processed by the model account for approximately 45% of total viewing time, delivering consistent improvements across key quality-of-experience metrics. Beyond perceptual enhancement, LPM delivers substantial system-level benefits: at comparable perceptual quality, it reduces bitrate by 20% relative to Kuaishou's in-house codec, yielding annual bandwidth cost savings on the order of hundreds of millions. Its low serving cost also enables integration into products such as Kling, demonstrating that generative restoration can be practical, scalable, and cost-effective for large-scale video processing.