cs.CV(2026-07-21)

📊 共 29 篇论文 | 🔗 6 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (11 🔗1) 支柱九:具身大模型 (Embodied Foundation Models) (9 🔗1) 支柱三:空间感知与语义 (Perception & Semantics) (6 🔗4) 支柱一:机器人控制 (Robot Control) (1) 支柱五:交互与反应 (Interaction & Reaction) (1) 支柱四:生成式动作 (Generative Motion) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (11 篇)

#题目一句话要点标签🔗
1 Latent Riemannian Flow Matching for Geometry-Grounded 3D Foundation Models 提出流匹配方法以解决3D生成模型的几何限制问题 flow matching VGGT foundation model
2 Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing 提出Mage-Flow以实现高效的图像生成与编辑 flow matching distillation foundation model
3 Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention 提出CURE框架以解决异构医疗数据的多模态融合问题 representation learning multimodal
4 FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling 提出FilmWorld以解决小说转电影生成的复杂问题 world model world models
5 Privileged Lesion-Context Relational Distillation for Mask-Free Skin Lesion Classification 提出特权病变上下文关系蒸馏以解决无掩膜皮肤病变分类问题 teacher-student distillation feature matching
6 Contrastive On-Policy Distillation 提出对比性在线蒸馏框架以提升推理效率 distillation multimodal
7 Norm or Direction? Decoding Vision Mambas for High-Resolution Vision 提出Vision Mamba以提高高分辨率视觉任务的性能 Mamba SSM state space model
8 Weakly Supervised Pathology-Informed Representation Learning for PET-Based Content Retrieval of Intra-Tumour Heterogeneity 提出弱监督病理信息引导的PET图像检索方法以解决肿瘤异质性问题 representation learning teacher-student
9 ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU 提出ABot-World-0以实现实时长时间闭环交互 world model world models distillation
10 OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation 提出OPD-IAD以解决工业异常检测中的像素级定位问题 distillation
11 Generative World Renderer at the Speed of Play 提出AlayaRenderer-Flash以解决实时生成世界渲染问题 world model world models

🔬 支柱九:具身大模型 (Embodied Foundation Models) (9 篇)

#题目一句话要点标签🔗
12 Appearance Pointers -- Multimodal Region Control of Diffusion Transformers 提出外观指针以解决多模态区域控制问题 multimodal
13 Delineate Anything v2: A Global Foundation Model for Field Delineation 提出Delineate Anything v2以解决农业领域边界划分问题 foundation model
14 TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering 提出TAP-RAG以解决长文档多模态问答中的证据使用问题 multimodal
15 In-Context Learning for Wound Classification with Small Multimodal Language Models 提出小型多模态语言模型以解决伤口分类问题 multimodal
16 Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval 提出自相似性基础的时刻提议与评分方法以解决零-shot视频时刻检索中的模态与语言风格差距问题 large language model multimodal
17 Continual Video-MLLM Adaptation over Evolving Domains 提出DAER框架以解决视频多模态大语言模型的持续适应问题 large language model multimodal
18 Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions 提出自适应视图检索以解决隐藏仇恨幻觉检测问题 multimodal
19 GLID: Gated Local Intrinsic Dimension Repairs the Blind Spots of Face-Forgery Detectors 提出GLID以解决人脸伪造检测盲点问题 foundation model
20 Bounding Boxes to Improve Small Language Model Performance on Vision-Based Grading Tasks 通过边界框提升小型语言模型在视觉评分任务中的表现 chain-of-thought

🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)

#题目一句话要点标签🔗
21 ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting 提出ZeroSplat以解决多目标动态分割问题 3D gaussian splatting 3DGS gaussian splatting
22 Open-Vocabulary Gaze Object Prediction: Benchmark and Method 提出开放词汇的注视目标预测方法以解决现有方法的局限性 open-vocabulary open vocabulary
23 IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer 提出IGGT4D以解决动态场景下的实例基础几何理解问题 3D reconstruction scene understanding open-vocabulary
24 FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility 提出FlexiAvatar以解决单目视频中3D人类头像重建问题 3D gaussian splatting gaussian splatting splatting
25 Seeing Before Generating: Object Perception Enhances Single-View 3D Reconstruction 提出基于感知信号的单视图3D重建方法以提升重建精度 3D reconstruction
26 UVFaceFusion: Fast Multi-view Topologically Consistent Face Reconstruction in the Wild via UV-space Neural Fusion 提出UVFaceFusion以解决高保真面部重建问题 VGGT

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
27 Masked Visual Actions for Unified World Modeling 提出Masked Visual Actions以解决机器人世界建模问题 manipulation world model world models

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
28 Sarus: Privacy-Preserving Multi-Vendor Perception Fusion via Homomorphic Encryption 提出Sarus框架以解决多供应商感知融合中的隐私问题 OMOMO

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
29 Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation 提出Moving Alphabet以研究文本到视频生成中的数据影响 classifier-free guidance

⬅️ 返回 cs.CV 首页 · 🏠 返回主页