cs.CV(2026-07-07)

📊 共 35 篇论文 | 🔗 8 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (11 🔗4) 支柱九:具身大模型 (Embodied Foundation Models) (11 🔗1) 支柱三:空间感知与语义 (Perception & Semantics) (6 🔗1) 支柱四:生成式动作 (Generative Motion) (2 🔗1) 支柱七:动作重定向 (Motion Retargeting) (2) 支柱六:视频提取与匹配 (Video Extraction) (1) 支柱一:机器人控制 (Robot Control) (1) 支柱八:物理动画 (Physics-based Animation) (1 🔗1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (11 篇)

#题目一句话要点标签🔗
1 GaussFusion: Towards Multimodal 3D Gaussian Pretraining 提出GaussFusion以解决3D高斯表示的多模态预训练问题 representation learning MAE 3D gaussian splatting
2 MoWorld: A Flash World Model 提出MoWorld以解决高效实时世界模型构建问题 world model world models distillation
3 What Images Cannot Say: Language-Guided Olfactory Representation Learning 提出SCENT框架以解决视觉与嗅觉信号对齐问题 representation learning multimodal
4 Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning 提出RNSIDNet以解决合成图像检测的跨模型泛化问题 representation learning contrastive learning
5 Few-Medoids: An Embarrassingly Simple Coreset Selection Method for Few-Shot Knowledge Distillation 提出Few-Medoids方法以解决少样本知识蒸馏中的核心集选择问题 teacher-student distillation
6 XRFormer: Multiscale Tokenization for XRF Representation Learning 提出XRFormer以解决XRF光谱自动学习挑战 representation learning
7 Bridging Diffusion Pruning and Step Distillation with Teacher-Aligned Repair 提出教师对齐修复方法以连接扩散剪枝与步骤蒸馏 distillation
8 Straight-Path Flow Matching for Incomplete Multi-View Clustering 提出流匹配框架以解决不完整多视图聚类问题 flow matching
9 Progressive Reasoning with Primitive Correction for Compositional Zero-Shot Learning 提出PRPC框架以解决组合零-shot学习中的错误传播问题 reinforcement learning chain-of-thought
10 AlayaWorld: Long-Horizon and Playable Video World Generation 提出AlayaWorld以解决游戏世界生成的高成本与低灵活性问题 world model world models
11 MobileWan: Closing the Quality Gap for Mobile Video Diffusion 提出MobileWan以解决移动视频生成质量不足问题 linear attention distillation

🔬 支柱九:具身大模型 (Embodied Foundation Models) (11 篇)

#题目一句话要点标签🔗
12 Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models 提出场景图思维以增强多模态大语言模型的结构化视觉推理能力 large language model multimodal
13 Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context 提出Harrison.Rad 1.5以解决放射科报告生成效率问题 large language model foundation model multimodal
14 Structured Data Extraction from Real Estate Documents using Clustering, Classification, and Large Language Models 提出基于聚类、分类和大语言模型的房地产文档结构化数据提取方法 large language model
15 LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding 提出Legato 2以解决乐谱图像识别与理解问题 multimodal
16 Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention 提出多令牌局部注意力以解决多模态大语言模型的幻觉问题 large language model multimodal
17 Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning 提出SegAnswer以解决视觉推理中的像素定位问题 large language model multimodal
18 ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation 提出ELSA3D以解决3D理解与生成中的文本-3D交互问题 foundation model
19 Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders 提出Analysis-by-Proxy以解决VLM作为条件编码器的局限性 multimodal
20 VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery 提出VaseMuseum以解决古希腊陶器数字博物馆中的VLM辅助问题 multimodal
21 AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring 提出AVA-VLM以解决施工现场监控中的效率与可靠性问题 chain-of-thought
22 SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs 提出SAMPLe以解决VLM中提示学习的性能与泛化问题 foundation model

🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)

#题目一句话要点标签🔗
23 CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models 提出CAIRN以解决多房间3D场景理解问题 scene understanding large language model multimodal
24 Vision as Unified Multimodal Generation 提出统一多模态生成方法以解决计算机视觉任务整合问题 depth estimation foundation model multimodal
25 TRIG: Trajectory-Rig Decoupled Metric Geometry Learning 提出TRIG以解决多摄像头自驾车几何感知问题 metric depth 3D reconstruction motion estimation
26 HoloCount: A Holistic Visual Counting Benchmark for MLLMs 提出HoloCount以解决多模态大语言模型的计数精度问题 scene understanding large language model multimodal
27 PhyMRI-SR: Toward Physics-Aware MRI Image Super-Resolution 提出PhyMRI-SR以解决MRI图像超分辨率问题 gaussian splatting splatting physically plausible
28 Why does Deep Learning Improve Visual SLAM? 提出深度学习方法以提升视觉SLAM在复杂环境中的表现 visual SLAM

🔬 支柱四:生成式动作 (Generative Motion) (2 篇)

#题目一句话要点标签🔗
29 DeSeG: Decoupling Semantic Intent and Geometric Constraints for Physically Plausible Human-Scene Interaction 提出DeSeG以解决人类场景交互中的语义与几何约束纠缠问题 motion synthesis motion generation physically plausible
30 SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation 提出SparseCtrl-HOI以解决人机交互视频生成中的稀疏控制问题 motion synthesis physically plausible human-object interaction

🔬 支柱七:动作重定向 (Motion Retargeting) (2 篇)

#题目一句话要点标签🔗
31 Synthetic-to-Real Translation for Class-Agnostic Motion Prediction 提出合成到真实的运动预测方法以解决数据标签获取成本高的问题 motion prediction
32 ARMS: Anchor-Relational Motion Streaming for Seamless Solo-Social Motion Transitions 提出ARMS框架以解决人类动作生成中的社交过渡问题 human motion

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
33 EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage 提出EgoPolice以解决警用身体摄像头视频理解问题 egocentric

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
34 Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator 提出Image2Sim以解决高保真互动环境构建问题 sim-to-real multimodal

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
35 EeveeDark: A Binary Neural Framework for Low-Light Video Enhancement via Event-Guided Sensor-Level Fusion 提出EeveeDark以解决极低光照视频增强问题 spatiotemporal

⬅️ 返回 cs.CV 首页 · 🏠 返回主页