cs.CV(2026-07-16)

📊 共 36 篇论文 | 🔗 11 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (13 🔗5) 支柱二:RL算法与架构 (RL & Architecture) (10 🔗4) 支柱三:空间感知与语义 (Perception & Semantics) (6) 支柱一:机器人控制 (Robot Control) (3 🔗2) 支柱六:视频提取与匹配 (Video Extraction) (1) 支柱四:生成式动作 (Generative Motion) (1) 支柱七:动作重定向 (Motion Retargeting) (1) 支柱八:物理动画 (Physics-based Animation) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (13 篇)

#题目一句话要点标签🔗
1 Symbal: Detecting Systematic Misalignments in Model-Generated Captions 提出Symbal以检测模型生成标题中的系统性错位问题 large language model foundation model multimodal
2 Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI Screening 提出基于自适应焦点损失的高效提示调优方法以解决轻度认知障碍检测问题 foundation model
3 Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality 提出多模态自监督学习以优化特定测试环境中的模型表现 multimodal
4 Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach 提出单一文本中心的多模态方法以解决模棱两可情感识别问题 multimodal
5 Video = World + Event Stream 提出Wan-Streamer v0.3以实现视频世界与事件流的实时交互 vision-language-action multimodal
6 ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships 提出ReBind框架以解决多参考视频编辑中的信息协调问题 large language model multimodal
7 VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance 提出VIABench以解决视觉障碍者辅助问题 large language model multimodal
8 Knowing You at First Glance: Inferring Apparent Personality from Faces 提出GlanceFace以解决面部表情与个性推断问题 large language model multimodal
9 Motion-Conditioned Multi-View Fusion for Myocardial Infarction Localization from Echocardiography 提出MCF-Net以解决心肌梗死定位中的视图依赖性模糊问题 foundation model
10 SceneBind: Binding What and Where Across Vision, Audio and Language 提出SceneBind以解决多模态场景理解中的空间结构缺失问题 zero-shot transfer
11 GeoDetect: Geometric Adversarial Detection for VLPs 提出GeoDetect以解决VLP模型的对抗攻击检测问题 multimodal
12 AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents 提出AdaTurn以解决预算感知的主动视觉感知代理问题 multimodal
13 Uni-AdaVD: Universal Concept Erasure for Visual Generation via Orthogonal Value Decomposition 提出Uni-AdaVD以解决视觉生成中的概念消除问题 multimodal

🔬 支柱二:RL算法与架构 (RL & Architecture) (10 篇)

#题目一句话要点标签🔗
14 AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning 提出AlphaWiSE以解决多模态持续学习中的对齐问题 representation learning multimodal
15 FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models 提出FoMoVLA以解决视觉预测与运动指导的结合问题 policy learning vision-language-action VLA
16 Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models 提出基于蒸馏的MIL预训练框架以解决病理图像分析中的挑战 distillation foundation model
17 3D Geometric Tooth Alignment Planning via Deep Reinforcement Learning 提出深度强化学习框架以自动化3D牙齿对齐规划 reinforcement learning deep reinforcement learning DRL
18 Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding 提出ViPS框架以提升多模态大语言模型的空间理解能力 VIP large language model foundation model
19 HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents 提出HyMobileAgent以解决移动GUI代理的高效交互问题 reinforcement learning reward design foundation model
20 From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting 提出D2DF框架以解决视频对象移除中的草稿依赖问题 distillation optical flow
21 Hierarchical Denoising For Multi-Step Visual Reasoning 提出HDR框架以解决视频多步推理中的逻辑一致性问题 world model world models foundation model
22 WanSong v1.0 Technical Report 提出WanSong以解决长篇音乐生成的效率与可控性问题 distillation foundation model
23 VideoSEMA: a scalable and efficient Mamba-like attention for video understanding 提出VideoSEMA以解决视频理解中的计算效率问题 Mamba

🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)

#题目一句话要点标签🔗
24 Compression of 3D Gaussian Splatting Data Using GPU-friendly Graphics Texture Coding 提出GPU友好的纹理编码方法以压缩3D高斯点数据 3D gaussian splatting 3DGS gaussian splatting
25 JADE-GS: Joint Alternating Deblurring Guided by Events in 3D Gaussian Splatting 提出JADE-GS以解决快速移动相机导致的模糊问题 3D gaussian splatting gaussian splatting splatting
26 G$^2$SR: Geometric Methods for Fast and Memory-Efficient Gaussian-based Surface Reconstruction 提出G$^2$SR以解决快速且内存高效的表面重建问题 3D gaussian splatting 3DGS gaussian splatting
27 Immediate 3D Gaussian Splat Reconstruction of Unordered Input with Global Consistency 提出即时3D高斯点云重建方法以解决无序输入的全局一致性问题 3D gaussian splatting 3DGS gaussian splatting
28 MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos 提出MAGiSt3R框架以实现单目RGB视频的多代理3D重建 3D reconstruction
29 VTM-Nav: Hierarchical Visual-Topological Memory for Cross-Episode Object-Goal Navigation 提出VTM-Nav以解决跨情节目标导航问题 open-vocabulary open vocabulary

🔬 支柱一:机器人控制 (Robot Control) (3 篇)

#题目一句话要点标签🔗
30 SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment 提出SUFLECA以解决CAD与图像对齐问题 sim-to-real foundation model
31 An LLM-Based Automatic Sportscast Solution for Robot Soccer Matches 提出基于LLM的自动体育解说方案以解决机器人足球比赛解说问题 humanoid humanoid robot
32 Skeleton: Visual Authoring of Non-visual Data Experiences 提出Skeleton以解决无视觉数据体验的可访问性问题 manipulation

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
33 Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation 提出Ego Scene Augmentation以解决多模态大语言模型空间感知不足问题 egocentric large language model multimodal

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
34 Physics-Informed Diffusion for Biomechanically Plausible 3D Sign Language Generation 提出物理信息扩散模型以生成生物力学合理的3D手语 classifier-free guidance physics-informed diffusion

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
35 Online Neural Space Time Memory for Dynamic Novel View Synthesis 提出在线神经时空记忆以解决动态新视角合成问题 human motion

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
36 VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding 提出VideoChat3以解决视频理解中的通用性与效率问题 spatiotemporal

⬅️ 返回 cs.CV 首页 · 🏠 返回主页