cs.CV(2026-07-23)

📊 共 36 篇论文 | 🔗 10 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (12 🔗3) 支柱二:RL算法与架构 (RL & Architecture) (10 🔗2) 支柱三:空间感知与语义 (Perception & Semantics) (9 🔗3) 支柱七:动作重定向 (Motion Retargeting) (2 🔗1) 支柱一:机器人控制 (Robot Control) (1 🔗1) 支柱五:交互与反应 (Interaction & Reaction) (1) 支柱八:物理动画 (Physics-based Animation) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (12 篇)

#题目一句话要点标签🔗
1 HalluScope: Fine-grained Hallucination Diagnosis for Multimodal Large Language Models 提出HalluScope以解决多模态大语言模型的幻觉诊断问题 large language model multimodal
2 Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning 提出FAIRGET和FAUN以解决多模态LLM的不平衡遗忘问题 large language model multimodal
3 Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features 提出质量感知多模态融合以解决面部识别挑战 multimodal
4 When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation 提出ResponseGuard以解决实时内容审查中的效率问题 multimodal chain-of-thought
5 CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA 提出CRAG-MM-Diagnostics以解决KI-VQA分析不足问题 multimodal visual grounding
6 C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs 提出C-PTQ以解决多模态大语言模型量化性能下降问题 large language model multimodal
7 MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement 提出情感声明判断以解决多模态大语言模型的视觉情感智能问题 large language model multimodal
8 ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? 提出ViSTR-Bench以评估MLLMs在动态场景中的推理能力 large language model multimodal
9 3D-Aware VLMs with Implicit and Explicit Geometries 提出VLM-IE3D以解决3D视觉语言模型的空间理解问题 visual grounding
10 GroupVideo: Multi-Identity Customized Text-to-Video Generation 提出GroupVideo以解决多身份视频生成中的身份混淆问题 multimodal
11 DINO-VPT: Hierarchical Visual Prompt Tuning for Joint Physical-Digital Face Anti-Spoofing 提出DINO-VPT以解决统一人脸反欺诈问题 multimodal
12 Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation 提出Agentic Designer以解决室内布局生成中的结构约束问题 large language model

🔬 支柱二:RL算法与架构 (RL & Architecture) (10 篇)

#题目一句话要点标签🔗
13 HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving 提出HyWorldVLA以解决自主驾驶中的噪声鲁棒性问题 world model world models representation learning
14 Self-Supervised Learning of Structured Dynamics from Videos 提出结构化动态模型以解决视频运动理解问题 representation learning VGGT motion representation
15 Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving 提出自监督深度估计方法以应对恶劣天气挑战 distillation depth estimation
16 Visual Contrastive Self-Distillation 提出视觉对比自蒸馏方法以简化自蒸馏过程 distillation
17 SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation 提出SANA-Video 2.0以高效生成720p视频 linear attention
18 Webly Supervised Multi-Label Recognition: Evaluation Benchmark and Dual-Branch Multi-Label Contrastive Learning 提出双分支多标签对比学习框架以解决网络监督多标签识别问题 contrastive learning
19 Unified Video Dense Prediction from Disjoint Data 提出UniD模型以解决视频场景理解中的数据碎片化问题 distillation scene understanding
20 Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers 提出W^2模型以解决多代理视频生成中的状态一致性问题 world model world models
21 Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On 提出Oxygen-TryOn以解决多样化虚拟试穿问题 reinforcement learning foundation model
22 Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints 提出ConVBench和ConVLM以增强LVLM的视觉推理能力 reinforcement learning reward design

🔬 支柱三:空间感知与语义 (Perception & Semantics) (9 篇)

#题目一句话要点标签🔗
23 GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis 提出GrainGS以解决动态场景重建中的运动建模与结构稳定性问题 3D gaussian splatting gaussian splatting splatting
24 DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV 提出DAPM以解决无人机单目深度估计在多变视角下的问题 depth estimation monocular depth 3D reconstruction
25 DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation 提出DINOde以解决开放词汇语义分割中的文本对齐问题 open-vocabulary open vocabulary
26 WAT3R: Feedforward Underwater 3D Reconstruction 提出WAT3R以解决水下3D重建中的光衰减与反向散射问题 depth estimation monocular depth 3D reconstruction
27 SubSplat: High-Resolution Pixel-aligned 3DGS via Sub-pixel Gaussian Reparameterization 提出SubSplat以解决高分辨率像素对齐3D图像合成问题 3DGS gaussian splatting splatting
28 Future Rendering $\neq$ Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed Window 提出FutureSurf以解决动态表面重建的未来几何问题 3DGS scene reconstruction
29 Scene Parameter Saliency via Differentiable Light Transport 提出可微光传输的场景参数显著性方法以提升模型可解释性 scene understanding
30 GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis 提出GrainGS以解决动态场景重建中的运动建模与结构稳定性问题 3D gaussian splatting gaussian splatting splatting
31 Hash-QNeRF: Multiresolution Hash Encoding for Quantum Neural Radiance Fields 提出Hash-QNeRF以解决NeRF在高保真渲染中的计算效率问题 NeRF neural radiance field

🔬 支柱七:动作重定向 (Motion Retargeting) (2 篇)

#题目一句话要点标签🔗
32 Geo3R: Mitigating Spatial Reasoning Hallucination in Multimodal Large Language Models 提出Geo3R以解决多模态大语言模型中的空间推理幻觉问题 spatial relationship large language model multimodal
33 HGeo-TopoMap: Boosting Topological Mapping with Hierarchical Geometric Priors 提出HGeo-TopoMap以解决中心线检测难题 spatial relationship geometric consistency

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
34 TransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects 提出TransBiolab以解决透明生物医学物体识别与操作问题 manipulation depth estimation 6D pose estimation

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
35 Out of Sight, Still in Mind: Token Compression for Omni-LLMs 提出ReMo框架以降低Omni-LLMs的输入token成本 ReMoS large language model

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
36 Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos 提出多模态表示方法以优化城市步行视频的参与度评估 spatiotemporal

⬅️ 返回 cs.CV 首页 · 🏠 返回主页