cs.CV(2026-07-10)

📊 共 33 篇论文 | 🔗 11 篇有代码

🎯 兴趣领域导航

支柱三:空间感知与语义 (Perception & Semantics) (11 🔗4) 支柱二:RL算法与架构 (RL & Architecture) (11 🔗4) 支柱九:具身大模型 (Embodied Foundation Models) (7 🔗2) 支柱六:视频提取与匹配 (Video Extraction) (2) 支柱八:物理动画 (Physics-based Animation) (1 🔗1) 支柱七:动作重定向 (Motion Retargeting) (1)

🔬 支柱三:空间感知与语义 (Perception & Semantics) (11 篇)

#题目一句话要点标签🔗
1 AnythingReality: Robust Online Gaussian Splatting SLAM for Open-Vocabulary VR Scene Exploration 提出在线高斯点云SLAM以解决VR场景探索中的噪声问题 3D gaussian splatting gaussian splatting splatting
2 What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility 提出Co-VGGT以解决3D重建中的共视问题 3D reconstruction VGGT large language model
3 Glob3R: Global Structure-from-Motion with 3D Foundation Models 提出Glob3R以解决3D重建中的不一致性问题 3D reconstruction VGGT foundation model
4 Promptable Concept Segmentation from Above: Evaluating SAM 3's Zero-Shot and One-Shot Capabilities in Remote Sensing 提出SAM 3以解决遥感图像的零-shot与一-shot分割问题 open-vocabulary open vocabulary foundation model
5 Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models 提出复杂社会行为数据集以评估视觉语言模型的准确性与错误类型 scene understanding large language model multimodal
6 Rethinking Monocular Depth Embedding for Generalized Stereo Matching 提出单目深度嵌入方法以解决立体匹配的泛化问题 monocular depth
7 4D Human-Scene Reconstruction from Low-Overlap Captures 提出StudioRecon以解决低重叠相机下的人体场景重建问题 scene reconstruction
8 Hydra++: Real-Time Hierarchical 3D Scene Graph Construction With Object-Level Shape Estimation 提出Hydra++以解决3D场景图构建中的形状估计问题 sam 3D SAM 3D
9 C-GAP: Class-Aware and Online Prompting Improves Vision-Language Models on Imbalanced Classes 提出C-GAP以解决长尾类物体检测问题 open-vocabulary open vocabulary
10 Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation 提出层次化框架以解决长时间音乐到舞蹈生成问题 optical flow
11 DGSfM: Depth-Guided Scale-Aware Global Structure-from-Motion 提出DGSfM以解决全局结构光束法中的尺度模糊问题 monocular depth

🔬 支柱二:RL算法与架构 (RL & Architecture) (11 篇)

#题目一句话要点标签🔗
12 Toward Active Object Detection for UAVs in the Wild: A Large-Scale Dataset, Benchmark and Method 提出ATRNet-LUDO数据集以解决无人机主动目标检测问题 reinforcement learning deep reinforcement learning DRL
13 ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts 提出ALICE模型以整合多模态病理学知识 distillation foundation model multimodal
14 Joint-Embedding Predictive Architecture for Solar PV Panel Fault Classification 提出JEFFNet以解决太阳能光伏面板故障分类问题 JEPA Joint-Embedding Predictive Architecture joint-embedding predictive architecture
15 Video Generation Models are General-Purpose Vision Learners 提出GenCeption以实现通用视觉智能 JEPA MAE Depth Anything
16 Causally Debiased Latent Action Model for Embodied Action Conditioned World Models 提出CD-LAM以解决可控动作模型中的偏差问题 world model world models contrastive learning
17 Multimodal Scenario Similarity Search for Autonomous Driving 提出多模态框架以提升自动驾驶场景检索效率 contrastive learning multimodal
18 Scalable Visual Pretraining for Language Intelligence 提出可扩展视觉预训练以提升语言智能 visual pre-training foundation model
19 Weaving Light and Time: Unified Harmonic-Geometric Representation Learning for Dense RGB-Event Parsing 提出Evita以解决RGB-事件解析中的多模态融合问题 representation learning multimodal
20 IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation 提出信息瓶颈引导的CFG蒸馏方法以解决文本到图像生成中的推理延迟问题 distillation classifier-free guidance
21 Probing Diffusion Denoising Dynamics for Contrastive Representation Learning 提出D$^3$CL以支持对比表示学习的去噪动态适应 representation learning contrastive learning
22 MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models 提出MOSAIC以解决异构视觉语言模型的优化问题 linear attention distillation

🔬 支柱九:具身大模型 (Embodied Foundation Models) (7 篇)

#题目一句话要点标签🔗
23 SVF-CR: Synchronized Visual-Facial Cross-Refinement for Multimodal Ambivalence and Hesitancy Recognition 提出SVF-CR框架以解决多模态模糊与犹豫识别问题 multimodal
24 Integrating Large Language Models and Graph Convolutional Networks for Semi-Supervised Image Classification 提出结合大语言模型与图卷积网络的半监督图像分类方法 large language model
25 The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs 提出检测引导自我修正方法以解决VLM计数错误问题 multimodal
26 Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference 揭示边缘VLM推理中的真实能耗瓶颈 embodied AI
27 SigLIP-HD by Fine-to-Coarse Supervision 提出SigLIP-HD以解决低成本高质量视觉感知问题 multimodal
28 REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation 提出REBASE以解决训练无关的上下文分割问题 foundation model
29 OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents 提出OmniMapBench以解决视觉中心推理的基准问题 visual grounding

🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)

#题目一句话要点标签🔗
30 TSR-Ego: Temporally Guided Stereo Refinement Framework for Egocentric 3D Human Pose Estimation 提出TSR-Ego以解决头戴立体相机下的3D人类姿态估计问题 egocentric
31 DETRAM: End-to-end DEtection, Tracking and Recovery of HumAn Meshes 提出DETRAM以解决多人物体检测与跟踪问题 human mesh recovery HMR

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
32 GeoTrace: Geometry-Aware Trajectory Token Compression for Video Large Language Models 提出GeoTrace以解决视频大语言模型中的视觉令牌压缩问题 spatiotemporal large language model

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
33 Adaptive Latent Trajectory Anchoring for Action Segmentation Dataset Condensation 提出自适应潜在轨迹锚定以解决动作分割数据集蒸馏问题 latent optimization

⬅️ 返回 cs.CV 首页 · 🏠 返回主页