cs.CV(2026-09-08)

📊 共 36 篇论文 | 🔗 11 篇有代码

🎯 兴趣领域导航

支柱三:空间感知与语义 (Perception & Semantics) (10 🔗5) 支柱九:具身大模型 (Embodied Foundation Models) (10 🔗2) 支柱二:RL算法与架构 (RL & Architecture) (8 🔗1) 支柱六:视频提取与匹配 (Video Extraction) (3 🔗1) 支柱七:动作重定向 (Motion Retargeting) (2 🔗1) 支柱五:交互与反应 (Interaction & Reaction) (1) 支柱一:机器人控制 (Robot Control) (1 🔗1) 支柱八:物理动画 (Physics-based Animation) (1)

🔬 支柱三:空间感知与语义 (Perception & Semantics) (10 篇)

#题目一句话要点标签🔗
1 EdMCGS: Event-Driven Markov Chain Gaussian Splatting for Extreme-Low-Frame-Rate Dynamic Scene Reconstruction 提出EdMCGS以解决极低帧率动态场景重建问题 gaussian splatting splatting scene reconstruction
2 CVT-GS: Learning to Simplify 3D Gaussian Splatting with Centroidal Voronoi Tessellation 提出CVT-GS以解决3D高斯点简化问题 3D gaussian splatting 3DGS gaussian splatting
3 GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting 提出GoDeep以解决无注释开放词汇3D场景理解问题 scene understanding open-vocabulary open vocabulary
4 From Coordinates to Candidate Regions: Temporal Change Localization via Region Selection in Remote Sensing Multimodal LLMs 提出区域选择方法以解决遥感图像中的时变变化定位问题 scene understanding large language model multimodal
5 PhysFlow: Physics-Aware Optical Flow for Motion Controllable Video Generation 提出PhysFlow以解决视频生成中的物理一致性问题 optical flow physically plausible motion representation
6 Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method 提出AGOS-Agent以解决城市环境中的空地协同物体搜索问题 scene understanding embodied AI
7 Segment Any Motion with Radar: Robust Multimodal Moving-Object Segmentation and Tracking 提出RGBTR-Motion基准与SAM-Radar框架以解决多模态运动目标分割与跟踪问题 optical flow multimodal
8 Spheriverse: 3D Scene Understanding from Spherical Observations in the Wild 提出Spheriverse以解决3D场景理解中的球面观测问题 scene understanding semantic mapping semantic map
9 Prior-free relative 6D pose estimation of multiple object instances 提出无先验相对6D姿态估计方法以解决多实例物体定位问题 6D pose estimation multimodal
10 FIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A Minute 提出FIRE3D以实现快速的3D场景重建 scene reconstruction

🔬 支柱九:具身大模型 (Embodied Foundation Models) (10 篇)

#题目一句话要点标签🔗
11 Dreaming in Flow: Generative Grounding Feedback for Self-Evolving Unified Multimodal Models 提出生成性基础反馈以提升统一多模态模型的自我演化能力 multimodal visual grounding
12 Studying Image Tokenizers as Visual Languages in Unified Multimodal Models 提出图像标记器作为视觉语言以优化多模态模型 multimodal
13 DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models 提出DXPR框架以解决多模态场景识别问题 foundation model
14 MorphoOrgaAgent: A Foundation-Model-Based Multi-Agent System for Autonomous Organoid Analysis 提出MorphoOrgaAgent以解决类器官分析中的自动化挑战 foundation model
15 "World Knowledge" in the Weights: Reading Concept Circuits of Vision Transformers 提出跨层转码器以解读视觉变换器中的概念电路 foundation model
16 DSE-VTG: Dual-Side Enhancement for Training-Free Video Temporal Grounding 提出DSE-VTG以解决视频时间定位中的信息瓶颈问题 large language model
17 Leveraging Visual and Geometric Priors for Metric-scale and Complete Vehicle Gaussian Reconstruction from Limited Views 提出一种方法以解决车辆重建中的尺度和完整性问题 foundation model
18 Effects of model architecture and learning strategies on deep learning-based recognition of activated sludge microscopic images and comparison with quantitative image analysis 提出基于变换器的深度学习方法以提升活性污泥显微图像识别精度 foundation model
19 Do Input-Level Defenses Transfer to Observation-Level Attacks on VideoLLMs? 提出DefTEval框架以评估视频LLMs对观察级攻击的防御能力 large language model
20 MARS-CLIP: Multi-Resolution and Attention Refined Zero-Shot Image Segmentation 提出MARS-CLIP以解决CLIP在零-shot图像分割中的局限性 zero-shot transfer

🔬 支柱二:RL算法与架构 (RL & Architecture) (8 篇)

#题目一句话要点标签🔗
21 ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation 提出ReMoMask-2以解决文本到运动生成中的检索与融合问题 contrastive learning text-to-motion motion generation
22 Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation 提出Marigold V2以解决单目深度估计中的泛化问题 flow matching depth estimation monocular depth
23 VeriScene: Reconstructing Crime Scenes from Legal Evidence via World-Model Agent 提出VeriScene以重建犯罪现场并解决证据整合问题 world model world models physically plausible
24 SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators 提出SyncWorld以解决机器人世界模型的可扩展性问题 world model world models
25 Hi-FLoop: Hierarchical State-Feedback Loops for Multi-Timescale World Modeling 提出HI-FLOOP以解决多智能体交通模拟中的时间尺度一致性问题 world model world models
26 ActionSplice: In-Flight Action Editing for Interactive World Models 提出ActionSplice以解决视频生成中的动作编辑问题 world model world models
27 From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video 提出Coherent4D与HIGFlow以解决4D交互预测问题 flow matching egocentric
28 Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout 提出Mask Forcing以解决视频生成中的模式崩溃问题 distillation

🔬 支柱六:视频提取与匹配 (Video Extraction) (3 篇)

#题目一句话要点标签🔗
29 Human-Centric Image Captioning with Subject-Centered Spatial Understanding 提出SPACE基准以解决人本图像描述中的空间理解问题 egocentric large language model multimodal
30 SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation 提出SynthGait-19K以解决步态参数估计数据不足问题 human mesh recovery HMR SMPL
31 GOLF: Global Observation with Local Focus for Calibration-Aware Stereo Interaction Field Estimation 提出GOLF以解决手部交互场景中的3D向量估计问题 egocentric

🔬 支柱七:动作重定向 (Motion Retargeting) (2 篇)

#题目一句话要点标签🔗
32 Point4D: Long-range 4D Motion Reconstruction 提出Point4D以解决长视频序列的4D重建问题 motion reconstruction
33 DriveMotion: A Large-Scale Multi-Source Benchmark for Driver Motion Sequence Modeling and Forecasting 提出DriveMotion以解决驾驶员动作序列建模与预测问题 human motion

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
34 AXS-Net: Interpretable Deep Unfolding for Hyperspectral Image Denoising via Spectral Basis Unmixing and Structured Noise Refinement 提出AXS-Net以解决高光谱图像去噪问题 HSI PULSE zero-shot transfer

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
35 FPicker: Topology-Guided Evolution for Filament Tracing in Low-SNR Microscopy 提出FPicker以解决低信噪比显微镜中细丝追踪问题 sim-to-real

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
36 WSPolypNet: Weakly Supervised Polyp Localization in Colonoscopy Videos 提出WSPolypNet以解决结肠镜视频中多发性息肉定位问题 spatiotemporal

⬅️ 返回 cs.CV 首页 · 🏠 返回主页