cs.CV(2026-09-04)

📊 共 30 篇论文 | 🔗 8 篇有代码

🎯 兴趣领域导航

支柱三:空间感知与语义 (Perception & Semantics) (8 🔗2) 支柱二:RL算法与架构 (RL & Architecture) (7 🔗5) 支柱九:具身大模型 (Embodied Foundation Models) (6) 支柱一:机器人控制 (Robot Control) (4) 支柱六:视频提取与匹配 (Video Extraction) (3) 支柱八:物理动画 (Physics-based Animation) (2 🔗1)

🔬 支柱三:空间感知与语义 (Perception & Semantics) (8 篇)

#题目一句话要点标签🔗
1 Compact Neural Appearance Models for Efficient Gaussian Splatting 提出紧凑神经外观模型以提高高斯点云渲染效率 3D gaussian splatting gaussian splatting splatting
2 Temporal Residual Neural Radiance Fields for Monocular Video Dynamic Human Body Reconstruction 提出时序残差神经辐射场以解决动态人类体重建问题 NeRF neural radiance field scene reconstruction
3 CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation 提出CrossDepth以解决多视角深度估计中的跨图像不一致问题 depth estimation
4 BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular Priors 提出BLASt3R以解决图像集的束调整问题 visual SLAM 3D reconstruction
5 WorldSculpt: Generating Compositional Worlds from Grounded Videos 提出WorldSculpt以解决复杂场景生成问题 3DGS
6 VoxelFix: Post-Hoc Semantic Correction of Completed 3D Voxel Maps 提出VoxelFix以解决3D体素地图语义修正问题 semantic map
7 CLON: Cue-Calibrated Linguistic Object Onboarding for Zero-Shot 6D Pose Front-Ends 提出CLON以解决零-shot 6D姿态估计中的前端问题 6D pose estimation
8 HiSfM: Disambiguating Structure-from-Motion via Scaffold-Anchored Hierarchical Reconstruction 提出HiSfM以解决结构光重建中的视觉歧义问题 3D reconstruction

🔬 支柱二:RL算法与架构 (RL & Architecture) (7 篇)

#题目一句话要点标签🔗
9 Weather-Conditioned Depth Anything 提出Weather-Conditioned Depth Anything以解决恶劣天气下深度估计问题 distillation depth estimation monocular depth
10 MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation 提出MEOX以解决地球观测中的多模态学习问题 representation learning masked autoencoder multimodal
11 Learning 3D Editing without Paired Supervision via Generative Prior Distillation 提出无配对监督的3D编辑学习方法以解决数据稀缺问题 distillation foundation model instruction following
12 First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves 提出FTF-rl以解决多优先级用户需求的推理问题 reinforcement learning large language model multimodal
13 Importance-Aware Low-Rank Distillation of Diffusion Transformers 提出SVDtrunc以解决扩散变换器的参数压缩问题 distillation large language model
14 BEAM3R: Beam's-eye-view architecture with Mamba-3 for implicit dose reconstruction 提出BEAM3R以解决剂量重建中的高效性与准确性问题 Mamba MAE
15 SeamFlow: Structure-Aware Flow Matching on Edge Probabilities for Artist-Like UV Unwrapping 提出SeamFlow以解决3D表面切割与UV展开问题 flow matching

🔬 支柱九:具身大模型 (Embodied Foundation Models) (6 篇)

#题目一句话要点标签🔗
16 MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression 提出MCPO以解决多模态推理链压缩问题 multimodal chain-of-thought
17 Conserved Immune Topology Improves Pathology Foundation Model Generalization for Cross-Cancer MSI-H Prediction 提出保守免疫拓扑以解决跨癌症MSI-H预测问题 foundation model
18 SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis 提出自解释多模态信息瓶颈以解决医疗诊断中的可解释性问题 multimodal
19 WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing 提出WeAgent-MMGenEdit以解决多模态图像生成与编辑中的知识依赖问题 multimodal
20 Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding 提出IVSGround以解决3D视觉定位中的视角选择问题 visual grounding
21 LetOccVote: Learning Weakly Supervised 3D Occupancy through Consensus 提出LetOccVote以解决弱监督3D占用预测问题 foundation model

🔬 支柱一:机器人控制 (Robot Control) (4 篇)

#题目一句话要点标签🔗
22 TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image 提出TourPhysics以解决长视频合成中的物理一致性问题 manipulation world model world models
23 UniMate: One Unified Model to Animate Diverse Skeletons 提出UniMate以解决多样化骨骼动画生成问题 quadruped bipedal biped
24 An Evaluation Framework for Generating Multi-View Images of a Person in a Scene 提出HSRD度量以解决多视角图像生成中的一致性问题 manipulation foundation model
25 LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Rendering 提出LensStyle以解决镜头效果渲染的多样性与可控性问题 manipulation

🔬 支柱六:视频提取与匹配 (Video Extraction) (3 篇)

#题目一句话要点标签🔗
26 MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision 提出MINT模型以解决世界坐标下相机与手部运动估计问题 egocentric hand reconstruction motion estimation
27 Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents 提出语言轨迹编码以解决长时间空间记忆问题 Ego4D
28 Hidden In Plain Gaze: Gaze Representations as Privacy Controls for Utility and Re-identification Risk in XR 提出基于注视表示的隐私控制以降低XR中的身份重识别风险 egocentric

🔬 支柱八:物理动画 (Physics-based Animation) (2 篇)

#题目一句话要点标签🔗
29 Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion 提出多特征编码与注意力融合以增强多模态情感识别 spatiotemporal multimodal
30 Sound-based Multi-Person 3D Pose Estimation 提出SoundMHPE以解决声学信号下的多人体姿态估计问题 AMP

⬅️ 返回 cs.CV 首页 · 🏠 返回主页