cs.CV(2026-09-04)
📊 共 30 篇论文 | 🔗 8 篇有代码
🎯 兴趣领域导航
支柱三:空间感知与语义 (Perception & Semantics) (8 🔗2)
支柱二:RL算法与架构 (RL & Architecture) (7 🔗5)
支柱九:具身大模型 (Embodied Foundation Models) (6)
支柱一:机器人控制 (Robot Control) (4)
支柱六:视频提取与匹配 (Video Extraction) (3)
支柱八:物理动画 (Physics-based Animation) (2 🔗1)
🔬 支柱三:空间感知与语义 (Perception & Semantics) (8 篇)
🔬 支柱二:RL算法与架构 (RL & Architecture) (7 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 9 | Weather-Conditioned Depth Anything | 提出Weather-Conditioned Depth Anything以解决恶劣天气下深度估计问题 | distillation depth estimation monocular depth | ✅ | |
| 10 | MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation | 提出MEOX以解决地球观测中的多模态学习问题 | representation learning masked autoencoder multimodal | ||
| 11 | Learning 3D Editing without Paired Supervision via Generative Prior Distillation | 提出无配对监督的3D编辑学习方法以解决数据稀缺问题 | distillation foundation model instruction following | ✅ | |
| 12 | First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves | 提出FTF-rl以解决多优先级用户需求的推理问题 | reinforcement learning large language model multimodal | ✅ | |
| 13 | Importance-Aware Low-Rank Distillation of Diffusion Transformers | 提出SVDtrunc以解决扩散变换器的参数压缩问题 | distillation large language model | ✅ | |
| 14 | BEAM3R: Beam's-eye-view architecture with Mamba-3 for implicit dose reconstruction | 提出BEAM3R以解决剂量重建中的高效性与准确性问题 | Mamba MAE | ||
| 15 | SeamFlow: Structure-Aware Flow Matching on Edge Probabilities for Artist-Like UV Unwrapping | 提出SeamFlow以解决3D表面切割与UV展开问题 | flow matching | ✅ |
🔬 支柱九:具身大模型 (Embodied Foundation Models) (6 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression | 提出MCPO以解决多模态推理链压缩问题 | multimodal chain-of-thought | ||
| 17 | Conserved Immune Topology Improves Pathology Foundation Model Generalization for Cross-Cancer MSI-H Prediction | 提出保守免疫拓扑以解决跨癌症MSI-H预测问题 | foundation model | ||
| 18 | SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis | 提出自解释多模态信息瓶颈以解决医疗诊断中的可解释性问题 | multimodal | ||
| 19 | WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing | 提出WeAgent-MMGenEdit以解决多模态图像生成与编辑中的知识依赖问题 | multimodal | ||
| 20 | Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding | 提出IVSGround以解决3D视觉定位中的视角选择问题 | visual grounding | ||
| 21 | LetOccVote: Learning Weakly Supervised 3D Occupancy through Consensus | 提出LetOccVote以解决弱监督3D占用预测问题 | foundation model |
🔬 支柱一:机器人控制 (Robot Control) (4 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 22 | TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image | 提出TourPhysics以解决长视频合成中的物理一致性问题 | manipulation world model world models | ||
| 23 | UniMate: One Unified Model to Animate Diverse Skeletons | 提出UniMate以解决多样化骨骼动画生成问题 | quadruped bipedal biped | ||
| 24 | An Evaluation Framework for Generating Multi-View Images of a Person in a Scene | 提出HSRD度量以解决多视角图像生成中的一致性问题 | manipulation foundation model | ||
| 25 | LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Rendering | 提出LensStyle以解决镜头效果渲染的多样性与可控性问题 | manipulation |
🔬 支柱六:视频提取与匹配 (Video Extraction) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 26 | MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision | 提出MINT模型以解决世界坐标下相机与手部运动估计问题 | egocentric hand reconstruction motion estimation | ||
| 27 | Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents | 提出语言轨迹编码以解决长时间空间记忆问题 | Ego4D | ||
| 28 | Hidden In Plain Gaze: Gaze Representations as Privacy Controls for Utility and Re-identification Risk in XR | 提出基于注视表示的隐私控制以降低XR中的身份重识别风险 | egocentric |
🔬 支柱八:物理动画 (Physics-based Animation) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 29 | Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion | 提出多特征编码与注意力融合以增强多模态情感识别 | spatiotemporal multimodal | ||
| 30 | Sound-based Multi-Person 3D Pose Estimation | 提出SoundMHPE以解决声学信号下的多人体姿态估计问题 | AMP | ✅ |