cs.CV(2026-09-09)
📊 共 19 篇论文 | 🔗 2 篇有代码
🎯 兴趣领域导航
支柱二:RL算法与架构 (RL & Architecture) (6)
支柱九:具身大模型 (Embodied Foundation Models) (5 🔗1)
支柱三:空间感知与语义 (Perception & Semantics) (4 🔗1)
支柱一:机器人控制 (Robot Control) (2)
支柱六:视频提取与匹配 (Video Extraction) (1)
支柱四:生成式动作 (Generative Motion) (1)
🔬 支柱二:RL算法与架构 (RL & Architecture) (6 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | RouteBridge: Reliability-Routed Bidirectional Distillation Between Neural Radiance Fields and 3D Gaussian Splatting | 提出RouteBridge以解决NeRF与3DGS间的蒸馏问题 | distillation 3D gaussian splatting 3DGS | ||
| 2 | Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis | 提出Arti-JEPA以解决实时MRI语音生产分析问题 | world model world models JEPA | ||
| 3 | Decoupled Self-Forcing Distillation for Streaming Talking Head Generation | 提出解耦自强蒸馏方法以提升流媒体人头生成质量 | distillation motion latent motion representation | ||
| 4 | Programmable World Model | 提出可编程世界模型以解决持久状态管理问题 | world model world models spatiotemporal | ||
| 5 | Multimodal Emotion Recognition in Conversations via Class-Wise Adaptive Modality Fusion and Affective Geometry | 提出多模态情感识别方法以解决对话中的情感动态问题 | distillation multimodal | ||
| 6 | Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs | 提出VIP-Router以解决视觉令牌剪枝中的适应性问题 | VIP large language model multimodal |
🔬 支柱九:具身大模型 (Embodied Foundation Models) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 7 | Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval | 提出基于基础模型的高效训练无关视频检索方法 | foundation model multimodal | ✅ | |
| 8 | BrainTaskonomy: Learning How to Pretrain and What to Transfer in fMRI Foundation Models | 提出BrainTaskonomy以优化fMRI预训练和任务迁移 | foundation model | ||
| 9 | AVSRBench: A Multi-Condition AVSR Benchmark | 提出AVSRBench以解决多条件下的AVSR泛化问题 | multimodal | ||
| 10 | Geometry Without Coordinates: LiDAR Diffusion as a 3D Feature Bridge | 提出基于LiDAR的扩散模型以解决3D特征提取问题 | foundation model | ||
| 11 | LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios | 提出LogiScope-VQA以解决工业场景中的物流危险识别问题 | multimodal |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (4 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 12 | LinearMask-GS: Stable-Mask Importance Pruning for Compact 3D Gaussian Splatting | 提出LinearMask-GS以解决3D高斯点云冗余问题 | 3D gaussian splatting 3DGS gaussian splatting | ||
| 13 | Shape-guided Gaussian Splatting for Sparse-View X-ray 3D Reconstruction | 提出形状引导的高斯点云方法以解决稀视角X射线3D重建问题 | 3D gaussian splatting 3D reconstruction gaussian splatting | ✅ | |
| 14 | Vague2Detect: Handling Ambiguous Prompts in Knowledge-Based Open-World Detection | 提出Vague2Detect以解决模糊提示在开放世界检测中的问题 | open-vocabulary open vocabulary large language model | ||
| 15 | Guiding Image-to-3D Generation with Test-Time Partial Observations | 提出无训练框架以利用测试时部分几何观测指导图像到3D生成 | sam 3D SAM 3D |
🔬 支柱一:机器人控制 (Robot Control) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | Isotropic Embedding Perturbations for Robust Vision Language Encoders | 提出Aether方法以解决多模态模型数据增强不足问题 | manipulation multimodal | ||
| 17 | PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving | 提出PACE框架以优化对话服务中的感知延迟问题 | humanoid humanoid robot |
🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 18 | 3rd Place Solution to Human Motion Challenges in Real-World and Clinical Settings (MoCha) @ECCV2026: Language-Aligned Motion Representations for Domain-Generalizable UPDRS-Gait Severity Estimation | 提出语言对齐运动表示以解决UPDRS步态严重性估计问题 | SMPL human motion motion representation |
🔬 支柱四:生成式动作 (Generative Motion) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 19 | SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination | 提出SceneHI框架以解决高分辨率3D纹理合成问题 | physically plausible |