cs.CV(2023-10-02)
📊 共 21 篇论文 | 🔗 7 篇有代码
🎯 兴趣领域导航
支柱三:空间感知与语义 (Perception & Semantics) (7 🔗3)
支柱九:具身大模型 (Embodied Foundation Models) (5 🔗1)
支柱一:机器人控制 (Robot Control) (3)
支柱二:RL算法与架构 (RL & Architecture) (3 🔗2)
支柱四:生成式动作 (Generative Motion) (2 🔗1)
支柱七:动作重定向 (Motion Retargeting) (1)
🔬 支柱三:空间感知与语义 (Perception & Semantics) (7 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | PC-NeRF: Parent-Child Neural Radiance Fields under Partial Sensor Data Loss in Autonomous Driving Environments | 提出PC-NeRF以解决自动驾驶环境中部分传感器数据丢失问题 | 3D reconstruction NeRF neural radiance field | ✅ | |
| 2 | CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction | 提出CLIPSelf以解决CLIP在开放词汇密集预测中的区域语言对齐问题 | open-vocabulary open vocabulary | ✅ | |
| 3 | DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection | 提出DST-Det以解决开放词汇目标检测问题 | open-vocabulary open vocabulary | ✅ | |
| 4 | Task-guided Domain Gap Reduction for Monocular Depth Prediction in Endoscopy | 提出任务引导的领域间差距缩减方法以改善内窥镜单目深度预测 | monocular depth | ||
| 5 | Adaptive Visual Scene Understanding: Incremental Scene Graph Generation | 提出增量场景图生成方法以解决动态视觉理解问题 | scene understanding | ||
| 6 | Segmenting the motion components of a video: A long-term unsupervised model | 提出长时间无监督模型以实现视频运动分割 | optical flow | ||
| 7 | Multi-task Learning with 3D-Aware Regularization | 提出3D感知正则化以解决多任务学习中的噪声问题 | depth estimation |
🔬 支柱九:具身大模型 (Embodied Foundation Models) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 8 | DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model | 提出DriveGPT4以解决自主驾驶中的可解释性问题 | large language model multimodal | ||
| 9 | HyMNet: a Multimodal Deep Learning System for Hypertension Classification using Fundus Photographs and Cardiometabolic Risk Factors | 提出HyMNet以解决单一数据源对高血压分类的局限性 | foundation model multimodal | ✅ | |
| 10 | Less is More: Toward Zero-Shot Local Scene Graph Generation via Foundation Models | 提出ELEGANT框架以解决零样本局部场景图生成问题 | foundation model | ||
| 11 | Making LLaMA SEE and Draw with SEED Tokenizer | 提出SEED图像分词器以解决多模态统一处理问题 | large language model multimodal | ||
| 12 | NEUCORE: Neural Concept Reasoning for Composed Image Retrieval | 提出NEUCORE以解决复合图像检索问题 | multimodal |
🔬 支柱一:机器人控制 (Robot Control) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 13 | Intelligent Knee Sleeves: A Real-time Multimodal Dataset for 3D Lower Body Motion Estimation Using Smart Textile | 提出智能膝盖护套以解决3D下肢运动估计问题 | locomotion motion estimation multimodal | ||
| 14 | STARS: Zero-shot Sim-to-Real Transfer for Segmentation of Shipwrecks in Sonar Imagery | 提出STARS以解决零样本船舶残骸分割问题 | sim-to-real | ||
| 15 | Trained Latent Space Navigation to Prevent Lack of Photorealism in Generated Images on Style-based Models | 提出无监督方法以解决生成图像缺乏真实感的问题 | manipulation |
🔬 支柱二:RL算法与架构 (RL & Architecture) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | Toward effective protection against diffusion based mimicry through score distillation | 提出基于得分蒸馏的策略以有效防护扩散模型的模仿攻击 | distillation | ✅ | |
| 17 | Strength in Diversity: Multi-Branch Representation Learning for Vehicle Re-Identification | 提出轻量级多分支深度架构以提升车辆重识别性能 | representation learning | ✅ | |
| 18 | Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality | 提出可学习的跨模态知识蒸馏以解决多模态学习中的缺失模态问题 | distillation |
🔬 支柱四:生成式动作 (Generative Motion) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 19 | Reconstructing 3D Human Pose from RGB-D Data with Occlusions | 提出新方法以解决RGB-D数据中人体重建的遮挡问题 | physically plausible penetration ReMoS | ||
| 20 | Mirror Diffusion Models for Constrained and Watermarked Generation | 提出镜面扩散模型以解决约束数据生成问题 | MDM | ✅ |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 21 | Leveraging Cutting Edge Deep Learning Based Image Matching for Reconstructing a Large Scene from Sparse Images | 提出基于深度学习的图像匹配方法以重建稀疏图像的大场景 | motion reconstruction |