cs.CV(2023-10-10)
📊 共 12 篇论文 | 🔗 1 篇有代码
🎯 兴趣领域导航
支柱三:空间感知与语义 (Perception & Semantics) (4 🔗1)
支柱九:具身大模型 (Embodied Foundation Models) (3)
支柱二:RL算法与架构 (RL & Architecture) (3)
支柱一:机器人控制 (Robot Control) (2)
🔬 支柱三:空间感知与语义 (Perception & Semantics) (4 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | Leveraging Neural Radiance Fields for Uncertainty-Aware Visual Localization | 提出利用神经辐射场提升视觉定位的不确定性感知 | NeRF neural radiance field | ||
| 2 | High-Fidelity 3D Head Avatars Reconstruction through Spatially-Varying Expression Conditioned Neural Radiance Field | 提出空间变化表情条件神经辐射场以解决3D头像重建中的细节保留问题 | NeRF neural radiance field | ||
| 3 | SketchBodyNet: A Sketch-Driven Multi-faceted Decoder Network for 3D Human Reconstruction | 提出SketchBodyNet以解决从自由手绘草图重建3D人形的问题 | 3D reconstruction SMPL | ||
| 4 | TextPSG: Panoptic Scene Graph Generation from Textual Descriptions | 提出TextPSG以解决从文本描述生成全景场景图的问题 | scene understanding | ✅ |
🔬 支柱九:具身大模型 (Embodied Foundation Models) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 5 | CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding | 提出CoT3DRef以解决3D视觉定位中的可解释性与数据效率问题 | visual grounding chain-of-thought | ||
| 6 | Uni3D: Exploring Unified 3D Representation at Scale | 提出Uni3D以解决3D对象和场景统一表示的挑战 | foundation model | ||
| 7 | On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets | 提出VLIT数据集评估与优化方法以提升多模态指令调优模型性能 | multimodal |
🔬 支柱二:RL算法与架构 (RL & Architecture) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 8 | Computational Pathology at Health System Scale -- Self-Supervised Foundation Models from Three Billion Images | 提出自监督基础模型以解决医学病理数据稀缺问题 | masked autoencoder MAE foundation model | ||
| 9 | Data Distillation Can Be Like Vodka: Distilling More Times For Better Quality | 提出渐进式数据蒸馏以提升深度网络训练性能 | distillation | ||
| 10 | Distillation Improves Visual Place Recognition for Low Quality Images | 提出知识蒸馏方法以提升低质量图像的视觉位置识别 | distillation |
🔬 支柱一:机器人控制 (Robot Control) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 11 | Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models | 提出一种新方法以实现开放词汇的视频目标跟踪 | manipulation scene understanding open-vocabulary | ||
| 12 | Perceptual MAE for Image Manipulation Localization: A High-level Vision Learner Focusing on Low-level Features | 提出感知MAE以解决图像篡改定位问题 | manipulation masked autoencoder MAE |