cs.CV(2023-10-11)
📊 共 14 篇论文 | 🔗 7 篇有代码
🎯 兴趣领域导航
支柱三:空间感知与语义 (Perception & Semantics) (7 🔗2)
支柱四:生成式动作 (Generative Motion) (2 🔗2)
支柱九:具身大模型 (Embodied Foundation Models) (2 🔗2)
支柱二:RL算法与架构 (RL & Architecture) (2 🔗1)
支柱六:视频提取与匹配 (Video Extraction) (1)
🔬 支柱三:空间感知与语义 (Perception & Semantics) (7 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | Ferret: Refer and Ground Anything Anywhere at Any Granularity | 提出Ferret以解决多模态图像理解中的空间引用与定位问题 | open-vocabulary open vocabulary large language model | ✅ | |
| 2 | rpcPRF: Generalizable MPI Neural Radiance Field for Satellite Camera | 提出rpcPRF以解决卫星图像新视角合成问题 | NeRF neural radiance field implicit representation | ||
| 3 | Dynamic Appearance Particle Neural Radiance Field | 提出动态外观粒子神经辐射场以解决动态场景建模问题 | NeRF neural radiance field | ✅ | |
| 4 | Echocardiography video synthesis from end diastolic semantic map via diffusion model | 提出基于扩散模型的心脏超声视频合成方法以解决数据集不足问题 | semantic map | ||
| 5 | PoRF: Pose Residual Field for Accurate Neural Surface Reconstruction | 提出Pose Residual Field以解决神经表面重建中的姿态噪声问题 | NeRF implicit representation | ||
| 6 | DESTINE: Dynamic Goal Queries with Temporal Transductive Alignment for Trajectory Prediction | 提出DESTINE以解决多智能体轨迹预测中的动态目标查询问题 | semantic map | ||
| 7 | S4C: Self-Supervised Semantic Scene Completion with Neural Fields | 提出S4C以解决自监督语义场景补全问题 | scene understanding |
🔬 支柱四:生成式动作 (Generative Motion) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 8 | Guided Attention for Interpretable Motion Captioning | 提出引导注意力机制以提升可解释的人体动作字幕生成 | motion generation human motion human motion generation | ✅ | |
| 9 | DeepSimHO: Stable Pose Estimation for Hand-Object Interaction via Physics Simulation | 提出DeepSimHO以解决手-物体交互中的姿态估计不稳定问题 | penetration | ✅ |
🔬 支柱九:具身大模型 (Embodied Foundation Models) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 10 | Uni-paint: A Unified Framework for Multimodal Image Inpainting with Pretrained Diffusion Model | 提出Uni-paint以解决多模态图像修复控制不足的问题 | multimodal | ✅ | |
| 11 | VeCLIP: Improving CLIP Training via Visual-enriched Captions | 提出VeCLIP以解决图像文本对齐中的噪声问题 | large language model | ✅ |
🔬 支柱二:RL算法与架构 (RL & Architecture) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 12 | Heuristic Vision Pre-Training with Self-Supervised and Supervised Multi-Task Learning | 提出一种多任务学习框架以提升视觉模型的预训练效果 | representation learning contrastive learning foundation model | ||
| 13 | A Discrepancy Aware Framework for Robust Anomaly Detection | 提出一种差异感知框架以解决缺陷检测的鲁棒性问题 | teacher-student distillation | ✅ |
🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 14 | LangNav: Language as a Perceptual Representation for Navigation | 提出LangNav以解决低数据环境下的视觉语言导航问题 | egocentric VLN |