cs.CV(2023-10-20)
📊 共 13 篇论文 | 🔗 4 篇有代码
🎯 兴趣领域导航
支柱三:空间感知与语义 (Perception & Semantics) (4 🔗3)
支柱九:具身大模型 (Embodied Foundation Models) (3)
支柱二:RL算法与架构 (RL & Architecture) (3 🔗1)
支柱七:动作重定向 (Motion Retargeting) (1)
支柱六:视频提取与匹配 (Video Extraction) (1)
支柱八:物理动画 (Physics-based Animation) (1)
🔬 支柱三:空间感知与语义 (Perception & Semantics) (4 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | OpenAnnotate3D: Open-Vocabulary Auto-Labeling System for Multi-modal 3D Data | 提出OpenAnnotate3D以解决多模态3D数据的自动标注问题 | open-vocabulary open vocabulary embodied AI | ||
| 2 | UE4-NeRF:Neural Radiance Field for Real-Time Rendering of Large-Scale Scene | 提出UE4-NeRF以解决大规模场景实时渲染问题 | 3D reconstruction NeRF neural radiance field | ✅ | |
| 3 | Sync-NeRF: Generalizing Dynamic NeRFs to Unsynchronized Videos | 提出Sync-NeRF以解决动态NeRF在非同步视频中的重建问题 | NeRF neural radiance field scene reconstruction | ✅ | |
| 4 | ManifoldNeRF: View-dependent Image Feature Supervision for Few-shot Neural Radiance Fields | 提出ManifoldNeRF以解决少量图像下的视角依赖特征监督问题 | NeRF neural radiance field | ✅ |
🔬 支柱九:具身大模型 (Embodied Foundation Models) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 5 | Benchmarking Sequential Visual Input Reasoning and Prediction in Multimodal Large Language Models | 提出多模态大语言模型预测推理基准以解决能力不足问题 | large language model multimodal | ||
| 6 | Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine Translation | 提出一种方法以解决合成图像与真实图像之间的差距问题 | multimodal | ||
| 7 | Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds | 提出Steve-Eye以解决LLM代理视觉感知不足问题 | large language model multimodal |
🔬 支柱二:RL算法与架构 (RL & Architecture) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 8 | SILC: Improving Vision Language Pretraining with Self-Distillation | 提出SILC框架以提升视觉语言预训练效果 | contrastive learning distillation open-vocabulary | ||
| 9 | Data-Free Knowledge Distillation Using Adversarially Perturbed OpenGL Shader Images | 提出一种新方法以解决无数据知识蒸馏问题 | distillation | ||
| 10 | Learning with Unmasked Tokens Drives Stronger Vision Learners | 通过引入未掩码标记提升视觉学习者的表现 | masked autoencoder MAE | ✅ |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 11 | PACE: Human and Camera Motion Estimation from in-the-wild Videos | 提出一种新方法以解决动态视频中的人类与相机运动估计问题 | human motion motion estimation |
🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 12 | FMRT: Learning Accurate Feature Matching with Reconciliatory Transformer | 提出FMRT以解决特征匹配精度不足问题 | feature matching |
🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 13 | Dance Your Latents: Consistent Dance Generation through Spatial-temporal Subspace Attention Guided by Motion Flow | 提出Dance-Your-Latents以解决舞蹈生成中的时空一致性问题 | spatiotemporal |