cs.CV(2023-10-14)
📊 共 8 篇论文 | 🔗 1 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (3 🔗1)
支柱三:空间感知与语义 (Perception & Semantics) (2)
支柱二:RL算法与架构 (RL & Architecture) (1)
支柱八:物理动画 (Physics-based Animation) (1)
支柱七:动作重定向 (Motion Retargeting) (1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning | 提出MiniGPT-v2以解决多模态任务统一接口问题 | large language model visual grounding | ||
| 2 | JM3D & JM3D-LLM: Elevating 3D Understanding with Joint Multi-modal Cues | 提出JM3D以解决3D理解中的信息降解与协同不足问题 | large language model multimodal | ✅ | |
| 3 | Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity? | 探讨CLIP的泛化性能是否主要源于高训练-测试相似性 | foundation model |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 4 | Detecting Moving Objects Using a Novel Optical-Flow-Based Range-Independent Invariant | 提出一种新颖的光流基础方法以解决相机运动下的移动物体检测问题 | optical flow | ||
| 5 | Time-based Mapping of Space Using Visual Motion Invariants | 提出基于视觉运动不变量的时间映射方法以解决环境形状不变性问题 | optical flow |
🔬 支柱二:RL算法与架构 (RL & Architecture) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 6 | PaintHuman: Towards High-fidelity Text-to-3D Human Texturing via Denoised Score Distillation | 提出PaintHuman以解决高保真文本到3D人类纹理生成问题 | distillation SMPL |
🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 7 | OBSUM: An object-based spatial unmixing model for spatiotemporal fusion of remote sensing images | 提出OBSUM以解决遥感图像时空融合中的对象信息缺失问题 | spatiotemporal |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 8 | Point-DynRF: Point-based Dynamic Radiance Fields from a Monocular Video | 提出Point-DynRF以解决动态辐射场全局几何表示问题 | geometric consistency |