cs.CV(2023-10-30)
📊 共 16 篇论文 | 🔗 6 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (8 🔗4)
支柱二:RL算法与架构 (RL & Architecture) (3 🔗2)
支柱三:空间感知与语义 (Perception & Semantics) (2)
支柱七:动作重定向 (Motion Retargeting) (1)
支柱四:生成式动作 (Generative Motion) (1)
支柱一:机器人控制 (Robot Control) (1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (8 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | Promise:Prompt-driven 3D Medical Image Segmentation Using Pretrained Image Foundation Models | 提出ProMISe以解决医学图像分割中的数据获取与标签可用性问题 | foundation model | ✅ | |
| 2 | Harvest Video Foundation Models via Efficient Post-Pretraining | 提出高效后预训练框架以获取视频基础模型 | foundation model | ✅ | |
| 3 | Are Natural Domain Foundation Models Useful for Medical Image Classification? | 探讨自然领域基础模型在医学图像分类中的有效性 | foundation model | ||
| 4 | MM-VID: Advancing Video Understanding with GPT-4V(ision) | 提出MM-VID以解决长视频理解与复杂任务挑战 | large language model multimodal | ||
| 5 | RGB-X Object Detection via Scene-Specific Fusion Modules | 提出RGB-X融合网络以解决多模态传感器融合问题 | multimodal | ✅ | |
| 6 | Res-Tuning: A Flexible and Efficient Tuning Paradigm via Unbinding Tuner from Backbone | 提出Res-Tuning以解决现有调优方法的灵活性不足问题 | foundation model | ||
| 7 | Intra-Modal Proxy Learning for Zero-Shot Visual Categorization with CLIP | 提出InMaP以解决视觉分类中的模态差距问题 | zero-shot transfer | ✅ | |
| 8 | VideoCrafter1: Open Diffusion Models for High-Quality Video Generation | 提出开源扩散模型以解决高质量视频生成问题 | foundation model |
🔬 支柱二:RL算法与架构 (RL & Architecture) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 9 | Text-to-3D with Classifier Score Distillation | 提出分类器得分蒸馏方法以提升文本到3D生成效果 | distillation classifier-free guidance | ✅ | |
| 10 | One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation | 提出OFA-KD框架以解决异构架构间知识蒸馏问题 | teacher-student distillation | ✅ | |
| 11 | MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval | 提出MCAD以提升图像-文本检索效率 | distillation |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 12 | Dynamic Gaussian Splatting from Markerless Motion Capture can Reconstruct Infants Movements | 提出动态高斯点云重建技术以解决婴儿运动捕捉问题 | gaussian splatting splatting markerless motion capture | ||
| 13 | SeamlessNeRF: Stitching Part NeRFs with Gradient Propagation | 提出SeamlessNeRF以解决多NeRF无缝合成问题 | NeRF neural radiance field |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 14 | GC-MVSNet: Multi-View, Multi-Scale, Geometrically-Consistent Multi-View Stereo | 提出GC-MVSNet以解决多视角立体视觉中的几何一致性问题 | geometric consistency |
🔬 支柱四:生成式动作 (Generative Motion) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 15 | LinFlo-Net: A two-stage deep learning method to generate simulation ready meshes of the heart | 提出LinFlo-Net以解决心脏模型生成中的自穿透问题 | penetration |
🔬 支柱一:机器人控制 (Robot Control) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | Deep-learning-based decomposition of overlapping-sparse images: application at the vertex of neutrino interactions | 提出深度学习方法以解决重叠稀疏图像分解问题 | manipulation |