cs.CV(2023-10-29)
📊 共 14 篇论文 | 🔗 7 篇有代码
🎯 兴趣领域导航
支柱三:空间感知与语义 (Perception & Semantics) (6 🔗2)
支柱二:RL算法与架构 (RL & Architecture) (5 🔗3)
支柱九:具身大模型 (Embodied Foundation Models) (2 🔗2)
支柱七:动作重定向 (Motion Retargeting) (1)
🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic Segmentation | 提出非可学习原型正则化以解决弱开放词汇语义分割问题 | open-vocabulary open vocabulary | ✅ | |
| 2 | Video Frame Interpolation with Many-to-many Splatting and Spatial Selective Refinement | 提出多对多溅射框架以解决视频帧插值问题 | splatting motion estimation | ||
| 3 | TivNe-SLAM: Dynamic Mapping and Tracking via Time-Varying Neural Radiance Fields | 提出TivNe-SLAM以解决动态场景映射与跟踪问题 | NeRF neural radiance field | ||
| 4 | Dynamo-Depth: Fixing Unsupervised Depth Estimation for Dynamical Scenes | 提出Dynamo-Depth以解决动态场景下的无监督深度估计问题 | depth estimation monocular depth | ||
| 5 | DynPoint: Dynamic Neural Point For View Synthesis | 提出DynPoint以解决长视频视图合成问题 | neural radiance field scene flow | ||
| 6 | 3DMiner: Discovering Shapes from Large-Scale Unannotated Image Datasets | 提出3DMiner以从大规模无标注图像数据集中挖掘3D形状 | 3D reconstruction | ✅ |
🔬 支柱二:RL算法与架构 (RL & Architecture) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 7 | Dynamic Task and Weight Prioritization Curriculum Learning for Multimodal Imagery | 提出动态任务与权重优先级课程学习以提升多模态图像分析性能 | curriculum learning multimodal | ✅ | |
| 8 | BirdSAT: Cross-View Contrastive Masked Autoencoders for Bird Species Classification and Mapping | 提出BirdSAT框架以解决鸟类物种分类与生态映射问题 | masked autoencoder contrastive learning | ✅ | |
| 9 | Reward Finetuning for Faster and More Accurate Unsupervised Object Discovery | 提出基于奖励微调的方法以加速无监督物体发现 | reinforcement learning RLHF large language model | ||
| 10 | Towards Generalized Multi-stage Clustering: Multi-view Self-distillation | 提出多视角自蒸馏方法以解决多阶段聚类中的伪标签偏差问题 | contrastive learning distillation | ||
| 11 | Identifiable Contrastive Learning with Automatic Feature Importance Discovery | 提出三因子对比学习以解决特征可识别性问题 | contrastive learning | ✅ |
🔬 支柱九:具身大模型 (Embodied Foundation Models) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 12 | Myriad: Large Multimodal Model by Applying Vision Experts for Industrial Anomaly Detection | 提出Myriad以解决工业异常检测中的模态差距问题 | multimodal instruction following | ✅ | |
| 13 | Multimodal ChatGPT for Medical Applications: an Experimental Study of GPT-4V | 评估GPT-4V在医学视觉问答中的应用潜力 | large language model multimodal | ✅ |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 14 | TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding | 提出TESTA以解决长视频编码效率瓶颈问题 | spatial relationship spatiotemporal |