cs.CV(2023-10-08)
📊 共 13 篇论文 | 🔗 4 篇有代码
🎯 兴趣领域导航
支柱三:空间感知与语义 (Perception & Semantics) (7 🔗2)
支柱九:具身大模型 (Embodied Foundation Models) (3)
支柱二:RL算法与架构 (RL & Architecture) (2 🔗1)
支柱七:动作重定向 (Motion Retargeting) (1 🔗1)
🔬 支柱三:空间感知与语义 (Perception & Semantics) (7 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching | 提出开放词汇动物关键点检测方法以解决零-shot检测问题 | open-vocabulary open vocabulary feature matching | ||
| 2 | Building an Open-Vocabulary Video CLIP Model with Better Architectures, Optimization and Data | 提出Open-VCLIP++以解决零-shot视频识别问题 | open-vocabulary open vocabulary large language model | ✅ | |
| 3 | OV-PARTS: Towards Open-Vocabulary Part Segmentation | 提出OV-PARTS以解决开放词汇的部件分割问题 | open-vocabulary open vocabulary | ✅ | |
| 4 | Compositional Semantics for Open Vocabulary Spatio-semantic Representations | 提出潜在组合语义嵌入以解决复杂任务推理问题 | open-vocabulary open vocabulary | ||
| 5 | LocoNeRF: A NeRF-based Approach for Local Structure from Motion for Precise Localization | 提出LocoNeRF以解决视觉定位中的高延迟问题 | NeRF neural radiance field | ||
| 6 | Geometry Aware Field-to-field Transformations for 3D Semantic Segmentation | 提出基于NeRF的几何感知场到场转换以解决3D语义分割问题 | NeRF neural radiance field | ||
| 7 | Improving Discriminative Multi-Modal Learning with Large-Scale Pre-Trained Models | 提出MMLoRA以解决多模态学习中的特征适应问题 | optical flow |
🔬 支柱九:具身大模型 (Embodied Foundation Models) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 8 | UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model | 提出UReader以解决视觉场景下的无OCR语言理解问题 | large language model multimodal | ||
| 9 | Lightweight In-Context Tuning for Multimodal Unified Models | 提出M$^2$IXT以增强多模态统一模型的上下文学习能力 | multimodal visual grounding | ||
| 10 | Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling | 提出Video-Teller以提升视频到文本生成的效果 | large language model foundation model |
🔬 支柱二:RL算法与架构 (RL & Architecture) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 11 | VisionFM: a Multi-Modal Multi-Task Vision Foundation Model for Generalist Ophthalmic Artificial Intelligence | 提出VisionFM以解决眼科人工智能多任务问题 | representation learning foundation model | ||
| 12 | Symmetrical Linguistic Feature Distillation with CLIP for Scene Text Recognition | 提出CLIP-OCR以提升场景文本识别性能 | distillation | ✅ |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 13 | AANet: Aggregation and Alignment Network with Semi-hard Positive Sample Mining for Hierarchical Place Recognition | 提出AANet以解决层次化地点识别中的效率与准确性问题 | geometric consistency | ✅ |