cs.CV(2023-10-08)

📊 共 13 篇论文 | 🔗 4 篇有代码

🎯 兴趣领域导航

支柱三:空间感知与语义 (Perception & Semantics) (7 🔗2) 支柱九:具身大模型 (Embodied Foundation Models) (3) 支柱二:RL算法与架构 (RL & Architecture) (2 🔗1) 支柱七:动作重定向 (Motion Retargeting) (1 🔗1)

🔬 支柱三:空间感知与语义 (Perception & Semantics) (7 篇)

#题目一句话要点标签🔗
1 Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching 提出开放词汇动物关键点检测方法以解决零-shot检测问题 open-vocabulary open vocabulary feature matching
2 Building an Open-Vocabulary Video CLIP Model with Better Architectures, Optimization and Data 提出Open-VCLIP++以解决零-shot视频识别问题 open-vocabulary open vocabulary large language model
3 OV-PARTS: Towards Open-Vocabulary Part Segmentation 提出OV-PARTS以解决开放词汇的部件分割问题 open-vocabulary open vocabulary
4 Compositional Semantics for Open Vocabulary Spatio-semantic Representations 提出潜在组合语义嵌入以解决复杂任务推理问题 open-vocabulary open vocabulary
5 LocoNeRF: A NeRF-based Approach for Local Structure from Motion for Precise Localization 提出LocoNeRF以解决视觉定位中的高延迟问题 NeRF neural radiance field
6 Geometry Aware Field-to-field Transformations for 3D Semantic Segmentation 提出基于NeRF的几何感知场到场转换以解决3D语义分割问题 NeRF neural radiance field
7 Improving Discriminative Multi-Modal Learning with Large-Scale Pre-Trained Models 提出MMLoRA以解决多模态学习中的特征适应问题 optical flow

🔬 支柱九:具身大模型 (Embodied Foundation Models) (3 篇)

#题目一句话要点标签🔗
8 UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model 提出UReader以解决视觉场景下的无OCR语言理解问题 large language model multimodal
9 Lightweight In-Context Tuning for Multimodal Unified Models 提出M$^2$IXT以增强多模态统一模型的上下文学习能力 multimodal visual grounding
10 Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling 提出Video-Teller以提升视频到文本生成的效果 large language model foundation model

🔬 支柱二:RL算法与架构 (RL & Architecture) (2 篇)

#题目一句话要点标签🔗
11 VisionFM: a Multi-Modal Multi-Task Vision Foundation Model for Generalist Ophthalmic Artificial Intelligence 提出VisionFM以解决眼科人工智能多任务问题 representation learning foundation model
12 Symmetrical Linguistic Feature Distillation with CLIP for Scene Text Recognition 提出CLIP-OCR以提升场景文本识别性能 distillation

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
13 AANet: Aggregation and Alignment Network with Semi-hard Positive Sample Mining for Hierarchical Place Recognition 提出AANet以解决层次化地点识别中的效率与准确性问题 geometric consistency

⬅️ 返回 cs.CV 首页 · 🏠 返回主页