cs.CV(2023-10-13)
📊 共 13 篇论文 | 🔗 1 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (5)
支柱三:空间感知与语义 (Perception & Semantics) (3 🔗1)
支柱二:RL算法与架构 (RL & Architecture) (3)
支柱一:机器人控制 (Robot Control) (2)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models | 提出COMM策略以提升多模态大语言模型的视觉能力 | large language model visual grounding | ||
| 2 | Transformer-based Multimodal Change Detection with Multitask Consistency Constraints | 提出基于Transformer的多模态变化检测以解决多任务一致性问题 | multimodal | ||
| 3 | Vision-by-Language for Training-Free Compositional Image Retrieval | 提出CIReVL以解决训练依赖的组合图像检索问题 | large language model | ||
| 4 | PaLI-3 Vision Language Models: Smaller, Faster, Stronger | 提出PaLI-3以提升多模态理解能力 | multimodal | ||
| 5 | Learning to Adapt SAM for Segmenting Cross-domain Point Clouds | 提出一种方法以解决3D点云跨域分割问题 | foundation model |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 6 | SAIR: Learning Semantic-aware Implicit Representation | 提出SAIR以解决图像重建中的语义信息缺失问题 | implicit representation | ||
| 7 | A Spatial-Temporal Dual-Mode Mixed Flow Network for Panoramic Video Salient Object Detection | 提出空间-时间双模混合流网络以解决全景视频显著目标检测问题 | optical flow | ||
| 8 | TIDE: Temporally Incremental Disparity Estimation via Pattern Flow in Structured Light System | 提出TIDE-Net以解决单目相机结构光系统中的视差估计问题 | optical flow | ✅ |
🔬 支柱二:RL算法与架构 (RL & Architecture) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 9 | Timestamp-supervised Wearable-based Activity Segmentation and Recognition with Contrastive Learning and Order-Preserving Optimal Transport | 提出基于时间戳监督的可穿戴活动分割与识别方法以解决多类窗口问题 | contrastive learning | ||
| 10 | UniParser: Multi-Human Parsing with Unified Correlation Representation Learning | 提出UniParser以解决多人体解析中的信息处理效率问题 | representation learning | ||
| 11 | Extending Multi-modal Contrastive Representations | 提出Ex-MCR以解决多模态对比表示学习中的数据依赖问题 | representation learning multimodal |
🔬 支柱一:机器人控制 (Robot Control) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 12 | Leveraging Image Augmentation for Object Manipulation: Towards Interpretable Controllability in Object-Centric Learning | 提出SlotAug方法以解决对象中心学习中的可解释性控制问题 | manipulation | ||
| 13 | pose-format: Library for Viewing, Augmenting, and Handling .pose Files | 提出pose-format以解决姿态数据管理与分析问题 | manipulation |