cs.CV(2023-10-19)
📊 共 16 篇论文 | 🔗 4 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (5 🔗3)
支柱二:RL算法与架构 (RL & Architecture) (5)
支柱七:动作重定向 (Motion Retargeting) (2 🔗1)
支柱一:机器人控制 (Robot Control) (2)
支柱四:生成式动作 (Generative Motion) (1)
支柱三:空间感知与语义 (Perception & Semantics) (1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering | 提出RSAdapter以解决遥感视觉问答中的资源消耗问题 | multimodal | ✅ | |
| 2 | CLAIR: Evaluating Image Captions with Large Language Models | 提出CLAIR以解决图像描述评估的挑战 | large language model | ✅ | |
| 3 | Weakly-Supervised Semantic Segmentation with Image-Level Labels: from Traditional Models to Foundation Models | 提出弱监督语义分割方法以解决图像级标签的挑战 | foundation model | ||
| 4 | 2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level Supervision | 提出多模态交错变换器以解决弱监督点云分割问题 | multimodal | ✅ | |
| 5 | Query-aware Long Video Localization and Relation Discrimination for Deep Video Understanding | 提出查询感知方法以解决长视频定位与关系辨别问题 | multimodal |
🔬 支柱二:RL算法与架构 (RL & Architecture) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 6 | LoMAE: Low-level Vision Masked Autoencoders for Low-dose CT Denoising | 提出LoMAE以解决低剂量CT图像去噪问题 | masked autoencoder MAE | ||
| 7 | Exploiting Low-confidence Pseudo-labels for Source-free Object Detection | 提出低置信度伪标签利用方法以提升源无关目标检测性能 | contrastive learning spatial relationship | ||
| 8 | Representation Learning via Consistent Assignment of Views over Random Partitions | 提出CARP方法以解决自监督聚类中的一致性问题 | representation learning | ||
| 9 | WeedCLR: Weed Contrastive Learning through Visual Representations with Class-Optimized Loss in Long-Tailed Datasets | 提出WeedCLR以解决长尾数据集中的杂草分类问题 | contrastive learning | ||
| 10 | SAM Meets UAP: Attacking Segment Anything Model With Universal Adversarial Perturbation | 提出基于自监督对比学习的UAP生成方法以攻击SAM模型 | contrastive learning foundation model |
🔬 支柱七:动作重定向 (Motion Retargeting) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 11 | Real-Time Motion Prediction via Heterogeneous Polyline Transformer with Relative Pose Encoding | 提出KNARPE机制以解决自主驾驶中的运动预测效率问题 | motion prediction | ✅ | |
| 12 | Multiscale Motion-Aware and Spatial-Temporal-Channel Contextual Coding Network for Learned Video Compression | 提出MASTC-VC以解决视频压缩中的运动估计不准确问题 | motion estimation motion prediction |
🔬 支柱一:机器人控制 (Robot Control) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 13 | 3D-GPT: Procedural 3D Modeling with Large Language Models | 提出3D-GPT以解决程序化3D建模的复杂性问题 | manipulation large language model | ||
| 14 | CycleNet: Rethinking Cycle Consistency in Text-Guided Diffusion for Image Manipulation | 提出CycleNet以解决无配对图像翻译中的一致性问题 | manipulation |
🔬 支柱四:生成式动作 (Generative Motion) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 15 | HumanTOMATO: Text-aligned Whole-body Motion Generation | 提出HumanTOMATO以解决文本驱动的全身动作生成问题 | text-driven motion motion generation VQ-VAE |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | FSD: Fast Self-Supervised Single RGB-D to Categorical 3D Objects | 提出快速自监督方法以解决单RGB-D图像的3D物体识别问题 | 6D pose estimation |