cs.CV(2023-10-23)
📊 共 23 篇论文 | 🔗 5 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (6 🔗1)
支柱二:RL算法与架构 (RL & Architecture) (5 🔗1)
支柱三:空间感知与语义 (Perception & Semantics) (4 🔗1)
支柱六:视频提取与匹配 (Video Extraction) (3)
支柱一:机器人控制 (Robot Control) (2)
支柱四:生成式动作 (Generative Motion) (1)
支柱五:交互与反应 (Interaction & Reaction) (1 🔗1)
支柱八:物理动画 (Physics-based Animation) (1 🔗1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (6 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | Large Language Models are Visual Reasoning Coordinators | 提出Cola以协调多模态视觉推理模型 | large language model multimodal | ||
| 2 | SyncFusion: Multimodal Onset-synchronized Video-to-Audio Foley Synthesis | 提出SyncFusion以解决视频与音频同步问题 | multimodal | ||
| 3 | Large Language Models can Share Images, Too! | 提出DribeR框架以评估大型语言模型的图像共享能力 | large language model | ✅ | |
| 4 | Videoprompter: an ensemble of foundational models for zero-shot video understanding | 提出Videoprompter以解决零-shot视频理解问题 | large language model | ||
| 5 | Vision-Enhanced Semantic Entity Recognition in Document Images via Visually-Asymmetric Consistency Learning | 提出视觉不对称一致性学习以提升文档图像中的实体识别 | multimodal | ||
| 6 | Leveraging Image-Text Similarity and Caption Modification for the DataComp Challenge: Filtering Track and BYOD Track | 提出基于图像-文本相似性和标题修改的数据过滤方案 | multimodal |
🔬 支柱二:RL算法与架构 (RL & Architecture) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 7 | MAS: Multi-view Ancestral Sampling for 3D motion generation using 2D diffusion | 提出多视角祖先采样方法以解决3D运动生成问题 | distillation motion generation | ✅ | |
| 8 | SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding | 提出SAM-CLIP以融合视觉基础模型解决语义与空间理解问题 | distillation foundation model | ||
| 9 | DREAM+: Efficient Dataset Distillation by Bidirectional Representative Matching | 提出DREAM+以解决数据集蒸馏中的样本选择问题 | distillation | ||
| 10 | CalibrationPhys: Self-supervised Video-based Heart and Respiratory Rate Measurements by Calibrating Between Multiple Cameras | 提出CalibrationPhys以解决无监督心率与呼吸率测量问题 | contrastive learning PULSE | ||
| 11 | Remote Heart Rate Monitoring in Smart Environments from Videos with Self-supervised Pre-training | 提出自监督对比学习以解决远程心率监测数据依赖问题 | representation learning contrastive learning |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (4 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 12 | CAwa-NeRF: Instant Learning of Compression-Aware NeRF Features | 提出CAwa-NeRF以解决NeRF特征压缩问题 | NeRF neural radiance field | ||
| 13 | VQ-NeRF: Vector Quantization Enhances Implicit Neural Representations | 提出VQ-NeRF以解决隐式神经表示的计算复杂性问题 | 3D reconstruction NeRF | ||
| 14 | RoboDepth: Robust Out-of-Distribution Depth Estimation under Corruptions | 提出RoboDepth以解决深度估计中的鲁棒性问题 | depth estimation | ||
| 15 | Converting Depth Images and Point Clouds for Feature-based Pose Estimation | 提出深度数据转换方法以提升姿态估计精度 | visual odometry | ✅ |
🔬 支柱六:视频提取与匹配 (Video Extraction) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | Localizing Active Objects from Egocentric Vision with Symbolic World Knowledge | 提出一种新方法以从自我中心视觉中定位主动对象 | egocentric egocentric vision Ego4D | ||
| 17 | 3M-TRANSFORMER: A Multi-Stage Multi-Stream Multimodal Transformer for Embodied Turn-Taking Prediction | 提出3M-TRANSFORMER以解决多方对话中的轮流发言预测问题 | egocentric multimodal | ||
| 18 | On Unsupervised Partial Shape Correspondence | 提出一种新方法以解决无监督部分形状对应问题 | feature matching |
🔬 支柱一:机器人控制 (Robot Control) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 19 | Interaction-Driven Active 3D Reconstruction with Object Interiors | 提出一种交互驱动的主动3D重建方法以解决物体内部结构获取问题 | manipulation 3D reconstruction | ||
| 20 | Manipulation Mask Generator: High-Quality Image Manipulation Mask Generation Method Based on Modified Total Variation Noise Reduction | 提出改进的全变差噪声减少方法以生成高质量图像篡改掩码 | manipulation |
🔬 支柱四:生成式动作 (Generative Motion) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 21 | Orientation-Aware Leg Movement Learning for Action-Driven Human Motion Prediction | 提出基于方向感知的腿部运动学习以解决人类动作预测问题 | motion diffusion model motion diffusion human motion |
🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 22 | Open-Set Image Tagging with Multi-Grained Text Supervision | 提出RAM++模型以解决开放集图像标记问题 | human-object interaction large language model | ✅ |
🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 23 | Novel-View Acoustic Synthesis from 3D Reconstructed Rooms | 提出结合3D重建与盲音频录音的创新视角声学合成方法 | PULSE | ✅ |