cs.CV(2023-10-23)

📊 共 23 篇论文 | 🔗 5 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (6 🔗1) 支柱二:RL算法与架构 (RL & Architecture) (5 🔗1) 支柱三:空间感知与语义 (Perception & Semantics) (4 🔗1) 支柱六:视频提取与匹配 (Video Extraction) (3) 支柱一:机器人控制 (Robot Control) (2) 支柱四:生成式动作 (Generative Motion) (1) 支柱五:交互与反应 (Interaction & Reaction) (1 🔗1) 支柱八:物理动画 (Physics-based Animation) (1 🔗1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (6 篇)

#题目一句话要点标签🔗
1 Large Language Models are Visual Reasoning Coordinators 提出Cola以协调多模态视觉推理模型 large language model multimodal
2 SyncFusion: Multimodal Onset-synchronized Video-to-Audio Foley Synthesis 提出SyncFusion以解决视频与音频同步问题 multimodal
3 Large Language Models can Share Images, Too! 提出DribeR框架以评估大型语言模型的图像共享能力 large language model
4 Videoprompter: an ensemble of foundational models for zero-shot video understanding 提出Videoprompter以解决零-shot视频理解问题 large language model
5 Vision-Enhanced Semantic Entity Recognition in Document Images via Visually-Asymmetric Consistency Learning 提出视觉不对称一致性学习以提升文档图像中的实体识别 multimodal
6 Leveraging Image-Text Similarity and Caption Modification for the DataComp Challenge: Filtering Track and BYOD Track 提出基于图像-文本相似性和标题修改的数据过滤方案 multimodal

🔬 支柱二:RL算法与架构 (RL & Architecture) (5 篇)

#题目一句话要点标签🔗
7 MAS: Multi-view Ancestral Sampling for 3D motion generation using 2D diffusion 提出多视角祖先采样方法以解决3D运动生成问题 distillation motion generation
8 SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding 提出SAM-CLIP以融合视觉基础模型解决语义与空间理解问题 distillation foundation model
9 DREAM+: Efficient Dataset Distillation by Bidirectional Representative Matching 提出DREAM+以解决数据集蒸馏中的样本选择问题 distillation
10 CalibrationPhys: Self-supervised Video-based Heart and Respiratory Rate Measurements by Calibrating Between Multiple Cameras 提出CalibrationPhys以解决无监督心率与呼吸率测量问题 contrastive learning PULSE
11 Remote Heart Rate Monitoring in Smart Environments from Videos with Self-supervised Pre-training 提出自监督对比学习以解决远程心率监测数据依赖问题 representation learning contrastive learning

🔬 支柱三:空间感知与语义 (Perception & Semantics) (4 篇)

#题目一句话要点标签🔗
12 CAwa-NeRF: Instant Learning of Compression-Aware NeRF Features 提出CAwa-NeRF以解决NeRF特征压缩问题 NeRF neural radiance field
13 VQ-NeRF: Vector Quantization Enhances Implicit Neural Representations 提出VQ-NeRF以解决隐式神经表示的计算复杂性问题 3D reconstruction NeRF
14 RoboDepth: Robust Out-of-Distribution Depth Estimation under Corruptions 提出RoboDepth以解决深度估计中的鲁棒性问题 depth estimation
15 Converting Depth Images and Point Clouds for Feature-based Pose Estimation 提出深度数据转换方法以提升姿态估计精度 visual odometry

🔬 支柱六:视频提取与匹配 (Video Extraction) (3 篇)

#题目一句话要点标签🔗
16 Localizing Active Objects from Egocentric Vision with Symbolic World Knowledge 提出一种新方法以从自我中心视觉中定位主动对象 egocentric egocentric vision Ego4D
17 3M-TRANSFORMER: A Multi-Stage Multi-Stream Multimodal Transformer for Embodied Turn-Taking Prediction 提出3M-TRANSFORMER以解决多方对话中的轮流发言预测问题 egocentric multimodal
18 On Unsupervised Partial Shape Correspondence 提出一种新方法以解决无监督部分形状对应问题 feature matching

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
19 Interaction-Driven Active 3D Reconstruction with Object Interiors 提出一种交互驱动的主动3D重建方法以解决物体内部结构获取问题 manipulation 3D reconstruction
20 Manipulation Mask Generator: High-Quality Image Manipulation Mask Generation Method Based on Modified Total Variation Noise Reduction 提出改进的全变差噪声减少方法以生成高质量图像篡改掩码 manipulation

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
21 Orientation-Aware Leg Movement Learning for Action-Driven Human Motion Prediction 提出基于方向感知的腿部运动学习以解决人类动作预测问题 motion diffusion model motion diffusion human motion

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
22 Open-Set Image Tagging with Multi-Grained Text Supervision 提出RAM++模型以解决开放集图像标记问题 human-object interaction large language model

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
23 Novel-View Acoustic Synthesis from 3D Reconstructed Rooms 提出结合3D重建与盲音频录音的创新视角声学合成方法 PULSE

⬅️ 返回 cs.CV 首页 · 🏠 返回主页