cs.CV(2023-10-12)

📊 共 21 篇论文 | 🔗 6 篇有代码

🎯 兴趣领域导航

支柱三:空间感知与语义 (Perception & Semantics) (6 🔗3) 支柱二:RL算法与架构 (RL & Architecture) (5) 支柱一:机器人控制 (Robot Control) (3 🔗1) 支柱九:具身大模型 (Embodied Foundation Models) (2 🔗1) 支柱八:物理动画 (Physics-based Animation) (2) 支柱四:生成式动作 (Generative Motion) (1) 支柱六:视频提取与匹配 (Video Extraction) (1 🔗1) 支柱七:动作重定向 (Motion Retargeting) (1)

🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)

#题目一句话要点标签🔗
1 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering 提出4D高斯点云渲染方法以解决动态场景实时渲染问题 gaussian splatting splatting TAMP
2 EC-Depth: Exploring the consistency of self-supervised monocular depth estimation in challenging scenes 提出EC-Depth以解决自监督单目深度估计在复杂场景中的一致性问题 depth estimation monocular depth
3 PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm 提出PonderV2以解决3D基础模型预训练的挑战 3D reconstruction foundation model
4 RT-SRTS: Angle-Agnostic Real-Time Simultaneous 3D Reconstruction and Tumor Segmentation from Single X-Ray Projection 提出RT-SRTS以解决实时肿瘤分割与三维重建问题 3D reconstruction
5 Consistent123: Improve Consistency for One Image to 3D Object Synthesis 提出Consistent123以解决3D对象合成中的视图一致性问题 3D reconstruction classifier-free guidance
6 Implicit Shape and Appearance Priors for Few-Shot Full Head Reconstruction 提出隐式形状与外观先验以解决少样本全头重建问题 3D reconstruction

🔬 支柱二:RL算法与架构 (RL & Architecture) (5 篇)

#题目一句话要点标签🔗
7 Multimodal Large Language Model for Visual Navigation 提出多模态大语言模型以解决视觉导航问题 behavior cloning large language model multimodal
8 GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models 提出GaussianDreamer以解决3D资产生成的速度与质量问题 dreamer 3D gaussian splatting gaussian splatting
9 Multimodal Variational Auto-encoder based Audio-Visual Segmentation 提出显式条件多模态变分自编码器解决音视频分割问题 representation learning multimodal
10 Aligning Data Selection with Performance: Performance-driven Reinforcement Learning for Active Learning in Object Detection 提出MGRAL以解决目标检测中的主动学习数据选择问题 reinforcement learning
11 Hyp-UML: Hyperbolic Image Retrieval with Uncertainty-aware Metric Learning 提出Hyp-UML以解决图像检索中的不确定性问题 representation learning contrastive learning

🔬 支柱一:机器人控制 (Robot Control) (3 篇)

#题目一句话要点标签🔗
12 Octopus: Embodied Vision-Language Programmer from Environmental Feedback 提出Octopus以解决高层规划与实际操作之间的鸿沟问题 manipulation reinforcement learning embodied AI
13 DUSA: Decoupled Unsupervised Sim2Real Adaptation for Vehicle-to-Everything Collaborative Perception 提出DUSA以解决V2X协作感知中的Sim2Real适应问题 sim2real
14 Visual Data-Type Understanding does not emerge from Scaling Vision-Language Models 提出视觉数据类型识别以解决视觉语言模型的盲点问题 manipulation

🔬 支柱九:具身大模型 (Embodied Foundation Models) (2 篇)

#题目一句话要点标签🔗
15 Generalized Logit Adjustment: Calibrating Fine-tuned Models by Removing Label Bias in Foundation Models 提出广义逻辑调整方法以消除基础模型中的标签偏差 foundation model zero-shot transfer
16 STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment 提出STELLA以解决音视频持续学习中的时空关联问题 multimodal

🔬 支柱八:物理动画 (Physics-based Animation) (2 篇)

#题目一句话要点标签🔗
17 A Deep Learning Framework for Spatiotemporal Ultrasound Localization Microscopy 提出深度学习框架以解决超声定位显微镜中的微血管重建问题 spatiotemporal
18 Im4D: High-Fidelity and Real-Time Novel View Synthesis for Dynamic Scenes 提出Im4D以解决动态场景中的视图合成问题 spatiotemporal

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
19 OmniControl: Control Any Joint at Any Time for Human Motion Generation 提出OmniControl以解决人类动作生成中的关节控制问题 motion generation OmniControl human motion

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
20 Mapping Memes to Words for Multimodal Hateful Meme Classification 提出ISSUES以解决多模态仇恨表情包分类问题 HuMoR multimodal

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
21 DiscoMatch: Fast Discrete Optimisation for Geometrically Consistent 3D Shape Matching 提出DiscoMatch以解决3D形状匹配中的几何一致性问题 geometric consistency

⬅️ 返回 cs.CV 首页 · 🏠 返回主页