cs.CV(2023-10-25)
📊 共 17 篇论文 | 🔗 6 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (6 🔗2)
支柱三:空间感知与语义 (Perception & Semantics) (6 🔗3)
支柱二:RL算法与架构 (RL & Architecture) (2 🔗1)
支柱一:机器人控制 (Robot Control) (1)
支柱七:动作重定向 (Motion Retargeting) (1)
支柱八:物理动画 (Physics-based Animation) (1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (6 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models | 提出DDCoT以解决多模态推理中的挑战 | large language model multimodal chain-of-thought | ||
| 2 | Diagnosing Alzheimer's Disease using Early-Late Multimodal Data Fusion with Jacobian Maps | 提出早晚多模态数据融合方法以诊断阿尔茨海默病 | multimodal | ||
| 3 | EdgeCalib: Multi-Frame Weighted Edge Features for Automatic Targetless LiDAR-Camera Calibration | 提出EdgeCalib以解决LiDAR与相机的自动标定问题 | multimodal | ||
| 4 | Exploring OCR Capabilities of GPT-4V(ision) : A Quantitative and In-depth Evaluation | 评估GPT-4V(ision)的OCR能力以解决多语言识别问题 | multimodal | ✅ | |
| 5 | ConvNets Match Vision Transformers at Scale | 提出高效ConvNet架构以匹敌Vision Transformer性能 | foundation model | ||
| 6 | Context Does Matter: End-to-end Panoptic Narrative Grounding with Deformable Attention Refined Matching Network | 提出可变注意力精细匹配网络以解决全景叙事定位问题 | visual grounding | ✅ |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 7 | Metrically Scaled Monocular Depth Estimation through Sparse Priors for Underwater Robots | 提出基于稀疏先验的单目深度估计方法以解决水下机器人深度预测问题 | depth estimation monocular depth | ✅ | |
| 8 | CoDet: Co-Occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection | 提出CoDet以解决开放词汇物体检测中的区域-词对齐问题 | open-vocabulary open vocabulary | ✅ | |
| 9 | Towards Explainability in Monocular Depth Estimation | 提出单目深度估计的可解释性研究以解决深度感知问题 | depth estimation monocular depth | ||
| 10 | PERF: Panoramic Neural Radiance Field from a Single Panorama | 提出PERF框架以解决单幅全景图生成3D场景的问题 | monocular depth NeRF neural radiance field | ✅ | |
| 11 | UAV-Sim: NeRF-based Synthetic Data Generation for UAV-based Perception | 提出UAV-Sim以解决无人机图像合成数据不足问题 | NeRF | ||
| 12 | LightSpeed: Light and Fast Neural Light Fields on Mobile Devices | 提出轻量快速的神经光场方法以解决移动设备实时图像合成问题 | NeRF |
🔬 支柱二:RL算法与架构 (RL & Architecture) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 13 | 4D-Editor: Interactive Object-level Editing in Dynamic Neural Radiance Fields via Semantic Distillation | 提出4D-Editor以解决动态场景中的对象级交互编辑问题 | distillation NeRF neural radiance field | ✅ | |
| 14 | GraFT: Gradual Fusion Transformer for Multimodal Re-Identification | 提出Gradual Fusion Transformer以解决多模态重识别问题 | representation learning multimodal |
🔬 支柱一:机器人控制 (Robot Control) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 15 | Open-NeRF: Towards Open Vocabulary NeRF Decomposition | 提出Open-NeRF以解决开放词汇NeRF分解问题 | manipulation 3D reconstruction NeRF |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | DSAM-GN:Graph Network based on Dynamic Similarity Adjacency Matrices for Vehicle Re-identification | 提出DSAM-GN以解决车辆重识别中的背景干扰问题 | spatial relationship |
🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 17 | ChimpACT: A Longitudinal Dataset for Understanding Chimpanzee Behaviors | 提出ChimpACT数据集以解决非人类灵长类行为研究不足问题 | spatiotemporal |