cs.CV(2023-10-04)
📊 共 17 篇论文 | 🔗 1 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (6 🔗1)
支柱三:空间感知与语义 (Perception & Semantics) (5)
支柱二:RL算法与架构 (RL & Architecture) (3)
支柱一:机器人控制 (Robot Control) (3)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (6 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | Comprehensive Multimodal Segmentation in Medical Imaging: Combining YOLOv8 with SAM and HQ-SAM Models | 提出YOLOv8与SAM结合的多模态医学影像分割方法 | multimodal | ||
| 2 | Improving Automatic VQA Evaluation Using Large Language Models | 利用大型语言模型改进自动视觉问答评估 | large language model | ||
| 3 | ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models | 提出ECoFLaP以解决大规模视觉语言模型的高能耗问题 | multimodal | ||
| 4 | Improving Vision Anomaly Detection with the Guidance of Language Modality | 提出跨模态引导方法以解决视觉异常检测中的信息冗余问题 | multimodal | ||
| 5 | GET: Group Event Transformer for Event-Based Vision | 提出Group Event Transformer以解决事件摄像头信息提取不足问题 | TAMP | ✅ | |
| 6 | ShaSTA-Fuse: Camera-LiDAR Sensor Fusion to Model Shape and Spatio-Temporal Affinities for 3D Multi-Object Tracking | 提出ShaSTA-Fuse以解决3D多目标跟踪中的传感器融合问题 | multimodal |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 7 | USB-NeRF: Unrolling Shutter Bundle Adjusted Neural Radiance Fields | 提出USB-NeRF以解决滚动快门图像合成问题 | NeRF neural radiance field motion estimation | ||
| 8 | CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection | 提出CoDA以解决开放词汇3D物体检测中的新物体定位与分类问题 | open-vocabulary open vocabulary | ||
| 9 | Shielding the Unseen: Privacy Protection through Poisoning NeRF with Spatial Deformation | 提出隐私保护方法以应对NeRF模型的生成能力 | NeRF neural radiance field scene reconstruction | ||
| 10 | Efficient-3DiM: Learning a Generalizable Single-image Novel-view Synthesizer in One Day | 提出Efficient-3DiM以解决单图像新视角合成问题 | neural radiance field | ||
| 11 | T$^3$Bench: Benchmarking Current Progress in Text-to-3D Generation | 提出T$^3$Bench以解决文本到3D生成评估问题 | NeRF |
🔬 支柱二:RL算法与架构 (RL & Architecture) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 12 | Kosmos-G: Generating Images in Context with Multimodal Large Language Models | 提出Kosmos-G以解决多模态图像生成中的输入限制问题 | distillation large language model multimodal | ||
| 13 | Reinforcement Learning-based Mixture of Vision Transformers for Video Violence Recognition | 提出基于强化学习的混合视觉变换器以解决视频暴力识别问题 | reinforcement learning | ||
| 14 | SweetDreamer: Aligning Geometric Priors in 2D Diffusion for Consistent Text-to-3D | 提出SweetDreamer以解决2D到3D生成中的几何不一致问题 | dreamer |
🔬 支柱一:机器人控制 (Robot Control) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 15 | Continuous 3D Myocardial Motion Tracking via Echocardiography | 提出神经心脏运动场以解决心肌运动追踪不准确问题 | motion tracking ReMoS motion estimation | ||
| 16 | ED-NeRF: Efficient Text-Guided Editing of 3D Scene with Latent Space NeRF | 提出ED-NeRF以解决3D场景编辑速度慢的问题 | manipulation distillation NeRF | ||
| 17 | A Grammatical Compositional Model for Video Action Detection | 提出语法组合模型以解决视频动作检测中的复杂动态问题 | manipulation |