cs.CV(2023-10-24)
📊 共 16 篇论文 | 🔗 7 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (6 🔗3)
支柱三:空间感知与语义 (Perception & Semantics) (5 🔗3)
支柱二:RL算法与架构 (RL & Architecture) (2)
支柱一:机器人控制 (Robot Control) (1)
支柱七:动作重定向 (Motion Retargeting) (1 🔗1)
支柱八:物理动画 (Physics-based Animation) (1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (6 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | Woodpecker: Hallucination Correction for Multimodal Large Language Models | 提出Woodpecker以解决多模态大语言模型的幻觉问题 | large language model multimodal | ✅ | |
| 2 | Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs | 提出自动视觉裁剪方法以提升零-shot视觉问答性能 | large language model multimodal | ||
| 3 | CPSeg: Finer-grained Image Semantic Segmentation via Chain-of-Thought Language Prompting | 提出CPSeg框架以解决图像语义分割精细化问题 | chain-of-thought | ||
| 4 | Large Language Models are Temporal and Causal Reasoners for Video Question Answering | 提出Flipped-VQA框架以解决视频问答中的语言偏见问题 | large language model | ✅ | |
| 5 | TiC-CLIP: Continual Training of CLIP Models | 提出TiC-CLIP以解决CLIP模型持续更新问题 | foundation model | ✅ | |
| 6 | ShARc: Shape and Appearance Recognition for Person Identification In-the-wild | 提出ShARc以解决复杂环境下的人物识别问题 | multimodal |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 7 | Breaking of brightness consistency in optical flow with a lightweight CNN network | 提出轻量级CNN网络以解决光流中的亮度一致性问题 | optical flow | ✅ | |
| 8 | Cross-view Self-localization from Synthesized Scene-graphs | 提出混合场景模型以解决跨视角自定位问题 | NeRF neural radiance field | ||
| 9 | GNeSF: Generalizable Neural Semantic Fields | 提出GNeSF以解决3D场景分割中的泛化问题 | implicit representation semantic map | ✅ | |
| 10 | G2-MonoDepth: A General Framework of Generalized Depth Inference from Monocular RGB+X Data | 提出G2-MonoDepth以解决单目深度推断问题 | depth estimation monocular depth | ||
| 11 | Salient Object Detection in RGB-D Videos | 提出DCTNet+以解决RGB-D视频中的显著目标检测问题 | optical flow | ✅ |
🔬 支柱二:RL算法与架构 (RL & Architecture) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 12 | I$^2$MD: 3D Action Representation Learning with Inter- and Intra-modal Mutual Distillation | 提出I$^2$MD框架以解决3D动作表示学习中的模态互补性不足问题 | representation learning contrastive learning distillation | ||
| 13 | Privacy Protection in MRI Scans Using 3D Masked Autoencoders | 提出CP-MAE以解决MRI扫描隐私保护问题 | masked autoencoder MAE |
🔬 支柱一:机器人控制 (Robot Control) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 14 | What's Left? Concept Grounding with Logic-Enhanced Foundation Models | 提出逻辑增强基础模型LEFT以解决概念跨领域基础问题 | manipulation human motion large language model |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 15 | Language-driven Scene Synthesis using Multi-conditional Diffusion Model | 提出多条件扩散模型以解决语言驱动场景合成问题 | human motion | ✅ |
🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | Pix2HDR -- A pixel-wise acquisition and deep learning-based synthesis approach for high-speed HDR videos | 提出Pix2HDR以解决高速度HDR视频捕获问题 | spatiotemporal |