cs.CV(2023-10-24)

📊 共 16 篇论文 | 🔗 7 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (6 🔗3) 支柱三:空间感知与语义 (Perception & Semantics) (5 🔗3) 支柱二:RL算法与架构 (RL & Architecture) (2) 支柱一:机器人控制 (Robot Control) (1) 支柱七:动作重定向 (Motion Retargeting) (1 🔗1) 支柱八:物理动画 (Physics-based Animation) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (6 篇)

#题目一句话要点标签🔗
1 Woodpecker: Hallucination Correction for Multimodal Large Language Models 提出Woodpecker以解决多模态大语言模型的幻觉问题 large language model multimodal
2 Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs 提出自动视觉裁剪方法以提升零-shot视觉问答性能 large language model multimodal
3 CPSeg: Finer-grained Image Semantic Segmentation via Chain-of-Thought Language Prompting 提出CPSeg框架以解决图像语义分割精细化问题 chain-of-thought
4 Large Language Models are Temporal and Causal Reasoners for Video Question Answering 提出Flipped-VQA框架以解决视频问答中的语言偏见问题 large language model
5 TiC-CLIP: Continual Training of CLIP Models 提出TiC-CLIP以解决CLIP模型持续更新问题 foundation model
6 ShARc: Shape and Appearance Recognition for Person Identification In-the-wild 提出ShARc以解决复杂环境下的人物识别问题 multimodal

🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)

#题目一句话要点标签🔗
7 Breaking of brightness consistency in optical flow with a lightweight CNN network 提出轻量级CNN网络以解决光流中的亮度一致性问题 optical flow
8 Cross-view Self-localization from Synthesized Scene-graphs 提出混合场景模型以解决跨视角自定位问题 NeRF neural radiance field
9 GNeSF: Generalizable Neural Semantic Fields 提出GNeSF以解决3D场景分割中的泛化问题 implicit representation semantic map
10 G2-MonoDepth: A General Framework of Generalized Depth Inference from Monocular RGB+X Data 提出G2-MonoDepth以解决单目深度推断问题 depth estimation monocular depth
11 Salient Object Detection in RGB-D Videos 提出DCTNet+以解决RGB-D视频中的显著目标检测问题 optical flow

🔬 支柱二:RL算法与架构 (RL & Architecture) (2 篇)

#题目一句话要点标签🔗
12 I$^2$MD: 3D Action Representation Learning with Inter- and Intra-modal Mutual Distillation 提出I$^2$MD框架以解决3D动作表示学习中的模态互补性不足问题 representation learning contrastive learning distillation
13 Privacy Protection in MRI Scans Using 3D Masked Autoencoders 提出CP-MAE以解决MRI扫描隐私保护问题 masked autoencoder MAE

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
14 What's Left? Concept Grounding with Logic-Enhanced Foundation Models 提出逻辑增强基础模型LEFT以解决概念跨领域基础问题 manipulation human motion large language model

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
15 Language-driven Scene Synthesis using Multi-conditional Diffusion Model 提出多条件扩散模型以解决语言驱动场景合成问题 human motion

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
16 Pix2HDR -- A pixel-wise acquisition and deep learning-based synthesis approach for high-speed HDR videos 提出Pix2HDR以解决高速度HDR视频捕获问题 spatiotemporal

⬅️ 返回 cs.CV 首页 · 🏠 返回主页