cs.CV(2023-10-03)

📊 共 16 篇论文 | 🔗 3 篇有代码

🎯 兴趣领域导航

支柱三:空间感知与语义 (Perception & Semantics) (7 🔗1) 支柱九:具身大模型 (Embodied Foundation Models) (5 🔗1) 支柱四:生成式动作 (Generative Motion) (2 🔗1) 支柱二:RL算法与架构 (RL & Architecture) (2)

🔬 支柱三:空间感知与语义 (Perception & Semantics) (7 篇)

#题目一句话要点标签🔗
1 Adaptive Multi-NeRF: Exploit Efficient Parallelism in Adaptive Multiple Scale Neural Radiance Field Rendering 提出自适应多NeRF以解决神经渲染效率问题 NeRF neural radiance field
2 MIMO-NeRF: Fast Neural Rendering with Multi-input Multi-output Neural Radiance Fields 提出MIMO-NeRF以解决NeRF渲染速度慢的问题 NeRF neural radiance field
3 EvDNeRF: Reconstructing Event Data with Dynamic Neural Radiance Fields 提出EvDNeRF以重建动态场景中的事件数据 NeRF neural radiance field TAMP
4 Skin the sheep not only once: Reusing Various Depth Datasets to Drive the Learning of Optical Flow 提出一种新方法以重用深度数据集提升光流估计精度 depth estimation monocular depth optical flow
5 Robust deformable image registration using cycle-consistent implicit representations 提出循环一致性隐式表示以解决医学图像配准的鲁棒性问题 implicit representation
6 CLIP Is Also a Good Teacher: A New Learning Framework for Inductive Zero-shot Semantic Segmentation 提出CLIP-ZSS框架以解决零-shot语义分割问题 open-vocabulary open vocabulary
7 Selective Feature Adapter for Dense Vision Transformers 提出选择性特征适配器以解决视觉变换器参数冗余问题 depth estimation

🔬 支柱九:具身大模型 (Embodied Foundation Models) (5 篇)

#题目一句话要点标签🔗
8 MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts 提出MathVista基准以评估视觉上下文中的数学推理能力 large language model foundation model multimodal
9 Multi-Prompt Fine-Tuning of Foundation Models for Enhanced Medical Image Segmentation 提出多提示微调框架以提升医学图像分割性能 foundation model
10 Sieve: Multimodal Dataset Pruning Using Image Captioning Models 提出Sieve以解决多模态数据集修剪问题 multimodal
11 Improved Automatic Diabetic Retinopathy Severity Classification Using Deep Multimodal Fusion of UWF-CFP and OCTA Images 提出多模态融合方法以提升糖尿病视网膜病变分类准确性 multimodal
12 HallE-Control: Controlling Object Hallucination in Large Multimodal Models 提出HallE-Control以解决多模态模型中的对象幻觉问题 multimodal

🔬 支柱四:生成式动作 (Generative Motion) (2 篇)

#题目一句话要点标签🔗
13 Hierarchical Generation of Human-Object Interactions with Diffusion Probabilistic Models 提出层次生成框架以解决人机交互长程运动合成问题 motion generation human-object interaction
14 MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens 提出MiniGPT-5以解决多模态生成中的图文一致性问题 classifier-free guidance large language model multimodal

🔬 支柱二:RL算法与架构 (RL & Architecture) (2 篇)

#题目一句话要点标签🔗
15 Understanding Masked Autoencoders From a Local Contrastive Perspective 提出局部对比视角的LC-MAE以解析Masked AutoEncoder的有效性 masked autoencoder MAE contrastive learning
16 ScaleNet: An Unsupervised Representation Learning Method for Limited Information 提出ScaleNet以解决有限信息下的无监督表示学习问题 representation learning

⬅️ 返回 cs.CV 首页 · 🏠 返回主页