cs.CV(2023-10-16)

📊 共 28 篇论文 | 🔗 4 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (12 🔗1) 支柱三:空间感知与语义 (Perception & Semantics) (8 🔗2) 支柱二:RL算法与架构 (RL & Architecture) (4 🔗1) 支柱六:视频提取与匹配 (Video Extraction) (2) 支柱七:动作重定向 (Motion Retargeting) (1) 支柱一:机器人控制 (Robot Control) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (12 篇)

#题目一句话要点标签🔗
1 VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools 提出VidCoM以解决视频理解与用户指令响应问题 large language model multimodal
2 LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation 提出LLM4SGG以解决弱监督场景图生成中的语义简化与低密度问题 large language model chain-of-thought
3 Few-shot Action Recognition with Captioning Foundation Models 提出CapFSAR框架以解决少样本动作识别问题 foundation model multimodal
4 BiomedJourney: Counterfactual Biomedical Image Generation by Instruction-Learning from Multimodal Patient Journeys 提出BiomedJourney以解决生物医学图像生成中的反事实问题 multimodal
5 Interpreting and Controlling Vision Foundation Models via Text Explanations 提出一种框架以通过文本解释理解和控制视觉基础模型 foundation model
6 Multimodal Object Query Initialization for 3D Object Detection 提出EfficientQ3M以解决3D目标检测中的查询初始化问题 multimodal
7 Using Global Land Cover Product as Prompt for Cropland Mapping via Visual Foundation Model 提出基于全球土地覆盖产品的提示学习方法以解决农田映射问题 foundation model
8 Automated Natural Language Explanation of Deep Visual Neurons with Large Models 提出自动化框架以生成深度视觉神经元的语义解释 foundation model
9 Loci-Segmented: Improving Scene Segmentation Learning 提出Loci-Segmented以解决场景分割学习中的背景依赖问题 foundation model
10 A Multi-Scale Spatial Transformer U-Net for Simultaneously Automatic Reorientation and Segmentation of 3D Nuclear Cardiac Images 提出多尺度空间变换U-Net以解决3D核心脏图像的自动重定向与分割问题 multimodal
11 Black-box Targeted Adversarial Attack on Segment Anything (SAM) 提出黑箱针对性对抗攻击方法以评估SAM模型的鲁棒性 foundation model
12 Towards Unified and Effective Domain Generalization 提出UniDG框架以提升领域泛化性能 foundation model

🔬 支柱三:空间感知与语义 (Perception & Semantics) (8 篇)

#题目一句话要点标签🔗
13 TraM-NeRF: Tracing Mirror and Near-Perfect Specular Reflections through Neural Radiance Fields 提出TraM-NeRF以解决镜面反射渲染问题 NeRF neural radiance field implicit representation
14 Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting 提出4D高斯点云以解决动态场景重建与渲染问题 gaussian splatting splatting
15 Self-supervised Fetal MRI 3D Reconstruction Based on Radiation Diffusion Generation Model 提出辐射扩散生成模型以解决胎儿MRI高质量重建问题 3D reconstruction NeRF
16 AutoDIR: Automatic All-in-One Image Restoration with Latent Diffusion 提出AutoDIR以解决图像恢复中的多种未知退化问题 open-vocabulary open vocabulary foundation model
17 Multi-Body Neural Scene Flow 提出多体神经场流以解决场景流中的刚体运动识别问题 scene flow
18 Long-term Dependency for 3D Reconstruction of Freehand Ultrasound Without External Tracker 提出长时依赖参数化方法以解决无外部跟踪的3D超声重建问题 3D reconstruction
19 Effortless Cross-Platform Video Codec: A Codebook-Based Method 提出基于码本的方法以解决跨平台视频编码问题 optical flow
20 Flow Dynamics Correction for Action Recognition 提出光流动态校正方法以提升动作识别性能 optical flow

🔬 支柱二:RL算法与架构 (RL & Architecture) (4 篇)

#题目一句话要点标签🔗
21 MoConVQ: Unified Physics-Based Motion Control via Scalable Discrete Representations 提出MoConVQ框架以实现统一的物理基础运动控制 reinforcement learning motion generation VQ-VAE
22 Temporal Embeddings: Scalable Self-Supervised Temporal Representation Learning from Spatiotemporal Data for Multimodal Computer Vision 提出自监督时序嵌入以解决多模态计算机视觉中的时序表示学习问题 representation learning spatiotemporal multimodal
23 DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing 提出DynVideo-E以解决长视频人类中心编辑问题 distillation NeRF neural radiance field
24 AST: Effective Dataset Distillation through Alignment with Smooth and High-Quality Expert Trajectories 提出AST框架以通过专家轨迹对数据蒸馏进行有效对齐 distillation

🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)

#题目一句话要点标签🔗
25 A Novel Benchmarking Paradigm and a Scale- and Motion-Aware Model for Egocentric Pedestrian Trajectory Prediction 提出新基准与运动感知模型以解决行人轨迹预测问题 egocentric multimodal
26 EAR-Net: Pursuing End-to-End Absolute Rotations from Multi-View Images 提出EAR-Net以解决多视角图像的绝对旋转估计问题 feature matching

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
27 Expression Domain Translation Network for Cross-domain Head Reenactment 提出表达域转换网络以解决跨域头部重演问题 geometric consistency human motion

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
28 Video Language Planning 提出视频语言规划以解决复杂长时间任务的规划问题 manipulation dexterous manipulation multimodal

⬅️ 返回 cs.CV 首页 · 🏠 返回主页