cs.CV(2023-10-17)

📊 共 11 篇论文 | 🔗 5 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (5 🔗2) 支柱三:空间感知与语义 (Perception & Semantics) (3 🔗1) 支柱二:RL算法与架构 (RL & Architecture) (3 🔗2)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (5 篇)

#题目一句话要点标签🔗
1 Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V 提出Set-of-Mark方法以提升GPT-4V的视觉定位能力 multimodal visual grounding
2 Towards Training-free Open-world Segmentation via Image Prompt Foundation Models 提出无训练开放世界分割方法以解决灵活性不足问题 large language model foundation model
3 Towards Automatic Satellite Images Captions Generation Using Large Language Models 提出自动遥感图像字幕生成方法以解决数据集不足问题 large language model
4 EvalCrafter: Benchmarking and Evaluating Large Video Generation Models 提出EvalCrafter以解决视频生成模型评估不足问题 large language model
5 NICE: Improving Panoptic Narrative Detection and Segmentation with Cascading Collaborative Learning 提出NICE框架以解决全景叙事检测与分割问题 visual grounding

🔬 支柱三:空间感知与语义 (Perception & Semantics) (3 篇)

#题目一句话要点标签🔗
6 Self-Supervised 3D Scene Flow Estimation and Motion Prediction using Local Rigidity Prior 提出自监督学习方法以解决3D场景流估计与运动预测问题 scene flow motion estimation motion prediction
7 FocDepthFormer: Transformer with latent LSTM for Depth Estimation from Focal Stack 提出FocDepthFormer以解决深度估计中焦点堆栈的局限性问题 depth estimation
8 Learning Neural Implicit through Volume Rendering with Attentive Depth Fusion Priors 提出注意力深度融合先验以解决神经隐式表示的深度缺失问题 3D reconstruction implicit representation

🔬 支柱二:RL算法与架构 (RL & Architecture) (3 篇)

#题目一句话要点标签🔗
9 Unsupervised Pre-Training Using Masked Autoencoders for ECG Analysis 提出无监督预训练方法以提升心电图分析精度 masked autoencoder MAE
10 MonoSKD: General Distillation Framework for Monocular 3D Object Detection via Spearman Correlation Coefficient 提出MonoSKD框架以解决单目3D目标检测中的知识蒸馏问题 distillation
11 Enhancing Plasticity for First Session Adaptation Continual Learning 提出PLASTIC以解决异构任务分布下的持续学习问题 teacher-student distillation

⬅️ 返回 cs.CV 首页 · 🏠 返回主页