cs.CV(2023-10-26)
📊 共 17 篇论文 | 🔗 7 篇有代码
🎯 兴趣领域导航
支柱二:RL算法与架构 (RL & Architecture) (5 🔗1)
支柱三:空间感知与语义 (Perception & Semantics) (5 🔗3)
支柱九:具身大模型 (Embodied Foundation Models) (2 🔗1)
支柱六:视频提取与匹配 (Video Extraction) (2 🔗2)
支柱七:动作重定向 (Motion Retargeting) (1)
支柱四:生成式动作 (Generative Motion) (1)
支柱一:机器人控制 (Robot Control) (1)
🔬 支柱二:RL算法与架构 (RL & Architecture) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | Three Pillars improving Vision Foundation Model Distillation for Lidar | 提出三大支柱以改善激光雷达的视觉基础模型蒸馏 | distillation foundation model | ||
| 2 | HyperFields: Towards Zero-Shot Generation of NeRFs from Text | 提出HyperFields以实现文本条件下的NeRF零-shot生成 | distillation NeRF neural radiance field | ||
| 3 | Noise-Free Score Distillation | 提出无噪声评分蒸馏以优化文本到图像生成 | distillation classifier-free guidance | ||
| 4 | Understanding the Effects of Projectors in Knowledge Distillation | 提出投影器集成方法以提升知识蒸馏性能 | teacher-student distillation | ||
| 5 | Prototypical Contrastive Learning-based CLIP Fine-tuning for Object Re-identification | 提出基于原型对比学习的CLIP微调方法以提升物体重识别性能 | contrastive learning | ✅ |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 6 | LP-OVOD: Open-Vocabulary Object Detection by Linear Probing | 提出LP-OVOD以解决开放词汇物体检测问题 | open-vocabulary open vocabulary | ✅ | |
| 7 | Detection Defenses: An Empty Promise against Adversarial Patch Attacks on Optical Flow | 研究检测防御机制对光流预测的影响 | optical flow motion prediction | ✅ | |
| 8 | Masked Space-Time Hash Encoding for Efficient Dynamic Scene Reconstruction | 提出Masked Space-Time Hash编码以解决动态场景重建问题 | scene reconstruction | ✅ | |
| 9 | Learning depth from monocular video sequences | 提出新训练损失以解决单目视频序列深度估计问题 | depth estimation monocular depth | ||
| 10 | RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open Environments | 提出RIO数据集以解决开放环境中的意图导向物体检测问题 | affordance |
🔬 支柱九:具身大模型 (Embodied Foundation Models) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 11 | Task-driven Prompt Evolution for Foundation Models | 提出SAMPOT以优化基础模型的图像分割任务 | foundation model | ||
| 12 | ControlLLM: Augment Language Models with Tools by Searching on Graphs | 提出ControlLLM框架以解决多模态工具调用问题 | large language model | ✅ |
🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 13 | Learning Temporal Sentence Grounding From Narrated EgoVideos | 提出Clip Merging方法以解决长视频中的时间句子定位问题 | egocentric Ego4D TAMP | ✅ | |
| 14 | IndustReal: A Dataset for Procedure Step Recognition Handling Execution Errors in Egocentric Videos in an Industrial-Like Setting | 提出IndustReal数据集以解决工业场景中的程序步骤识别问题 | egocentric | ✅ |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 15 | Affective Video Content Analysis: Decade Review and New Perspectives | 综述情感视频内容分析的进展与未来研究方向 | motion representation multimodal |
🔬 支柱四:生成式动作 (Generative Motion) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | CADS: Unleashing the Diversity of Diffusion Models through Condition-Annealed Sampling | 提出CADS以解决扩散模型输出多样性不足问题 | classifier-free guidance |
🔬 支柱一:机器人控制 (Robot Control) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 17 | Lookup Table meets Local Laplacian Filter: Pyramid Reconstruction Network for Tone Mapping | 提出基于局部拉普拉斯滤波的3D查找表以解决HDR图像色调映射问题 | manipulation |