cs.CV(2023-10-07)
📊 共 11 篇论文 | 🔗 2 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (5 🔗1)
支柱三:空间感知与语义 (Perception & Semantics) (2)
支柱二:RL算法与架构 (RL & Architecture) (2 🔗1)
支柱六:视频提取与匹配 (Video Extraction) (1)
支柱一:机器人控制 (Robot Control) (1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | HowToCaption: Prompting LLMs to Transform Video Annotations at Scale | 提出利用大型语言模型生成高质量视频字幕以解决视频标注不足问题 | large language model multimodal TAMP | ||
| 2 | Tree-GPT: Modular Large Language Model Expert System for Forest Remote Sensing Image Understanding and Interactive Analysis | 提出Tree-GPT以解决森林遥感图像理解问题 | large language model | ||
| 3 | VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models | 提出VLATTACK以解决多模态对抗攻击问题 | multimodal | ✅ | |
| 4 | Towards Long-Range 3D Object Detection for Autonomous Vehicles | 提出双网络结合与虚拟点增强以解决长距离3D目标检测问题 | multimodal | ||
| 5 | DISCOVER: Making Vision Networks Interpretable via Competition and Dissection | 提出DISCOVER框架以提高视觉网络的可解释性 | multimodal |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 6 | UFD-PRiME: Unsupervised Joint Learning of Optical Flow and Stereo Depth through Pixel-Level Rigid Motion Estimation | 提出UFD-PRiME以解决光流与立体深度联合学习问题 | stereo depth optical flow motion estimation | ||
| 7 | Federated Self-Supervised Learning of Monocular Depth Estimators for Autonomous Vehicles | 提出FedSCDepth以解决单目深度估计的隐私与效率问题 | depth estimation monocular depth |
🔬 支柱二:RL算法与架构 (RL & Architecture) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 8 | Reinforced UI Instruction Grounding: Towards a Generic UI Task Automation API | 提出一种多模态模型以解决UI任务自动化问题 | reinforcement learning large language model multimodal | ||
| 9 | HalluciDet: Hallucinating RGB Modality for Person Detection Through Privileged Information | 提出HalluciDet以解决红外到可见光行人检测问题 | privileged information | ✅ |
🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 10 | 1st Place Solution of Egocentric 3D Hand Pose Estimation Challenge 2023 Technical Report:A Concise Pipeline for Egocentric Hand Pose Reconstruction | 提出基于ViT的回归方法以解决单视角3D手势估计问题 | egocentric |
🔬 支柱一:机器人控制 (Robot Control) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 11 | Exploiting Facial Relationships and Feature Aggregation for Multi-Face Forgery Detection | 提出多面孔伪造检测框架以解决现有方法的局限性 | manipulation |