cs.CV(2023-10-01)
📊 共 14 篇论文 | 🔗 4 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (6 🔗3)
支柱三:空间感知与语义 (Perception & Semantics) (3)
支柱一:机器人控制 (Robot Control) (2)
支柱二:RL算法与架构 (RL & Architecture) (1 🔗1)
支柱五:交互与反应 (Interaction & Reaction) (1)
支柱六:视频提取与匹配 (Video Extraction) (1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (6 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants | 提出Muffin框架和UniMM-Chat数据集以提升多模态助手性能 | large language model foundation model multimodal | ✅ | |
| 2 | HOH: Markerless Multimodal Human-Object-Human Handover Dataset with Large Object Count | 提出HOH数据集以解决无标记人机交互中的物体交接问题 | multimodal | ||
| 3 | RegBN: Batch Normalization of Multimodal Data with Regularization | 提出RegBN以解决多模态数据归一化问题 | multimodal | ✅ | |
| 4 | LiveChat: Video Comment Generation from Audio-Visual Multimodal Contexts | 提出LiveChat以解决视频直播评论生成问题 | multimodal | ||
| 5 | Propagating Semantic Labels in Video Data | 提出视频数据语义标签传播方法以减少手动标注工作量 | foundation model | ||
| 6 | Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs | 提出新框架以提升多模态大语言模型的细粒度图像理解能力 | large language model | ✅ |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 7 | Multi-tiling Neural Radiance Field (NeRF) -- Geometric Assessment on Large-scale Aerial Datasets | 提出多切片神经辐射场以解决大规模航空数据集的几何评估问题 | 3D reconstruction NeRF neural radiance field | ||
| 8 | Logical Bias Learning for Object Relation Prediction | 提出基于因果推理的逻辑偏见学习以改善对象关系预测 | scene understanding foundation model | ||
| 9 | Win-Win: Training High-Resolution Vision Transformers from Two Windows | 提出Win-Win方法以高效训练高分辨率视觉变换器 | monocular depth optical flow |
🔬 支柱一:机器人控制 (Robot Control) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 10 | A Hierarchical Graph-based Approach for Recognition and Description Generation of Bimanual Actions in Videos | 提出层次图模型以提升双手动作视频的识别与描述生成 | manipulation bi-manual | ||
| 11 | GhostEncoder: Stealthy Backdoor Attacks with Dynamic Triggers to Pre-trained Encoders in Self-supervised Learning | 提出GhostEncoder以解决自监督学习中的隐形后门攻击问题 | manipulation |
🔬 支柱二:RL算法与架构 (RL & Architecture) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 12 | Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning | 提出多模态ICL变体以解决大型多模态模型的局限性 | RLHF generalist agent large language model | ✅ |
🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 13 | Scene-aware Human Motion Forecasting via Mutual Distance Prediction | 提出基于互距预测的场景感知人类动作预测方法 | human-scene interaction human motion |
🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 14 | Comics for Everyone: Generating Accessible Text Descriptions for Comic Strips | 提出可访问的漫画文本描述生成方法以解决视觉障碍者的阅读问题 | HuMoR large language model multimodal |