cs.CV(2023-10-31)
📊 共 16 篇论文 | 🔗 2 篇有代码
🎯 兴趣领域导航
支柱三:空间感知与语义 (Perception & Semantics) (5)
支柱九:具身大模型 (Embodied Foundation Models) (4)
支柱四:生成式动作 (Generative Motion) (2 🔗1)
支柱一:机器人控制 (Robot Control) (1)
支柱五:交互与反应 (Interaction & Reaction) (1)
支柱七:动作重定向 (Motion Retargeting) (1)
支柱八:物理动画 (Physics-based Animation) (1 🔗1)
支柱六:视频提取与匹配 (Video Extraction) (1)
🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | FLODCAST: Flow and Depth Forecasting via Multimodal Recurrent Architectures | 提出FLODCAST以解决光流和深度预测问题 | optical flow multimodal | ||
| 2 | FPO++: Efficient Encoding and Rendering of Dynamic Neural Radiance Fields by Analyzing and Enhancing Fourier PlenOctrees | 提出FPO++以解决动态神经辐射场渲染中的伪影问题 | NeRF neural radiance field | ||
| 3 | NeRF Revisited: Fixing Quadrature Instability in Volume Rendering | 提出新方法解决NeRF体积渲染中的四重积分不稳定性问题 | NeRF neural radiance field | ||
| 4 | Refined Equivalent Pinhole Model for Large-scale 3D Reconstruction from Spaceborne CCD Imagery | 提出等效针孔模型以解决大规模3D重建问题 | 3D reconstruction | ||
| 5 | Spuriosity Rankings for Free: A Simple Framework for Last Layer Retraining Based on Object Detection | 提出基于物体检测的排名框架以解决深度学习模型的虚假特征问题 | open-vocabulary open vocabulary |
🔬 支柱九:具身大模型 (Embodied Foundation Models) (4 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 6 | A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis | 评估GPT-4V在医学图像分析中的多模态能力 | large language model multimodal visual grounding | ||
| 7 | A Multi-Modal Foundation Model to Assist People with Blindness and Low Vision in Environmental Interaction | 提出多模态基础模型以帮助盲人和低视力人士进行环境交互 | foundation model | ||
| 8 | Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts | 提出知识增强提示以提升多模态语言模型在视觉问答中的表现 | multimodal | ||
| 9 | CapsFusion: Rethinking Image-Text Data at Scale | 提出CapsFusion以解决多模态数据质量与可扩展性问题 | large language model multimodal |
🔬 支柱四:生成式动作 (Generative Motion) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 10 | SemanticBoost: Elevating Motion Generation with Augmented Textual Cues | 提出SemanticBoost以解决复杂语义描述下的运动生成问题 | motion generation large language model | ||
| 11 | Pose-to-Motion: Cross-Domain Motion Retargeting with Pose Prior | 提出Pose-to-Motion以解决动作合成数据不足问题 | motion synthesis motion retargeting | ✅ |
🔬 支柱一:机器人控制 (Robot Control) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 12 | StairNet: Visual Recognition of Stairs for Human-Robot Locomotion | 提出StairNet以解决人机步态在楼梯环境中的识别问题 | locomotion egocentric egocentric vision |
🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 13 | Object-centric Video Representation for Long-term Action Anticipation | 提出对象中心视频表示以解决长期动作预测问题 | human-object interaction Ego4D |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 14 | GACE: Geometry Aware Confidence Enhancement for Black-Box 3D Object Detectors on LiDAR-Data | 提出GACE以解决LiDAR数据中3D目标检测信心估计不足问题 | spatial relationship |
🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 15 | ZoomNeXt: A Unified Collaborative Pyramid Network for Camouflaged Object Detection | 提出ZoomNeXt以解决伪装物体检测中的复杂性问题 | spatiotemporal | ✅ |
🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | Team I2R-VI-FF Technical Report on EPIC-KITCHENS VISOR Hand Object Segmentation Challenge 2023 | 提出结合PointRend与SAM以解决手部与物体分割问题 | egocentric |