cs.CV(2023-10-22)
📊 共 7 篇论文 | 🔗 1 篇有代码
🎯 兴趣领域导航
支柱三:空间感知与语义 (Perception & Semantics) (4 🔗1)
支柱六:视频提取与匹配 (Video Extraction) (1)
支柱九:具身大模型 (Embodied Foundation Models) (1)
支柱八:物理动画 (Physics-based Animation) (1)
🔬 支柱三:空间感知与语义 (Perception & Semantics) (4 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 1 | OV-VG: A Benchmark for Open-Vocabulary Visual Grounding | 提出OV-VG基准以解决开放词汇视觉定位问题 | open-vocabulary open vocabulary visual grounding | ✅ | |
| 2 | Mobile AR Depth Estimation: Challenges & Prospects -- Extended Version | 提出移动AR中的单目深度估计方法以解决深度感知挑战 | depth estimation monocular depth metric depth | ||
| 3 | A Quantitative Evaluation of Dense 3D Reconstruction of Sinus Anatomy from Monocular Endoscopic Video | 提出自监督方法以提高单目内窥镜视频的鼻窦三维重建精度 | depth estimation monocular depth 3D reconstruction | ||
| 4 | Guidance system for Visually Impaired Persons using Deep Learning and Optical flow | 提出深度学习与光流结合的导航系统以帮助视觉障碍人士 | depth estimation optical flow |
🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 5 | Research on Key Technologies of Infrastructure Digitalization based on Multimodal Spatial Data | 提出多模态空间数据基础设施数字化关键技术以解决交通网络构建问题 | feature matching multimodal |
🔬 支柱九:具身大模型 (Embodied Foundation Models) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 6 | MMTF-DES: A Fusion of Multimodal Transformer Models for Desire, Emotion, and Sentiment Analysis of Social Media Data | 提出MMTF-DES框架以解决社交媒体人类欲望分析问题 | multimodal |
🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 7 | ConViViT -- A Deep Neural Network Combining Convolutions and Factorized Self-Attention for Human Activity Recognition | 提出ConViViT以解决人类活动识别中的信息融合问题 | spatiotemporal |