cs.CV(2026-07-08)
📊 共 23 篇论文 | 🔗 6 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (9 🔗4)
支柱二:RL算法与架构 (RL & Architecture) (6)
支柱三:空间感知与语义 (Perception & Semantics) (5 🔗1)
支柱一:机器人控制 (Robot Control) (2 🔗1)
支柱七:动作重定向 (Motion Retargeting) (1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (9 篇)
🔬 支柱二:RL算法与架构 (RL & Architecture) (6 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 10 | BUS: Brain-Inspired Unsupervised Self-Reflection for Advanced Multimodal Reasoning | 提出BUS框架以解决多模态推理中的自反性不足问题 | reinforcement learning multimodal | ||
| 11 | Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation | 综述放射学中的视觉基础模型,提升临床应用潜力 | contrastive learning foundation model | ||
| 12 | CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training | 提出CarbonCLIP以解决城市碳排放预测问题 | contrastive learning distillation multimodal | ||
| 13 | A Theory of Contrastive Learning with Natural Images | 提出对比学习理论以优化自然图像表示 | contrastive learning | ||
| 14 | ColorFM: An Optimization-to-Learning Framework for Color Transfer via Flow Matching | 提出ColorFM框架以解决颜色传输中的不一致性问题 | flow matching | ||
| 15 | Infinite Worlds with Versatile Interactions | 提出LingBot-World 2.0以实现无限交互的虚拟世界 | world model world models |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction | 提出NoDrift3R以解决长序列重建中的位姿漂移问题 | 3D gaussian splatting 3DGS 3D reconstruction | ||
| 17 | Sparse Attention for Dense Open-Vocabulary Prediction in CLIP | 提出稀疏注意力机制以解决CLIP中的密集开放词汇预测问题 | open-vocabulary open vocabulary | ||
| 18 | Two-Stage Multi-Modal Fusion with Adaptive Alignment for Action Quality Assessment | 提出DualAlign以解决多模态融合中的对齐与稳定性问题 | optical flow | ||
| 19 | `Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation | 提出自监督框架以解决视频目标分割中的时空一致性问题 | optical flow | ||
| 20 | PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation | 提出PUF框架以解决在线3D场景图生成中的不确定性问题 | scene understanding | ✅ |
🔬 支柱一:机器人控制 (Robot Control) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 21 | Ego-Human Motion Prediction with 3D-Aware LLM | 提出Ego3DLM以解决人类动作预测中的3D语义理解问题 | motion tracking physically plausible egocentric | ✅ | |
| 22 | Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence | 提出LingBot-Video以解决视频生成模型在机器人控制中的领域不匹配问题 | manipulation egocentric foundation model |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 23 | SpiS-GAN: Spiral-Modulated Handwriting Synthesis with Star Operation | 提出SpiS-GAN以解决手写体合成中的结构细节缺失问题 | spatial relationship |