cs.CV(2026-07-08)

📊 共 23 篇论文 | 🔗 6 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (9 🔗4) 支柱二:RL算法与架构 (RL & Architecture) (6) 支柱三:空间感知与语义 (Perception & Semantics) (5 🔗1) 支柱一:机器人控制 (Robot Control) (2 🔗1) 支柱七:动作重定向 (Motion Retargeting) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (9 篇)

#题目一句话要点标签🔗
1 MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models 提出MedPMC框架以解决医疗多模态数据不足问题 large language model foundation model multimodal
2 Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning 提出树状思维推理框架以解决文本到图像的上下文学习问题 large language model multimodal chain-of-thought
3 AT-Attn: Temporal-Aware Cross-Attention for Longitudinal Multimodal Alzheimer's Disease Diagnosis 提出AT-Attn以解决阿尔茨海默病诊断中的多模态信息融合问题 multimodal
4 General Incomplete Multimodal Learning via Dynamic Quality Perception 提出通用不完全多模态学习框架以解决缺失模态问题 multimodal
5 LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models 提出LoCA以解决视觉基础模型适应中的空间通道耦合问题 foundation model
6 HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models 提出HIVE以研究视觉语言模型中的后幻觉推理问题 multimodal
7 Prototype-Anchored Generalized Manifold Regression for Unknown-Domain Object Detection 提出原型锚定广义流形回归以解决未知领域物体检测问题 chain-of-thought
8 AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning 提出AnchorPrune以解决视觉模型中冗余视觉标记问题 multimodal
9 Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild 提出Video2Reaction以解决视频内容观众反应预测问题 multimodal

🔬 支柱二:RL算法与架构 (RL & Architecture) (6 篇)

#题目一句话要点标签🔗
10 BUS: Brain-Inspired Unsupervised Self-Reflection for Advanced Multimodal Reasoning 提出BUS框架以解决多模态推理中的自反性不足问题 reinforcement learning multimodal
11 Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation 综述放射学中的视觉基础模型,提升临床应用潜力 contrastive learning foundation model
12 CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training 提出CarbonCLIP以解决城市碳排放预测问题 contrastive learning distillation multimodal
13 A Theory of Contrastive Learning with Natural Images 提出对比学习理论以优化自然图像表示 contrastive learning
14 ColorFM: An Optimization-to-Learning Framework for Color Transfer via Flow Matching 提出ColorFM框架以解决颜色传输中的不一致性问题 flow matching
15 Infinite Worlds with Versatile Interactions 提出LingBot-World 2.0以实现无限交互的虚拟世界 world model world models

🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)

#题目一句话要点标签🔗
16 NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction 提出NoDrift3R以解决长序列重建中的位姿漂移问题 3D gaussian splatting 3DGS 3D reconstruction
17 Sparse Attention for Dense Open-Vocabulary Prediction in CLIP 提出稀疏注意力机制以解决CLIP中的密集开放词汇预测问题 open-vocabulary open vocabulary
18 Two-Stage Multi-Modal Fusion with Adaptive Alignment for Action Quality Assessment 提出DualAlign以解决多模态融合中的对齐与稳定性问题 optical flow
19 `Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation 提出自监督框架以解决视频目标分割中的时空一致性问题 optical flow
20 PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation 提出PUF框架以解决在线3D场景图生成中的不确定性问题 scene understanding

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
21 Ego-Human Motion Prediction with 3D-Aware LLM 提出Ego3DLM以解决人类动作预测中的3D语义理解问题 motion tracking physically plausible egocentric
22 Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence 提出LingBot-Video以解决视频生成模型在机器人控制中的领域不匹配问题 manipulation egocentric foundation model

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
23 SpiS-GAN: Spiral-Modulated Handwriting Synthesis with Star Operation 提出SpiS-GAN以解决手写体合成中的结构细节缺失问题 spatial relationship

⬅️ 返回 cs.CV 首页 · 🏠 返回主页