cs.CV(2026-07-14)

📊 共 34 篇论文 | 🔗 6 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (15 🔗2) 支柱三:空间感知与语义 (Perception & Semantics) (8 🔗2) 支柱二:RL算法与架构 (RL & Architecture) (8 🔗1) 支柱五:交互与反应 (Interaction & Reaction) (1 🔗1) 支柱七:动作重定向 (Motion Retargeting) (1) 支柱一:机器人控制 (Robot Control) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (15 篇)

#题目一句话要点标签🔗
1 EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval 提出EvoGraph-R1以解决静态知识图谱在多模态检索中的局限性 large language model multimodal
2 Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks 审计多模态基准中的数据泄露问题以提升WSI VQA评估准确性 foundation model multimodal
3 ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models 提出ViCo3D以解决LiDAR基础的协作3D目标检测问题 foundation model
4 VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression 提出VisCo以解决视觉标记压缩效率低的问题 large language model
5 Open-KNEAD: Knowledge-grounded Nutrition Estimation via Agentic Decomposition 提出Open-KNEAD框架以解决饮食营养估计问题 large language model multimodal
6 Hy-Embodied-VLM-1.0: Efficient Physical-World Agents 提出Hy-Embodied-VLM-1.0以提升物理世界代理的能力 foundation model multimodal
7 ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning 提出ReflectVLN以解决视觉语言导航中的闭环跟踪问题 VLN chain-of-thought
8 IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment 提出IQA-T1以解决开放世界图像质量评估问题 large language model multimodal
9 CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models 提出CoRe框架以解决跨图像比较推理问题 multimodal
10 HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and Ambivalence/Hesitancy Video Recognition 提出多任务学习方法以解决情感视频识别问题 multimodal
11 Text-Aided Multi-Modal Panoptic Symbol Spotting for CAD Floor Plan Drawings 提出TextCAD以解决CAD平面图符号识别中的文本注释利用不足问题 multimodal
12 Towards Vision-Free CIR: Attribute-Augmented Scoring and LLM-Based Reranking for Zero-Shot Composed Image Retrieval 提出视觉无关的CIR框架以解决复杂图像检索问题 multimodal
13 More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization 提出CROSS框架以解决极端视角下的跨视图定位问题 foundation model
14 UMSS: Towards Unsupervised Multi-modal Semantic Segmentation 提出UniM2以解决无监督多模态语义分割问题 multimodal
15 How to Realize Recursively Self-Improving Agents and Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture 提出多智能体架构以实现递归自我改进的代理与个人奇点 large language model

🔬 支柱三:空间感知与语义 (Perception & Semantics) (8 篇)

#题目一句话要点标签🔗
16 ExtraGS: Enhancing Endoscopic View Extrapolation via Diffusion-Guided 3D Gaussian Splatting 提出ExtraGS以解决内窥镜视图外推中的伪影问题 3D gaussian splatting gaussian splatting splatting
17 Implicit 4D Gaussian Splatting for Fast Motion with Large Inter-Frame Displacements 提出SPIN-4DGS以解决快速运动下的高质量重建问题 gaussian splatting splatting spatiotemporal
18 X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras 提出X-Lens以解决异构相机的实时度量深度估计问题 depth estimation metric depth
19 ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning 提出ARDepth以解决单目深度估计中的几何结构问题 depth estimation monocular depth
20 Semantic-Edge Response Decoding of SAM3 for Zero-Shot Crack Segmentation 提出语义边缘响应解码以解决零-shot裂缝分割问题 open-vocabulary open vocabulary foundation model
21 DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs 提出DynTrace以解决动态场景感知不足问题 scene understanding large language model multimodal
22 Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence 提出SIS-Bench以解决无人机自我意识与空间认知不足问题 optical flow large language model multimodal
23 DermDepth: Toward Monocular Metric Scale 3D Reconstruction Models for Dermatology 提出DermDepth以解决皮肤科单目三维重建问题 3D reconstruction

🔬 支柱二:RL算法与架构 (RL & Architecture) (8 篇)

#题目一句话要点标签🔗
24 RFMSR: Residual Flow Matching for Image Super-Resolution 提出RFMSR以解决图像超分辨率中的信息损失问题 flow matching foundation model
25 MambaPSA: A Mamba-based Replacement for C2PSA in YOLO26 提出MambaPSA以替代YOLO26中的C2PSA模块 Mamba SSM state space model
26 DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery 提出DM-KG以解决街景图像中视觉语言模型空间认知问题 MAE depth estimation metric depth
27 MobileSAM2: Lightweight Segment Anything for Spatial Intelligence 提出MobileSAM2以解决移动设备上的视频图像分割问题 distillation embodied AI foundation model
28 Adaptive Cross-Modal Fusion with Sparse Attention for Pedestrian Crossing Intention Prediction 提出ADAPT框架以解决行人过马路意图预测问题 Mamba semantic map multimodal
29 DiTailed: Ensuring Visual Object Consistency in Text-Image-to-Image Flow Matching Models 提出DiTailed以解决文本引导图像编辑中的对象一致性问题 flow matching
30 Contrastive-Augmented Flow Matching for Style-Content Disentanglement 提出对比增强流匹配以解决内容与风格分离问题 flow matching
31 The GEST-Engine: From Event Graphs to Synthetic Video. A Full Technical Report 提出GEST-Engine以实现从文本到合成视频的自动化生成 world model world models

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
32 MBTI: A Multi-Branch Efficient Fine-Tuning Framework for Hyperspectral Image Classification with Foundation Models 提出MBTI框架以解决高光谱图像分类中的模型迁移问题 HSI foundation model

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
33 Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation 提出Hallo4D以解决4D生成中的时空一致性问题 geometric consistency spatiotemporal multimodal

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
34 UniVR: Thinking in Visual Space for Unified Visual Reasoning 提出UniVR以解决视觉空间推理与规划问题 manipulation reinforcement learning multimodal

⬅️ 返回 cs.CV 首页 · 🏠 返回主页