cs.CV(2026-07-22)

📊 共 26 篇论文 | 🔗 5 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (10 🔗2) 支柱二:RL算法与架构 (RL & Architecture) (6 🔗1) 支柱三:空间感知与语义 (Perception & Semantics) (5) 支柱八:物理动画 (Physics-based Animation) (2 🔗1) 支柱五:交互与反应 (Interaction & Reaction) (1) 支柱一:机器人控制 (Robot Control) (1 🔗1) 支柱七:动作重定向 (Motion Retargeting) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (10 篇)

#题目一句话要点标签🔗
1 MV-Bench: Benchmarking Multimodal Large Language Models for Coordinated Multi-View Interface Construction 提出MV-Bench以解决多视图接口构建评估问题 large language model multimodal
2 Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs 提出SmartVL以解决多模态大语言模型的高推理成本问题 large language model multimodal
3 ETPDesigner: Multi-Agent Orchestration for Interactive Multimodal Electronic Theater Program 提出ETPDesigner以解决复杂电子剧场节目设计问题 large language model multimodal
4 Self-supervision drives representational convergence in medical foundation models more than clinical supervision 提出自监督学习以提升医学基础模型的表示收敛性 foundation model
5 Development of an automated, reliable, and clinically meaningful artificial intelligence (AI) tool for diagnosing cardiac disease from conventional cardiovascular magnetic resonance (CMR) images 开发自动化AI工具以诊断心脏疾病 large language model foundation model multimodal
6 Forecasting the Number of Harvest-ready Fruits of Sweet Peppers Using Multimodal Time-Series Data 提出多模态深度学习框架以提高甜椒果实产量预测精度 multimodal
7 MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation 提出MTVDiff以解决热成像到可见光人脸转换中的挑战 multimodal
8 Diverse-Intent Multi-Turn Fashion Image Retrieval 提出DIM-Fashion以解决多轮时尚图像检索中的意图多样性问题 multimodal
9 ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models 提出ENTRAP-VL以研究视觉语言模型中的上下文引导问题 multimodal
10 Lean-SAM2: Target-Anchored Memory and Encoder Acceleration for SAM2 提出Lean-SAM2以解决SAM2模型的内存和效率问题 TAMP

🔬 支柱二:RL算法与架构 (RL & Architecture) (6 篇)

#题目一句话要点标签🔗
11 STEREOFLOW: Progressive Stereo Matching with StereoDiT and Transition Flow Matching 提出StereoFlow以解决立体匹配中的模态分布问题 flow matching 3D reconstruction scene flow
12 Factor-Informed Uncertainty Distillation for Gaze Estimation 提出因子引导的不确定性蒸馏以提升注视估计精度 curriculum learning teacher-student distillation
13 Physics-Aware Complex-Valued State Space Model with Scattering-Prior Feature Modulation for PolSAR Image Classification 提出CV-SSMNet以解决PolSAR图像分类中的物理知识不足问题 SSM state space model representation learning
14 PercepCap: Video Captioner with Structured Spatio-Temporal Perception 提出PercepCap以解决视频字幕生成中的时空感知问题 reinforcement learning spatiotemporal TAMP
15 PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving 提出PerceptDrive以解决自主驾驶中的感知优先世界-动作建模问题 flow matching distillation foundation model
16 Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training 提出重要性感知采样以解决可见-红外对齐问题 representation learning contrastive learning curriculum learning

🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)

#题目一句话要点标签🔗
17 Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose? 系统评估多模态大语言模型在遥感图像理解中的应用 scene understanding large language model multimodal
18 Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos 提出Memory-Augmented MLLMs以解决无人机小目标理解问题 open-vocabulary open vocabulary large language model
19 ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion 提出ATSplat以解决3D高斯点云稀疏性问题 3D gaussian splatting 3DGS gaussian splatting
20 Look Before You Edit: Attention-Guided Camera Placement and Multi-View Alignment for 3D Gaussian Splatting Editing 提出LB-Edit以解决3D场景编辑中的相机放置与视图一致性问题 3D gaussian splatting 3DGS gaussian splatting
21 Extending a Large View Synthesis Model for Multi-view Panoptic Segmentation 提出一种新方法以扩展大视图合成模型用于多视角全景分割 3D reconstruction scene understanding

🔬 支柱八:物理动画 (Physics-based Animation) (2 篇)

#题目一句话要点标签🔗
22 ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment 提出ReFace以改善面部疼痛评估问题 spatiotemporal
23 Efficient Tracking and Understanding Object Transformations 提出FluxGraph以解决对象变换跟踪的高计算成本问题 spatiotemporal

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
24 StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation 提出StreamHOI以解决实时人机交互视频生成问题 HOI

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
25 Evolving Cache Schedules for Fast Diffusion Policy Inference 提出EVO以解决扩散策略推理中的计算效率问题 manipulation diffusion policy

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
26 RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs 提出RIM框架以解决无人机跨域视觉定位问题 geometric consistency foundation model

⬅️ 返回 cs.CV 首页 · 🏠 返回主页