cs.CV(2026-07-15)

📊 共 40 篇论文 | 🔗 11 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (14 🔗6) 支柱二:RL算法与架构 (RL & Architecture) (11 🔗2) 支柱三:空间感知与语义 (Perception & Semantics) (4) 支柱七:动作重定向 (Motion Retargeting) (4 🔗1) 支柱一:机器人控制 (Robot Control) (3 🔗1) 支柱六:视频提取与匹配 (Video Extraction) (3 🔗1) 支柱五:交互与反应 (Interaction & Reaction) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (14 篇)

#题目一句话要点标签🔗
1 Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models 提出切片式数据合成方法以提升医学多模态大语言模型的3D空间推理能力 large language model multimodal chain-of-thought
2 FM$^2$: Unified Federated Foundation Models for Heterogeneous Multimodal Medical Imaging 提出FM$^2$框架以解决异构多模态医学影像的联合学习问题 foundation model multimodal
3 A novel unsupervised machine learning strategy to handle multimodal cardiac PET/MRI data 提出一种新颖的无监督机器学习策略处理多模态心脏PET/MRI数据 multimodal
4 Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning 提出多模态深度学习框架以评估胰腺癌切除可能性 multimodal
5 VGIF-Score: Interpretable and Diagnostic Evaluation of Spatio-Temporal Instruction Following in Video Generation 提出VGIF-Score以解决视频生成模型指令遵循评估不足的问题 instruction following
6 TRACE-PCa: Predicting Prostate Cancer Progression from Longitudinal MRI During Active Surveillance 提出TRACE-PCa以解决前列腺癌监测中的病理进展预测问题 foundation model multimodal
7 WAVE-Stereo: Warp-Aligned Volume Encoding for Stereo Matching 提出WAVE-Stereo以解决立体匹配中的特征对齐问题 foundation model
8 CF-Net: Conflict Fusion with Speaker Normalisation and Certainty Weighting for Ambivalence/Hesitancy Recognition 提出CF-Net以解决视频中的模棱两可和犹豫识别问题 multimodal
9 Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge 提出多模态说话人识别方法以解决语言和模态限制问题 multimodal
10 UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets 提出UniPhysGen以解决3D资产物理语义统一问题 embodied AI
11 Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models 提出无注意力轻量级令牌减少方法以提升视觉语言模型效率 multimodal
12 LPM: Industrial-Scale Generative Video Restoration 提出LPM以解决工业规模视频恢复问题 foundation model
13 ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding 提出ScanFocus以解决时空视频定位中的边界精确性问题 multimodal
14 Audio-Text Cross-Attention with Psycholinguistic Support Features for Ambivalence/Hesitancy Recognition 提出音频-文本跨注意力模型以识别矛盾与犹豫情感 TAMP

🔬 支柱二:RL算法与架构 (RL & Architecture) (11 篇)

#题目一句话要点标签🔗
15 Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs 提出Groc-PO以解决多模态大语言模型的真实度问题 DPO direct preference optimization large language model
16 SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning 提出SIVA-RL以解决多模态强化学习中的视觉对齐问题 reinforcement learning multimodal
17 From Surface Forecasting to Observability Forecasting: A Latent World Model for Cloud-Aware EO Monitoring 提出云监测的可观测性预测模型以提升地球观测效率 world model worldmodel world models
18 VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders 提出VideoRAE以解决视频生成模型的表示学习问题 JEPA foundation model
19 From Pixels to States: Rethinking Interactive World Models as Game Engines 提出互动世界模型作为游戏引擎以解决游戏交互性问题 world model world models
20 Towards Spatial Supersensing in the Wild 提出VSI-Super-Wild以解决真实场景中的空间超感知问题 world model world models multimodal
21 OvisOCR2 Technical Report 提出OvisOCR2以解决文档解析问题 reinforcement learning distillation reward design
22 ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning 提出ThinkBLOX以解决3D室内场景生成中的交互编辑问题 reinforcement learning chain-of-thought
23 Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection 提出基于共生启发的知识蒸馏方法以解决增量目标检测问题 distillation
24 The 2nd International StepUP Competition for Biometric Footstep Recognition: From Steps to Strides 提出基于压力的脚步生物识别方法以解决用户识别挑战 representation learning spatiotemporal
25 Beyond Color Geometry: Evaluating Human-Like Color Representations in Vision Models 提出模糊感知模型以评估视觉模型中的人类色彩表现 masked autoencoder MAE

🔬 支柱三:空间感知与语义 (Perception & Semantics) (4 篇)

#题目一句话要点标签🔗
26 Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting 提出一种解耦辐射表示以加速3D高斯点云渲染 3D gaussian splatting 3DGS gaussian splatting
27 Calibrated Closed-Form Uncertainty for Radiative Gaussian Splatting in Sparse-View CT 提出闭式形式的不确定性校准以改善稀视角CT重建 gaussian splatting splatting
28 DreamSat-Pose: Spacecraft Pose Estimation from Single-View 3D Reconstructions and Learned 2D-3D Feature Matching 提出DreamSat-Pose以解决未知航天器姿态估计问题 3D reconstruction feature matching
29 CASA-SDF: Curriculum-Aware Spatial Adaptation with Curvature-Guided Density for Neural Implicit Surface Reconstruction 提出CASA-SDF以解决室内场景表面重建中的几何异质性问题 3D reconstruction implicit representation

🔬 支柱七:动作重定向 (Motion Retargeting) (4 篇)

#题目一句话要点标签🔗
30 GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding 提出GeoAnchor以解决3D空间理解中的多模态推理问题 spatial relationship large language model multimodal
31 MultiAnimate: A Unified Framework for Controllable Multi-Character Animation 提出MultiAnimate框架以解决多角色动画生成问题 spatial relationship character animation
32 Music-to-Dance Generation via Atomic Movements 提出基于原子动作的音乐驱动舞蹈生成框架以解决结构不一致问题 human motion large language model
33 RainDancer: RGB-Event Video Deraining with Rain-Oriented Spiking Dynamics 提出RainDancer以解决动态雨天视频去雨问题 structure preservation

🔬 支柱一:机器人控制 (Robot Control) (3 篇)

#题目一句话要点标签🔗
34 Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation 提出REAL框架以解决开放世界移动操控中的探索与交互问题 manipulation mobile manipulation dual-arm
35 M$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming 提出M$^4$World以解决自主驾驶模拟中的对象控制与长时间稳定性问题 manipulation world model world models
36 Detector Confidence Signals Presence Rather Than Occlusion in Cluttered Manipulation 提出新方法以解决检测器在遮挡场景中的信号误读问题 manipulation open-vocabulary open vocabulary

🔬 支柱六:视频提取与匹配 (Video Extraction) (3 篇)

#题目一句话要点标签🔗
37 T3HG-Editor: Text-driven 3D Human Garment Editing with Body Priors Embedded in SMPL-X 提出T3HG-Editor以解决文本驱动的3D人类服装编辑问题 SMPL SMPL-X
38 Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction 提出Human4K数据集以解决3D人类重建中的深度模糊问题 SMPL SMPL-X motion retargeting
39 EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent 提出EgoProceVQA以解决日常活动程序理解问题 egocentric

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
40 Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild 提出AgentHOI以解决训练无关的人-物交互检测问题 human-object interaction HOI large language model

⬅️ 返回 cs.CV 首页 · 🏠 返回主页