cs.CV(2026-07-27)

📊 共 38 篇论文 | 🔗 4 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (16) 支柱二:RL算法与架构 (RL & Architecture) (7 🔗1) 支柱三:空间感知与语义 (Perception & Semantics) (5) 支柱一:机器人控制 (Robot Control) (4) 支柱六:视频提取与匹配 (Video Extraction) (2) 支柱七:动作重定向 (Motion Retargeting) (2 🔗2) 支柱八:物理动画 (Physics-based Animation) (1) 支柱四:生成式动作 (Generative Motion) (1 🔗1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (16 篇)

#题目一句话要点标签🔗
1 ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding 提出ClinFusion以解决医疗领域多模态理解问题 large language model multimodal instruction following
2 Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding 提出Mixture-of-Thought-Tokens以解决多模态定位与推理统一问题 large language model multimodal
3 IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data 通过合成训练数据适应基础模型以提升人脸识别性能 foundation model
4 Disentangling Semantic Attention from Structural Bias in the Attention Manifold 提出SPAR方法以解决多模态语言模型中的视觉注意力偏差问题 large language model multimodal visual grounding
5 MMOE: Modernizing Diffusion Transformers with Efficient Expert Design 提出MMOE以平衡生成质量与训练成本问题 large language model foundation model
6 The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding 提出稀疏帧适应策略以解决视频定位中的性能瓶颈问题 large language model multimodal
7 ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation 提出ReflexTrack以解决训练无关的引用视频目标分割问题 large language model multimodal
8 Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking 提出时空条件去噪变换器以解决模态缺失的RGBT跟踪问题 multimodal
9 Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels 提出无坐标和区域标签的证据归属方法以解决文档理解问题 multimodal
10 CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding 提出CADER以解决长视频理解中的动态证据推理问题 chain-of-thought
11 DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes 提出DecoupleMix以解决VLM数据构建的系统性问题 multimodal
12 Rethinking Expert Training for Model Merging with Prompt Learning 提出双调专家以优化模型合并中的专家训练 foundation model
13 UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing 提出UMI3D以解决多图像输入下3D生成质量下降问题 foundation model
14 FilmBench: A Film-Grade Benchmark for Cinematic Video Generation 提出FilmBench以解决视频生成评估标准不足的问题 multimodal
15 DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection 提出DuoAD以解决训练无关的少样本异常检测问题 foundation model
16 Embeddings based Anomaly Detection for Cleaning Global Crop Type Reference Datasets 提出基于嵌入的异常检测以清理全球作物类型参考数据集 foundation model

🔬 支柱二:RL算法与架构 (RL & Architecture) (7 篇)

#题目一句话要点标签🔗
17 RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models 提出RP-OPSD以解决多模态大语言模型的自蒸馏问题 distillation privileged information large language model
18 MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning 提出多尺度自适应视觉编码器以提升细粒度视觉感知与多模态推理效率 distillation spatial relationship large language model
19 Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI 提出多模态AI框架以提升眼底摄影分析的临床应用 SSM state space model foundation model
20 Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation 提出正向方向匹配以解决负支路不对称问题 distillation privileged information classifier-free guidance
21 Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding 提出链式思维框架以解决交通规则理解问题 reinforcement learning chain-of-thought
22 Test-Time Adaptation via Dual Distillation for Videos Under Severe Distribution Shifts 提出双蒸馏的测试时适应方法以应对视频中的严重分布偏移问题 distillation
23 LCMamNet: A Lightweight Cross-scale Mamba Network for Infrared Small Target Detection 提出LCMamNet以解决红外小目标检测中的背景干扰问题 Mamba

🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)

#题目一句话要点标签🔗
24 GenSplatCodec: Feed-Forward Gaussian Splatting Compression via One-Step Diffusion 提出GenSplatCodec以解决低比特率高频纹理压缩问题 3D gaussian splatting 3DGS gaussian splatting
25 Multimodal Semantic-Probabilistic Objectness for Open World Object Detection 提出MSPO框架以提升开放世界物体检测的准确性 open-vocabulary open vocabulary multimodal
26 SILICA: Repurposing Diffusion Priors for Joint Glass Segmentation and Depth Estimation 提出SILICA以解决透明表面深度估计与分割问题 depth estimation monocular depth zero-shot transfer
27 MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction 提出MSVS-VAE以解决高保真3D重建中的重构质量瓶颈问题 3D reconstruction
28 NSL-SLAM: High-Fidelity Neural Structured-Light Depth for Practical SLAM and Reconstruction 提出NSL-SLAM以解决高保真结构光深度感知问题 monocular depth

🔬 支柱一:机器人控制 (Robot Control) (4 篇)

#题目一句话要点标签🔗
29 Face Age Verification Vulnerabilities Under Simple Appearance Manipulations 提出面部年龄验证系统的脆弱性研究以应对外观操控问题 manipulation large language model multimodal
30 CameraAnything: Refilming Videos with Arbitrary Camera Control 提出CameraAnything以解决视频编辑中的相机控制问题 manipulation 3D reconstruction
31 What Can I Edit? Open-Ended Strategy Discovery and the Emotion Editability Landscape 提出EmoScope以解决情感图像编辑的策略发现问题 manipulation affordance
32 DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models 提出DailyBench以解决AI生成与操控图像检测的评估问题 manipulation

🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)

#题目一句话要点标签🔗
33 EgoPlay: Event-Triggered Video Editing for Egocentric Streams 提出EgoPlay以解决自我中心视频编辑中的事件触发问题 egocentric Ego4D
34 Multiview Multi-Person Human Mesh Recovery Under Large Scenes with Occlusions 提出MVMP-HMR以解决大场景下多人重建与遮挡问题 human mesh recovery HMR

🔬 支柱七:动作重定向 (Motion Retargeting) (2 篇)

#题目一句话要点标签🔗
35 Accuracy potential of visual localization exploiting high-end street-level imagery 提出高精度视觉定位方法以解决GNSS局限性问题 motion reconstruction
36 DreamStyle3D: Efficient 3D Stylized Asset Generation via Dual-Attention Disentanglement 提出DreamStyle3D以解决高效生成3D风格化资产问题 geometric consistency

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
37 Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention 提出Gazette框架以解决人类注意力解码问题 spatiotemporal large language model multimodal

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
38 ViDS: Video Diffusion Shader using 3D Face Tracking 提出ViDS以解决生动且身份保留的肖像动画问题 motion latent

⬅️ 返回 cs.CV 首页 · 🏠 返回主页