cs.CV(2026-07-09)

📊 共 37 篇论文 | 🔗 9 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (15 🔗4) 支柱三:空间感知与语义 (Perception & Semantics) (8 🔗2) 支柱二:RL算法与架构 (RL & Architecture) (7 🔗2) 支柱一:机器人控制 (Robot Control) (3 🔗1) 支柱六:视频提取与匹配 (Video Extraction) (2) 支柱四:生成式动作 (Generative Motion) (2)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (15 篇)

#题目一句话要点标签🔗
1 LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action 提出LEEVLA以解决复杂动态场景中的视觉-语言-动作问题 vision-language-action VLA multimodal
2 CT-CLIP Representations for Multimodal Lung Cancer Survival Prediction 提出CT-CLIP表示以解决肺癌生存预测问题 foundation model multimodal
3 UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery 提出UniRef-UAV以解决无人机图像中的多模态指代问题 multimodal visual grounding
4 Predicting Viticulture Potential through an Ensemble of U-Net and a Geospatial Foundation Model 通过U-Net与地理基础模型集成预测葡萄种植潜力 foundation model
5 DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models 提出DeltaV以解决多模态模型中视觉状态冗余问题 multimodal
6 Multimodal 3D LUT Generation via StatLUT with Statistical Features for Photorealistic Style Transfer 提出StatLUT以解决光写实风格迁移中的色彩与结构问题 multimodal
7 VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness 提出VSRo-200数据集以研究罗马尼亚视觉语音识别的监督与多模态鲁棒性 multimodal
8 Post-Training in End-to-End Autonomous Driving 提出后训练技术以提升自动驾驶模型的可靠性 vision-language-action multimodal
9 LUMI: Tokenizer-Agnostic LLM-Based Lossless Image Compression 提出LUMI框架以解决图像无损压缩中的tokenizer依赖问题 large language model foundation model
10 OpenCoF: Learning to Reason Through Video Generation 提出OpenCoF框架以解决视频生成中的推理能力不足问题 chain-of-thought
11 When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities 提出结构稀疏自编码器以解决视觉语言模型中的概念一致性问题 multimodal
12 Dive Into the Implicit Biases of Low-rank Vision-language Alignment 提出低秩适应方法以优化视觉-语言对齐 large language model
13 Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark 提出双关联超图网络以解决RGB-T视频目标检测中的对齐问题 multimodal
14 Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing 提出混合探测器以解决多模态大语言模型训练中的特权模态问题 large language model multimodal
15 Is sub-metre resolution necessary for cocoa mapping? A landscape-stratified evaluation of very high resolution imagery, decametric Earth Observation inputs, and operational products in Cote d'Ivoire 通过高分辨率影像提升可可种植区映射精度 foundation model

🔬 支柱三:空间感知与语义 (Perception & Semantics) (8 篇)

#题目一句话要点标签🔗
16 On the Design of Mixture-of-Experts for Dynamic Gaussian Splatting 提出混合专家模型以解决动态高斯点云重建问题 3D gaussian splatting gaussian splatting splatting
17 UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation 提出UAV-OVVIS以解决无人机视频实例分割的开放词汇问题 open-vocabulary open vocabulary foundation model
18 VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval 提出VocaDet以解决开放词汇物体检测与分割问题 open-vocabulary open vocabulary feature matching
19 Geometry and Gradient-based Partitioning for Panoramic Outdoor Reconstruction 提出PanoLOG以解决大规模户外场景重建问题 monocular depth 3D gaussian splatting 3DGS
20 Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation 提出AR-ERIS框架以解决内窥镜图像的开放词汇指称分割问题 open-vocabulary open vocabulary
21 Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery 提出Track2Map以解决手术视频中缺乏准确相机轨迹的问题 3D gaussian splatting 3D reconstruction gaussian splatting
22 LTM: Large-scale Terrain Model for Wildfire-prone Landscapes 提出多模态重建框架以解决野火高危地区的3D地形建模问题 3D reconstruction feature matching
23 StereoSplat+: Feed-Forward Stereo Gaussian Splatting with Diffusion-Assisted Progressive Inference 提出StereoSplat+以解决单视角立体重建问题 3D gaussian splatting 3DGS gaussian splatting

🔬 支柱二:RL算法与架构 (RL & Architecture) (7 篇)

#题目一句话要点标签🔗
24 WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving 提出WCog-VLA以解决现有自动驾驶模型的认知不足问题 world model world models physically plausible
25 ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device 提出ZipDepth以解决轻量级单目深度估计问题 distillation depth estimation monocular depth
26 Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning 提出Switch-Reasoner以解决多任务混合中的推理效率问题 reinforcement learning large language model multimodal
27 Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing 提出认知结构多模态智能体以解决长时段多模态对话问题 reinforcement learning multimodal
28 ProsMAE: Multi-Source MAE Pretraining for ISUP Grade Classification 提出ProsMAE以解决组织切片图像分类中的多源数据挑战 representation learning masked autoencoder MAE
29 OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators 提出OPSD-V以解决长视频生成中的动态衰退问题 distillation
30 Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions 提出持续测试时适应方法以应对计算机视觉中的数据分布变化问题 teacher-student foundation model

🔬 支柱一:机器人控制 (Robot Control) (3 篇)

#题目一句话要点标签🔗
31 Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues? 提出新学习范式以解决手-物交互识别中的短路问题 manipulation HOI egocentric
32 Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio 提出时间比率以解决视频动作泛化差距问题 manipulation foundation model
33 Decoupled Illumination Priors for Spatially Controllable Multi-View Indoor Scene Relighting 提出Lume-Palette以解决室内场景重照明的多视角一致性问题 manipulation distillation

🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)

#题目一句话要点标签🔗
34 Whareformer: Learning to Track What is Where in Long Egocentric Videos 提出Whareformer以解决长时间自我中心视频中的物体追踪问题 egocentric
35 VEGAS: Human-Aligned Video Caption Evaluation via Gaze 提出VEGAS以解决视频字幕与观众注意力不匹配问题 egocentric

🔬 支柱四:生成式动作 (Generative Motion) (2 篇)

#题目一句话要点标签🔗
36 SkelGen4D: Weakly-Supervised Skeleton-Based 4D Generation for Text-Driven Mesh Animation 提出SkelGen4D以解决文本驱动的4D网格动画生成问题 physically plausible
37 Closing the Null Space: Guidance-Aware Quantization for Classifier-Free Diffusion 提出GAMP以解决CFG扩散模型量化问题 classifier-free guidance

⬅️ 返回 cs.CV 首页 · 🏠 返回主页