cs.CV(2026-07-24)

📊 共 19 篇论文 | 🔗 9 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (7 🔗4) 支柱二:RL算法与架构 (RL & Architecture) (5 🔗2) 支柱三:空间感知与语义 (Perception & Semantics) (5 🔗1) 支柱一:机器人控制 (Robot Control) (1 🔗1) 支柱八:物理动画 (Physics-based Animation) (1 🔗1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (7 篇)

#题目一句话要点标签🔗
1 Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification in the Foundation-Model Era 重新思考多分支与跨骨干融合以提升车辆重识别性能 foundation model
2 Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models 提出Medical-Checklist以评估多模态医学模型的理解能力 multimodal
3 Rethinking Layer-Wise Information Allocation for Vision Foundation Model Adaptation 提出Prompted Information Bottlenecks以解决视觉基础模型适应性问题 foundation model
4 Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs 提出能力差距驱动的后训练方法以解决遥感多模态大语言模型的场景专用性问题 large language model multimodal
5 CommandLM: Data driven behavior level descriptor for ego vehicles 提出CommandLM以解决自主驾驶行为描述透明性问题 large language model multimodal
6 FAIR: Feature-Augmented Implicit Regularization for AI-generated Fake Image Detection 提出FAIR以解决AI生成假图像检测中的泛化问题 zero-shot transfer
7 EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection 提出EVL-MCoT以解决有害MEME检测中的多视角理解问题 chain-of-thought

🔬 支柱二:RL算法与架构 (RL & Architecture) (5 篇)

#题目一句话要点标签🔗
8 Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning 提出视觉显著性引导蒸馏以解决多模态推理问题 distillation large language model multimodal
9 RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding 提出RadSight以解决医学多模态图像理解的可靠性问题 curriculum learning large language model multimodal
10 Twins: Learn to Predict Unified Representations with Focal Loss 提出Twins模型以解决多模态统一表示优化不平衡问题 flow matching classifier-free guidance multimodal
11 SiPhy: Single-Image Physical Property Reasoning 提出SiPhy以解决单图像物理属性推理问题 MAE embodied AI
12 Spectral Prior for Reducing Exposure Bias in Diffusion Models 提出谱先验以减少扩散模型中的曝光偏差问题 flow matching classifier-free guidance

🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)

#题目一句话要点标签🔗
13 Visual Relocalization from Sparse Views in Aliased and Low-Texture Environments via Novel View Synthesis 提出一种新方法以解决低纹理环境中的视觉重定位问题 3D gaussian splatting 3DGS gaussian splatting
14 Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment Interactions 提出时间反演成像以重建人类与环境的互动 scene understanding multimodal
15 SM4RT: Learning Structured Motion Geometry for 4D Reconstruction 提出SM4RT以解决4D动态理解中的运动结构问题 3D reconstruction motion reconstruction foundation model
16 JustDepth: Real-Time Radar-Camera Depth Estimation with Single-Scan LiDAR Supervision 提出JustDepth以解决雷达-摄像头深度估计问题 depth estimation
17 Deformable Triangle Splatting: Flexible Primitives for Real-Time Radiance Field Rendering 提出可变形三角形溅射技术以解决非凸形状渲染问题 splatting

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
18 AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment 提出AgentHOI以解决人机交互视频生成中的运动控制问题 motion planning implicit representation text-to-motion

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
19 The Lift Spectrum: How Measurement-to-Space Adaptivity Shapes Robustness in Image-Free Single-Pixel Sensing 提出Lift Spectrum以提升图像无关单像素传感的鲁棒性 spatiotemporal

⬅️ 返回 cs.CV 首页 · 🏠 返回主页