| 12 |
Toward Active Object Detection for UAVs in the Wild: A Large-Scale Dataset, Benchmark and Method |
提出ATRNet-LUDO数据集以解决无人机主动目标检测问题 |
reinforcement learning deep reinforcement learning DRL |
✅ |
|
| 13 |
ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts |
提出ALICE模型以整合多模态病理学知识 |
distillation foundation model multimodal |
✅ |
|
| 14 |
Joint-Embedding Predictive Architecture for Solar PV Panel Fault Classification |
提出JEFFNet以解决太阳能光伏面板故障分类问题 |
JEPA Joint-Embedding Predictive Architecture joint-embedding predictive architecture |
✅ |
|
| 15 |
Video Generation Models are General-Purpose Vision Learners |
提出GenCeption以实现通用视觉智能 |
JEPA MAE Depth Anything |
|
|
| 16 |
Causally Debiased Latent Action Model for Embodied Action Conditioned World Models |
提出CD-LAM以解决可控动作模型中的偏差问题 |
world model world models contrastive learning |
|
|
| 17 |
Multimodal Scenario Similarity Search for Autonomous Driving |
提出多模态框架以提升自动驾驶场景检索效率 |
contrastive learning multimodal |
|
|
| 18 |
Scalable Visual Pretraining for Language Intelligence |
提出可扩展视觉预训练以提升语言智能 |
visual pre-training foundation model |
|
|
| 19 |
Weaving Light and Time: Unified Harmonic-Geometric Representation Learning for Dense RGB-Event Parsing |
提出Evita以解决RGB-事件解析中的多模态融合问题 |
representation learning multimodal |
✅ |
|
| 20 |
IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation |
提出信息瓶颈引导的CFG蒸馏方法以解决文本到图像生成中的推理延迟问题 |
distillation classifier-free guidance |
|
|
| 21 |
Probing Diffusion Denoising Dynamics for Contrastive Representation Learning |
提出D$^3$CL以支持对比表示学习的去噪动态适应 |
representation learning contrastive learning |
|
|
| 22 |
MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models |
提出MOSAIC以解决异构视觉语言模型的优化问题 |
linear attention distillation |
|
|