| 14 |
AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning |
提出AlphaWiSE以解决多模态持续学习中的对齐问题 |
representation learning multimodal |
|
|
| 15 |
FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models |
提出FoMoVLA以解决视觉预测与运动指导的结合问题 |
policy learning vision-language-action VLA |
✅ |
|
| 16 |
Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models |
提出基于蒸馏的MIL预训练框架以解决病理图像分析中的挑战 |
distillation foundation model |
✅ |
|
| 17 |
3D Geometric Tooth Alignment Planning via Deep Reinforcement Learning |
提出深度强化学习框架以自动化3D牙齿对齐规划 |
reinforcement learning deep reinforcement learning DRL |
|
|
| 18 |
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding |
提出ViPS框架以提升多模态大语言模型的空间理解能力 |
VIP large language model foundation model |
✅ |
|
| 19 |
HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents |
提出HyMobileAgent以解决移动GUI代理的高效交互问题 |
reinforcement learning reward design foundation model |
|
|
| 20 |
From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting |
提出D2DF框架以解决视频对象移除中的草稿依赖问题 |
distillation optical flow |
|
|
| 21 |
Hierarchical Denoising For Multi-Step Visual Reasoning |
提出HDR框架以解决视频多步推理中的逻辑一致性问题 |
world model world models foundation model |
✅ |
|
| 22 |
WanSong v1.0 Technical Report |
提出WanSong以解决长篇音乐生成的效率与可控性问题 |
distillation foundation model |
|
|
| 23 |
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding |
提出VideoSEMA以解决视频理解中的计算效率问题 |
Mamba |
|
|