| 15 |
Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs |
提出Groc-PO以解决多模态大语言模型的真实度问题 |
DPO direct preference optimization large language model |
|
|
| 16 |
SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning |
提出SIVA-RL以解决多模态强化学习中的视觉对齐问题 |
reinforcement learning multimodal |
|
|
| 17 |
From Surface Forecasting to Observability Forecasting: A Latent World Model for Cloud-Aware EO Monitoring |
提出云监测的可观测性预测模型以提升地球观测效率 |
world model worldmodel world models |
|
|
| 18 |
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders |
提出VideoRAE以解决视频生成模型的表示学习问题 |
JEPA foundation model |
✅ |
|
| 19 |
From Pixels to States: Rethinking Interactive World Models as Game Engines |
提出互动世界模型作为游戏引擎以解决游戏交互性问题 |
world model world models |
|
|
| 20 |
Towards Spatial Supersensing in the Wild |
提出VSI-Super-Wild以解决真实场景中的空间超感知问题 |
world model world models multimodal |
|
|
| 21 |
OvisOCR2 Technical Report |
提出OvisOCR2以解决文档解析问题 |
reinforcement learning distillation reward design |
✅ |
|
| 22 |
ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning |
提出ThinkBLOX以解决3D室内场景生成中的交互编辑问题 |
reinforcement learning chain-of-thought |
|
|
| 23 |
Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection |
提出基于共生启发的知识蒸馏方法以解决增量目标检测问题 |
distillation |
|
|
| 24 |
The 2nd International StepUP Competition for Biometric Footstep Recognition: From Steps to Strides |
提出基于压力的脚步生物识别方法以解决用户识别挑战 |
representation learning spatiotemporal |
|
|
| 25 |
Beyond Color Geometry: Evaluating Human-Like Color Representations in Vision Models |
提出模糊感知模型以评估视觉模型中的人类色彩表现 |
masked autoencoder MAE |
|
|