| 1 |
LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action |
提出LEEVLA以解决复杂动态场景中的视觉-语言-动作问题 |
vision-language-action VLA multimodal |
✅ |
|
| 2 |
CT-CLIP Representations for Multimodal Lung Cancer Survival Prediction |
提出CT-CLIP表示以解决肺癌生存预测问题 |
foundation model multimodal |
|
|
| 3 |
UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery |
提出UniRef-UAV以解决无人机图像中的多模态指代问题 |
multimodal visual grounding |
|
|
| 4 |
Predicting Viticulture Potential through an Ensemble of U-Net and a Geospatial Foundation Model |
通过U-Net与地理基础模型集成预测葡萄种植潜力 |
foundation model |
✅ |
|
| 5 |
DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models |
提出DeltaV以解决多模态模型中视觉状态冗余问题 |
multimodal |
✅ |
|
| 6 |
Multimodal 3D LUT Generation via StatLUT with Statistical Features for Photorealistic Style Transfer |
提出StatLUT以解决光写实风格迁移中的色彩与结构问题 |
multimodal |
|
|
| 7 |
VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness |
提出VSRo-200数据集以研究罗马尼亚视觉语音识别的监督与多模态鲁棒性 |
multimodal |
|
|
| 8 |
Post-Training in End-to-End Autonomous Driving |
提出后训练技术以提升自动驾驶模型的可靠性 |
vision-language-action multimodal |
|
|
| 9 |
LUMI: Tokenizer-Agnostic LLM-Based Lossless Image Compression |
提出LUMI框架以解决图像无损压缩中的tokenizer依赖问题 |
large language model foundation model |
|
|
| 10 |
OpenCoF: Learning to Reason Through Video Generation |
提出OpenCoF框架以解决视频生成中的推理能力不足问题 |
chain-of-thought |
|
|
| 11 |
When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities |
提出结构稀疏自编码器以解决视觉语言模型中的概念一致性问题 |
multimodal |
|
|
| 12 |
Dive Into the Implicit Biases of Low-rank Vision-language Alignment |
提出低秩适应方法以优化视觉-语言对齐 |
large language model |
|
|
| 13 |
Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark |
提出双关联超图网络以解决RGB-T视频目标检测中的对齐问题 |
multimodal |
|
|
| 14 |
Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing |
提出混合探测器以解决多模态大语言模型训练中的特权模态问题 |
large language model multimodal |
✅ |
|
| 15 |
Is sub-metre resolution necessary for cocoa mapping? A landscape-stratified evaluation of very high resolution imagery, decametric Earth Observation inputs, and operational products in Cote d'Ivoire |
通过高分辨率影像提升可可种植区映射精度 |
foundation model |
|
|