| 1 |
Multimodal Taxonomic Conditioning for Generative Plankton Imagery |
提出多模态分类条件生成以解决浮游生物图像不足问题 |
multimodal |
|
|
| 2 |
Toward Interpretable Multimodal Fusion: Heat Conduction Modeling for Hyperspectral and LiDAR Joint Classification |
提出M2Heat以解决多模态融合中的长距离依赖建模问题 |
multimodal |
|
|
| 3 |
OmniKVQuant: KV Cache Quantization for Omni-LLMs |
提出OmniKVQuant以解决Omni-LLMs的KV缓存量化问题 |
large language model multimodal |
✅ |
|
| 4 |
DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis Staging |
提出DINO-Med框架以解决多模态医学影像分析问题 |
foundation model multimodal |
|
|
| 5 |
Language-Augmented Semantic Priors for B-Spline Surface Fitting |
提出语言增强的语义先验以解决B样条曲面拟合问题 |
large language model |
|
|
| 6 |
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation |
提出Vidu S2以实现实时交互、可编辑的空间视频生成 |
instruction following |
|
|
| 7 |
R4Tun: LLM-guided adaptive segmental tunnel lining segmentation in point clouds |
提出R4Tun以解决隧道衬砌分割的自适应问题 |
large language model |
|
|
| 8 |
HALDETECT at ImageEval 2026 Shared Tasks: Answer-First Contrastive Grounding with QLoRA |
提出HALDETECT以解决多模态模型的幻觉检测问题 |
multimodal |
|
|
| 9 |
CamPilot: A Multi-Agent Cinematic Assistant for Camera-Controlled Movie Generation |
提出CamPilot以解决专业电影生成中的镜头控制问题 |
large language model |
|
|