| 1 |
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding |
提出ClinFusion以解决医疗领域多模态理解问题 |
large language model multimodal instruction following |
|
|
| 2 |
Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding |
提出Mixture-of-Thought-Tokens以解决多模态定位与推理统一问题 |
large language model multimodal |
|
|
| 3 |
IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data |
通过合成训练数据适应基础模型以提升人脸识别性能 |
foundation model |
|
|
| 4 |
Disentangling Semantic Attention from Structural Bias in the Attention Manifold |
提出SPAR方法以解决多模态语言模型中的视觉注意力偏差问题 |
large language model multimodal visual grounding |
|
|
| 5 |
MMOE: Modernizing Diffusion Transformers with Efficient Expert Design |
提出MMOE以平衡生成质量与训练成本问题 |
large language model foundation model |
|
|
| 6 |
The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding |
提出稀疏帧适应策略以解决视频定位中的性能瓶颈问题 |
large language model multimodal |
|
|
| 7 |
ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation |
提出ReflexTrack以解决训练无关的引用视频目标分割问题 |
large language model multimodal |
|
|
| 8 |
Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking |
提出时空条件去噪变换器以解决模态缺失的RGBT跟踪问题 |
multimodal |
|
|
| 9 |
Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels |
提出无坐标和区域标签的证据归属方法以解决文档理解问题 |
multimodal |
|
|
| 10 |
CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding |
提出CADER以解决长视频理解中的动态证据推理问题 |
chain-of-thought |
|
|
| 11 |
DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes |
提出DecoupleMix以解决VLM数据构建的系统性问题 |
multimodal |
|
|
| 12 |
Rethinking Expert Training for Model Merging with Prompt Learning |
提出双调专家以优化模型合并中的专家训练 |
foundation model |
|
|
| 13 |
UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing |
提出UMI3D以解决多图像输入下3D生成质量下降问题 |
foundation model |
|
|
| 14 |
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation |
提出FilmBench以解决视频生成评估标准不足的问题 |
multimodal |
|
|
| 15 |
DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection |
提出DuoAD以解决训练无关的少样本异常检测问题 |
foundation model |
|
|
| 16 |
Embeddings based Anomaly Detection for Cleaning Global Crop Type Reference Datasets |
提出基于嵌入的异常检测以清理全球作物类型参考数据集 |
foundation model |
|
|