| 1 |
EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval |
提出EvoGraph-R1以解决静态知识图谱在多模态检索中的局限性 |
large language model multimodal |
|
|
| 2 |
Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks |
审计多模态基准中的数据泄露问题以提升WSI VQA评估准确性 |
foundation model multimodal |
|
|
| 3 |
ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models |
提出ViCo3D以解决LiDAR基础的协作3D目标检测问题 |
foundation model |
|
|
| 4 |
VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression |
提出VisCo以解决视觉标记压缩效率低的问题 |
large language model |
|
|
| 5 |
Open-KNEAD: Knowledge-grounded Nutrition Estimation via Agentic Decomposition |
提出Open-KNEAD框架以解决饮食营养估计问题 |
large language model multimodal |
|
|
| 6 |
Hy-Embodied-VLM-1.0: Efficient Physical-World Agents |
提出Hy-Embodied-VLM-1.0以提升物理世界代理的能力 |
foundation model multimodal |
|
|
| 7 |
ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning |
提出ReflectVLN以解决视觉语言导航中的闭环跟踪问题 |
VLN chain-of-thought |
✅ |
|
| 8 |
IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment |
提出IQA-T1以解决开放世界图像质量评估问题 |
large language model multimodal |
✅ |
|
| 9 |
CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models |
提出CoRe框架以解决跨图像比较推理问题 |
multimodal |
|
|
| 10 |
HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and Ambivalence/Hesitancy Video Recognition |
提出多任务学习方法以解决情感视频识别问题 |
multimodal |
|
|
| 11 |
Text-Aided Multi-Modal Panoptic Symbol Spotting for CAD Floor Plan Drawings |
提出TextCAD以解决CAD平面图符号识别中的文本注释利用不足问题 |
multimodal |
|
|
| 12 |
Towards Vision-Free CIR: Attribute-Augmented Scoring and LLM-Based Reranking for Zero-Shot Composed Image Retrieval |
提出视觉无关的CIR框架以解决复杂图像检索问题 |
multimodal |
|
|
| 13 |
More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization |
提出CROSS框架以解决极端视角下的跨视图定位问题 |
foundation model |
|
|
| 14 |
UMSS: Towards Unsupervised Multi-modal Semantic Segmentation |
提出UniM2以解决无监督多模态语义分割问题 |
multimodal |
|
|
| 15 |
How to Realize Recursively Self-Improving Agents and Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture |
提出多智能体架构以实现递归自我改进的代理与个人奇点 |
large language model |
|
|