| 1 |
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools |
提出VidCoM以解决视频理解与用户指令响应问题 |
large language model multimodal |
|
|
| 2 |
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation |
提出LLM4SGG以解决弱监督场景图生成中的语义简化与低密度问题 |
large language model chain-of-thought |
|
|
| 3 |
Few-shot Action Recognition with Captioning Foundation Models |
提出CapFSAR框架以解决少样本动作识别问题 |
foundation model multimodal |
|
|
| 4 |
BiomedJourney: Counterfactual Biomedical Image Generation by Instruction-Learning from Multimodal Patient Journeys |
提出BiomedJourney以解决生物医学图像生成中的反事实问题 |
multimodal |
|
|
| 5 |
Interpreting and Controlling Vision Foundation Models via Text Explanations |
提出一种框架以通过文本解释理解和控制视觉基础模型 |
foundation model |
|
|
| 6 |
Multimodal Object Query Initialization for 3D Object Detection |
提出EfficientQ3M以解决3D目标检测中的查询初始化问题 |
multimodal |
|
|
| 7 |
Using Global Land Cover Product as Prompt for Cropland Mapping via Visual Foundation Model |
提出基于全球土地覆盖产品的提示学习方法以解决农田映射问题 |
foundation model |
|
|
| 8 |
Automated Natural Language Explanation of Deep Visual Neurons with Large Models |
提出自动化框架以生成深度视觉神经元的语义解释 |
foundation model |
|
|
| 9 |
Loci-Segmented: Improving Scene Segmentation Learning |
提出Loci-Segmented以解决场景分割学习中的背景依赖问题 |
foundation model |
|
|
| 10 |
A Multi-Scale Spatial Transformer U-Net for Simultaneously Automatic Reorientation and Segmentation of 3D Nuclear Cardiac Images |
提出多尺度空间变换U-Net以解决3D核心脏图像的自动重定向与分割问题 |
multimodal |
|
|
| 11 |
Black-box Targeted Adversarial Attack on Segment Anything (SAM) |
提出黑箱针对性对抗攻击方法以评估SAM模型的鲁棒性 |
foundation model |
|
|
| 12 |
Towards Unified and Effective Domain Generalization |
提出UniDG框架以提升领域泛化性能 |
foundation model |
✅ |
|