| 1 |
Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities |
提出DataGovBench以解决LLM在真实数据分析中的不足问题 |
large language model |
✅ |
|
| 2 |
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability |
提出Pluralis v0.1以解决AI安全评估中的文化偏见问题 |
multimodal |
|
|
| 3 |
MemDefrag: Latent Memory Defragmentation for Large Language Models |
提出MemDefrag以解决大语言模型的潜在记忆碎片化问题 |
large language model |
|
|
| 4 |
SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation |
提出SpanUQ以解决大语言模型生成中的不确定性量化问题 |
large language model |
|
|
| 5 |
From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b |
提出针对生物医学问答的LLM框架以提升答案稳健性 |
large language model chain-of-thought |
|
|
| 6 |
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages |
提出PluraMath以解决数学推理评估的语言偏见问题 |
large language model instruction following |
|
|
| 7 |
LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis |
提出LongCrafter以解决长上下文理解的多样性问题 |
large language model |
|
|
| 8 |
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability |
提出LLM代理以解决部分可观测联合决策问题 |
large language model |
|
|