| 1 |
MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings |
提出MeetingToM以解决多方会议中的心智理论推理问题 |
large language model multimodal |
|
|
| 2 |
Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models |
提出大规模提示设计方法以优化指令遵循与模型记忆 |
large language model instruction following |
✅ |
|
| 3 |
CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness |
提出CASE框架以解决链式推理的忠实性问题 |
large language model chain-of-thought |
✅ |
|
| 4 |
Stochastic Meta-Unlearning: Bridging Language Backbone and Multimodal Unlearning |
提出随机元遗忘以解决多模态模型的遗忘问题 |
multimodal |
|
|
| 5 |
Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA |
提出FiT框架以优化小型LLM在网络安全问答中的应用 |
large language model instruction following |
|
|
| 6 |
DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning |
提出依赖感知中间问答监督以解决复杂推理问题 |
chain-of-thought |
|
|
| 7 |
AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism |
提出AutoJourn以解决自动化新闻中的多视角生成与偏见检测问题 |
large language model |
|
|
| 8 |
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness |
提出双层元评分标准以评估开放式生成的事实完整性 |
multimodal |
|
|
| 9 |
Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM |
提出SSC-GRPO以解决LLM推理中的幻觉问题 |
large language model |
|
|