| 1 |
Large Language Models as Unified Multimodal Learners for Clinical Prediction |
提出统一多模态学习模型以简化临床预测系统 |
large language model multimodal |
|
|
| 2 |
Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior |
提出可转移的思维链条以应对语言模型的有害行为 |
chain-of-thought |
|
|
| 3 |
AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models |
提出AuAu基准以审计大型语言模型中的威权主义倾向 |
large language model |
|
|
| 4 |
An MLIR-Based Compilation Method for Large Language Models |
提出基于MLIR的编译方法以解决大语言模型部署问题 |
large language model |
|
|
| 5 |
PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment |
提出PolyInterview以解决面试准备中的实践不足问题 |
multimodal |
|
|
| 6 |
Verbalizable Representations Form a Global Workspace in Language Models |
提出Jacobian透镜技术揭示语言模型的全球工作空间特征 |
large language model |
|
|
| 7 |
VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs |
提出VarRate以解决长上下文LLM推理中的KV缓存瓶颈问题 |
large language model |
|
|
| 8 |
DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods |
提出DECODEM以解决企业组织文件数据提取问题 |
large language model |
|
|
| 9 |
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability |
通过机制可解释性分析细调LLM中的道德偏见 |
large language model |
|
|
| 10 |
Scaling Point-in-Time Language Models |
提出点时语言模型以解决未来信息泄露问题 |
large language model |
|
|
| 11 |
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors |
提出人格向量审计开放权重LLM以揭示模型行为 |
chain-of-thought |
|
|
| 12 |
EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections |
提出EpiNarrate以解决公共卫生叙事生成问题 |
large language model |
|
|
| 13 |
From Plausible to Actionable: A Position on LLM Self-Explanations |
提出自我解释的可行性评估框架以提升LLM的可解释性 |
large language model |
|
|
| 14 |
BayesPO: Bayesian Prompt Optimization via Parallel-Tempered Gradient-Guided Discrete MCMC |
提出BayesPO以优化大语言模型的提示生成 |
large language model |
|
|
| 15 |
Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D |
提出2D-RoPE以解决语言模型复制问题 |
large language model |
|
|
| 16 |
From Articles to Premises: Building PrimeFacts, an Extraction Methodology and Resource for Fact-Checking Evidence |
提出PrimeFacts以解决事实核查证据提取问题 |
large language model |
|
|
| 17 |
What Should a Skill Remember? Quality--Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model Agents |
提出成本意识的技能重写方法以优化语言模型代理的性能 |
large language model |
|
|
| 18 |
It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO |
提出单例GRPO以揭示大语言模型的偏见脆弱性 |
large language model |
|
|
| 19 |
Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila |
提出Pancasila-Dilemmas评估数据集以解决印尼价值观对齐问题 |
large language model |
✅ |
|
| 20 |
Large Language Models for Citation Function Classification |
提出大型语言模型进行引文功能分类以提升文献分析 |
large language model |
|
|
| 21 |
How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks |
提出三维评估框架以提升多模态对话状态信号的可靠性 |
multimodal |
|
|
| 22 |
Benchmarking Resource-Efficient LLMs for Research Topic Ontology Generation in the Biomedical Field |
提出小型LLM以解决生物医学领域本体生成问题 |
large language model chain-of-thought |
|
|
| 23 |
PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning |
提出PPL-Factory以解决数据选择中的任务感知与预算感知问题 |
large language model |
|
|
| 24 |
DeLIVeR: Decomposed Learning for Information-grounded Veracity Recognition via Reinforced Knowledge Graph Exploration |
提出DeLIVeR以解决自动化事实核查中的查询脆弱性问题 |
large language model |
|
|
| 25 |
Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI |
提出HALO框架以解决企业AI的幻觉问题 |
large language model |
|
|
| 26 |
It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief |
提出一种评估EoB对LLM响应影响的分类法 |
large language model |
|
|
| 27 |
VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval |
提出VDAR-Router以解决LLM路由中的查询难度分析问题 |
large language model |
|
|
| 28 |
When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs |
提出文化纠缠同形异义词消歧义方法以解决低资源语言模型的偏见问题 |
chain-of-thought |
✅ |
|
| 29 |
ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions |
提出ESCUCHA基准以评估西班牙语在多样化声学条件下的理解能力 |
multimodal |
|
|
| 30 |
C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference |
提出C$^2$KV以解决长上下文LLM推理中的KV缓存存储与访问问题 |
large language model |
|
|
| 31 |
D-NOVA: In-Storage Retrieval Accelerator via Dual-Bound 3D NAND-Optimized Similarity Search with Vector Adaptation |
提出D-NOVA以解决存储检索加速中的性能瓶颈问题 |
large language model |
|
|