| 1 |
Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models |
提出预算感知的测试时间模型选择以优化大语言模型的响应质量 |
large language model |
|
|
| 2 |
Eigenvalue Calibration for Semantic Embeddings of Large Language Models |
提出新框架以校准大型语言模型的语义嵌入特征 |
large language model |
|
|
| 3 |
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization |
提出系统感知的KV缓存优化以提升大语言模型服务效率 |
large language model |
|
|
| 4 |
Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment |
提出小型视觉语言模型量化的新框架以优化边缘部署 |
large language model multimodal |
|
|
| 5 |
Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA |
提出自验证LLM风险分析工具以解决安全分析盲点问题 |
large language model |
|
|
| 6 |
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents |
提出统一的内存压缩框架以优化大型语言模型的上下文管理 |
large language model |
|
|
| 7 |
Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models |
提出敏感度感知阈值和令牌路由以优化大语言模型的激活稀疏化 |
large language model |
|
|
| 8 |
Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models |
提出预算感知的测试时模型选择方法以优化大语言模型的响应质量 |
large language model |
|
|
| 9 |
NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision |
提出NL-PAC框架以解决LLM监督中的规范模糊性问题 |
large language model |
|
|
| 10 |
TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning |
提出TSRouter以解决时间序列推理中的动态模态选择问题 |
large language model |
✅ |
|
| 11 |
BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving |
提出BlockServe以解决扩散大语言模型服务中的收敛异质性问题 |
large language model |
|
|
| 12 |
Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal |
提出激活引导的对抗后缀优化以增强模型安全性 |
large language model |
|
|