| 1 |
LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models |
提出LLM4EHR以解决临床时间序列与医疗事件序列对齐问题 |
representation learning large language model foundation model |
|
|
| 2 |
SODA: Semi On-Policy Black-Box Distillation for Large Language Models |
提出SODA以解决大语言模型蒸馏中的效率与稳定性问题 |
distillation large language model |
|
|
| 3 |
Dichotomous Diffusion Policy Optimization |
提出DIPOLE以解决扩散策略优化中的不稳定性问题 |
reinforcement learning diffusion policy vision-language-action |
|
|
| 4 |
CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward Approach |
提出情感增强深度强化学习以解决主动交易问题 |
reinforcement learning deep reinforcement learning PPO |
|
|
| 5 |
Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems |
提出PEARL以解决高维动态系统的实时最优控制问题 |
reinforcement learning policy learning sparse sensors |
|
|
| 6 |
ADS-C: Antidistillation Sampling for Classification |
提出ADS-C以解决分类任务中的蒸馏攻击问题 |
distillation large language model |
|
|
| 7 |
QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides |
提出QUADS以解决MoE强化学习中的NVFP4不稳定问题 |
reinforcement learning large language model |
|
|
| 8 |
SC-JEPA: Stabilizing Latent Predictive Learning for Time-Series Anomaly Prediction |
提出SC-JEPA以解决时间序列异常预测中的不稳定性问题 |
JEPA predictive model distillation |
|
|
| 9 |
The Terminal Representation in Reinforcement Learning |
提出终端表示法以改进强化学习中的状态表示 |
reinforcement learning representation learning reward shaping |
|
|
| 10 |
DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning |
提出DADiff以解决强化学习中的跨域策略适应问题 |
reinforcement learning representation learning |
|
|
| 11 |
A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model |
提出经典与量子算法以优化强化学习中的生成模型 |
reinforcement learning offline reinforcement learning |
|
|
| 12 |
Robust Peak-cost Constrained Reinforcement Learning |
提出稳健的峰值成本约束强化学习以解决安全关键问题 |
reinforcement learning |
|
|
| 13 |
On the Failure of Boundary-Seeking Distillation in Bottlenecked Generative Architectures |
提出边界寻求蒸馏方法以解决瓶颈生成架构的知识转移问题 |
distillation |
|
|
| 14 |
When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis |
提出模型合并方法以挑战联合多任务强化学习 |
reinforcement learning |
|
|
| 15 |
Understanding Reasoning from Pretraining to Post-Training |
提出棋类作为控制测试平台以研究预训练与后训练的推理关系 |
reinforcement learning large language model |
|
|
| 16 |
When Does Muon Help Agentic Reinforcement Learning? |
提出Muon优化算法以提升稀疏奖励强化学习性能 |
reinforcement learning |
|
|
| 17 |
Prediction-Only Distillation in Linear and Logistic Regression |
提出预测仅蒸馏方法以解决训练数据缺乏问题 |
distillation |
|
|
| 18 |
Learning Standard Model structure from LHC data with Riemannian flow matching |
提出ShellFlow模型以从LHC数据中学习标准模型结构 |
flow matching |
|
|
| 19 |
CTC: The Composite Task Challenge for Cooperative Multi-Agent Reinforcement Learning |
提出复合任务挑战以解决合作多智能体强化学习中的分工不足问题 |
reinforcement learning |
|
|
| 20 |
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone |
提出DiffuMamba以提升扩散语言模型的推理效率 |
Mamba |
|
|
| 21 |
Soft $Q(λ)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces |
提出Soft $Q(λ)$以解决多步离策略熵正则化强化学习问题 |
reinforcement learning |
|
|
| 22 |
Beyond Objective Expressivity: Geometry Preservation in Multimodal Contrastive Learning |
提出几何保持编码器以解决多模态对比学习中的优化问题 |
contrastive learning multimodal |
|
|
| 23 |
FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications |
提出FlashRT以优化多模态应用的实时部署 |
world model world models multimodal |
|
|
| 24 |
A Geometric Perspective on Stabilizing Value Conflict Resolution |
提出链式思维以解决大语言模型的价值冲突问题 |
reinforcement learning RLHF large language model |
|
|
| 25 |
MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models |
提出MADA-RL以解决紧凑模型推理效率低的问题 |
reinforcement learning large language model |
|
|
| 26 |
Information-Based Exploration via Random Features for Reinforcement Learning |
提出随机特征信息增益方法以优化强化学习探索策略 |
reinforcement learning deep reinforcement learning representation learning |
|
|
| 27 |
AGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models |
提出JAGG以解决扩散模型GRPO训练的计算瓶颈问题 |
reinforcement learning flow matching large language model |
|
|
| 28 |
OR Else: A Differentiable Trust Region for Policy Optimization |
提出可微信任区域方法以优化策略训练 |
PPO RLHF large language model |
|
|
| 29 |
Adaptive Mamba Neural Operators |
提出自适应Mamba神经算子以解决偏微分方程问题 |
Mamba SSM |
|
|
| 30 |
Planning with Transformers: Chain of Computation and Structured Context Windows |
提出链式计算架构以解决规划问题 |
world model world models large language model |
|
|
| 31 |
Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal Inference |
提出基于RAG的策略学习方法以解决因果推断中的邻近匹配问题 |
policy learning |
|
|
| 32 |
Enhancing Rubric-based RL via Self-Distillation |
提出CriPO以解决Rubric-based RL中的探索不足问题 |
distillation |
|
|
| 33 |
Theoretical Foundations of $\max$@$k$ Reinforcement Learning |
提出$ ext{max}@k$强化学习理论基础以解决评估问题 |
reinforcement learning |
|
|
| 34 |
A Weisfeiler-Leman Characterization of Global-Attention Graph Transformers for Mixed-Integer Linear Programs |
提出图基础模型以解决混合整数线性规划的表达能力问题 |
linear attention foundation model |
|
|
| 35 |
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks |
提出体验学习方法以提升非可验证任务的学习效果 |
reinforcement learning distillation |
|
|
| 36 |
Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization |
提出基于优势聚合的多智能体策略优化方法 |
reinforcement learning PPO |
|
|
| 37 |
Distributional Soft Bellman Operator under the Cramér Geometry |
提出基于Cramér几何的分布式软贝尔曼算子以优化策略评估 |
reinforcement learning DRL |
|
|