cs.CL(2026-07-16)
📊 共 15 篇论文 | 🔗 2 篇有代码
🎯 兴趣领域导航
🔬 支柱九:具身大模型 (Embodied Foundation Models) (10 篇)
🔬 支柱二:RL算法与架构 (RL & Architecture) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 11 | Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models | 提出新方法以解决大语言模型推理蒸馏中的数据降解问题 | distillation large language model | ||
| 12 | SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning | 提出SEED框架以解决稀疏奖励下的决策监督问题 | reinforcement learning policy learning distillation | ✅ | |
| 13 | Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents | 提出多代理框架以模拟和审计政治联盟形成 | reinforcement learning RLHF DPO | ||
| 14 | Gold-Guided Programmatic Distillation for Financial Reasoning over Hybrid Tables and Text | 提出金导向程序蒸馏以解决混合表格与文本的金融推理问题 | distillation large language model | ||
| 15 | Mask-Aware Policy Gradients for Diffusion Language Models | 提出Mask-Aware策略梯度以提升MDLM的推理能力 | reinforcement learning large language model |