cs.LG(2026-07-21)

📊 共 16 篇论文

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (11) 支柱九:具身大模型 (Embodied Foundation Models) (4) 支柱一:机器人控制 (Robot Control) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (11 篇)

#题目一句话要点标签🔗
1 From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning 提出LA-MAML以解决元强化学习中的轨迹收集问题 reinforcement learning language conditioned
2 H$^2$SD: Hybrid Hindsight Self-Distillation 提出H$^2$SD框架以解决稀疏监督问题 reinforcement learning distillation privileged information
3 Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information 提出Off-Context GRPO以解决强化学习中的学习信号缺失问题 reinforcement learning privileged information large language model
4 Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation 提出保守查询与自适应正则化以解决离线强化学习中的不确定性问题 reinforcement learning offline RL offline reinforcement learning
5 Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation 提出LLMol框架以解决分子生成中的优化问题 reinforcement learning large language model
6 Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning 提出Staleness-Adaptive Trust Region以解决异步强化学习中的更新不稳定问题 reinforcement learning PPO
7 Exposure-Based Reinforcement Learning to Rank 提出基于曝光的强化学习以优化排序问题 reinforcement learning distillation
8 A Reinforcement-Learning-Augmented Liquid-Fueled Reactor Network Model for Predicting Lean Blowout in Gas Turbine Combustors 提出强化学习框架以改善燃气涡轮燃烧器的低燃烧极限预测 reinforcement learning
9 S3: Stable Subgoal Selection by Constraining Uncertainty of Coarse Dynamics in Hierarchical Reinforcement Learning 提出基于粗糙动态约束不确定性的稳定子目标选择方法 reinforcement learning
10 AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters 提出AdaFlash以解决扩散草稿模型的高方差问题 distillation large language model
11 ISO: An RLVR-Native Optimization Stack 提出ISO框架以优化RLVR中的奖励反馈转换问题 reinforcement learning distillation

🔬 支柱九:具身大模型 (Embodied Foundation Models) (4 篇)

#题目一句话要点标签🔗
12 ATLAS: A Foundation Neural Sampler for Amorphous Materials 提出ATLAS以高效采样无定形材料 large language model foundation model
13 GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models 提出GUIDED以解决GNN模型空间转移性问题 multimodal
14 Breaking the Homogeneity Assumption: Specialized Multi-Generator Adversarial Learning for Rare Failure Detection in Predictive Maintenance 提出专用多生成对抗学习以解决预测性维护中的稀有故障检测问题 multimodal
15 Subject-Conditioned Glucose Forecasting in Type-1 Diabetes 提出SCGP以解决1型糖尿病个性化血糖预测问题 multimodal

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
16 Deep learning-based prediction of time-resolved adhesive forces in viscoelastic Hertzian contacts 提出深度学习模型以快速预测粘附性粘弹接触中的力学响应 manipulation

⬅️ 返回 cs.LG 首页 · 🏠 返回主页