| 1 |
Interpreting Learned Feedback Patterns in Large Language Models |
提出学习反馈模式以提高大语言模型的反馈一致性 |
reinforcement learning RLHF large language model |
|
|
| 2 |
Beyond Traditional DoE: Deep Reinforcement Learning for Optimizing Experiments in Model Identification of Battery Dynamics |
提出基于深度强化学习的实验优化方法以解决电池动态建模问题 |
reinforcement learning deep reinforcement learning |
|
|
| 3 |
Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias |
提出离线重训练方法以缓解在线强化学习中的探索偏差问题 |
reinforcement learning policy learning offline RL |
✅ |
|
| 4 |
Virtual Augmented Reality for Atari Reinforcement Learning |
提出虚拟增强现实方法以提升Atari强化学习代理性能 |
reinforcement learning foundation model |
|
|
| 5 |
Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining |
提出理论框架以提升变换器在上下文强化学习中的决策能力 |
reinforcement learning offline reinforcement learning distillation |
|
|
| 6 |
Reinforcement Learning of Display Transfer Robots in Glass Flow Control Systems: A Physical Simulation-Based Approach |
提出深度强化学习方法以优化玻璃流控制系统调度 |
reinforcement learning deep reinforcement learning |
|
|
| 7 |
Robustness to Multi-Modal Environment Uncertainty in MARL using Curriculum Learning |
提出基于课程学习的MARL方法以应对多模态环境不确定性问题 |
reinforcement learning curriculum learning |
|
|
| 8 |
MetaBox: A Benchmark Platform for Meta-Black-Box Optimization with Reinforcement Learning |
提出MetaBox以解决Meta-黑箱优化基准缺失问题 |
reinforcement learning DRL |
✅ |
|
| 9 |
Splicing Up Your Predictions with RNA Contrastive Learning |
提出对比学习方法以提升RNA序列预测精度 |
contrastive learning |
|
|
| 10 |
Cross-Episodic Curriculum for Transformer Agents |
提出跨情节课程算法以提升Transformer代理的学习效率 |
reinforcement learning imitation learning |
|
|
| 11 |
Learning RL-Policies for Joint Beamforming Without Exploration: A Batch Constrained Off-Policy Approach |
提出离线模型基于深度Q学习的联合波束成形优化方法 |
reinforcement learning deep reinforcement learning |
|
|
| 12 |
A Symmetry-Aware Exploration of Bayesian Neural Network Posteriors |
提出对深度贝叶斯神经网络后验分布的对称性探索 |
world model world models |
|
|