| 1 |
Large Language Models as Generalizable Policies for Embodied Tasks |
提出LLaRP以解决视觉任务中的通用策略问题 |
reinforcement learning egocentric embodied AI |
|
|
| 2 |
CQM: Curriculum Reinforcement Learning with a Quantized World Model |
提出CQM以解决高维空间中课程强化学习的目标生成问题 |
reinforcement learning world model world models |
|
|
| 3 |
FedPEAT: Convergence of Federated Learning, Parameter-Efficient Fine Tuning, and Emulator Assisted Tuning for Artificial Intelligence Foundation Models with Mobile Edge Computing |
提出FedPEAT以解决大规模模型在边缘计算中的调优问题 |
reinforcement learning deep reinforcement learning foundation model |
|
|
| 4 |
Understanding and Addressing the Pitfalls of Bisimulation-based Representations in Offline Reinforcement Learning |
提出期望算子以解决离线强化学习中的双模拟表示问题 |
reinforcement learning offline RL offline reinforcement learning |
✅ |
|
| 5 |
Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 Tricks |
将DreamerV3技巧应用于PPO以提升奖励规模鲁棒性 |
reinforcement learning PPO dreamer |
|
|
| 6 |
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-Free Reinforcement Learning Updates |
提出动态不变数据增强策略以提升无模型强化学习的数据效率 |
reinforcement learning |
|
|
| 7 |
Combating Representation Learning Disparity with Geometric Harmonization |
提出几何协调方法以解决长尾分布下的表示学习不均衡问题 |
representation learning |
✅ |
|
| 8 |
Fair collaborative vehicle routing: A deep multi-agent reinforcement learning approach |
提出深度多智能体强化学习解决公平协作车辆路径规划问题 |
reinforcement learning |
|
|
| 9 |
Coalitional Bargaining via Reinforcement Learning: An Application to Collaborative Vehicle Routing |
通过强化学习提出合作博弈解决协作车辆调度问题 |
reinforcement learning |
|
|
| 10 |
Demonstration-Regularized RL |
提出演示正则化强化学习以提高样本效率 |
reinforcement learning behavior cloning RLHF |
|
|
| 11 |
CROP: Conservative Reward for Model-based Offline Policy Optimization |
提出CROP以解决离线强化学习中的过估计问题 |
reinforcement learning offline RL offline reinforcement learning |
|
|
| 12 |
Spatio-Temporal Meta Contrastive Learning |
提出CL4ST框架以解决时空图神经网络数据稀缺问题 |
contrastive learning |
|
|
| 13 |
Learning Regularized Graphon Mean-Field Games with Unknown Graphons |
提出GMFG-PPO算法以解决未知图谱的均衡学习问题 |
reinforcement learning PPO |
|
|
| 14 |
DSAC-C: Constrained Maximum Entropy for Robust Discrete Soft-Actor Critic |
提出DSAC-C以增强离散软演员评论家的鲁棒性 |
reinforcement learning SAC |
|
|