cs.LG(2023-10-09)

📊 共 25 篇论文 | 🔗 2 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (11 🔗1) 支柱九:具身大模型 (Embodied Foundation Models) (10) 支柱一:机器人控制 (Robot Control) (2) 支柱八:物理动画 (Physics-based Animation) (2 🔗1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (11 篇)

#题目一句话要点标签🔗
1 Reinforcement Learning in the Era of LLMs: What is Essential? What is needed? An RL Perspective on RLHF, Prompting, and Beyond 探讨RLHF在大语言模型中的应用与未来发展 reinforcement learning policy learning PPO
2 DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning 提出DiffCPS以解决离线强化学习中的约束策略搜索问题 reinforcement learning offline RL offline reinforcement learning
3 Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement Learning 提出奖励一致性动态模型以解决离线强化学习的泛化问题 reinforcement learning offline reinforcement learning
4 Distributional Soft Actor-Critic with Three Refinements 提出三项改进的分布式软演员评论家以解决价值估计不准确问题 reinforcement learning PPO SAC
5 Planning to Go Out-of-Distribution in Offline-to-Online Reinforcement Learning 提出PTGOOD以解决离线到在线强化学习中的探索问题 reinforcement learning offline RL
6 Multi-timestep models for Model-based Reinforcement Learning 提出多时步模型以解决模型基础强化学习中的预测误差问题 reinforcement learning SAC
7 On Double Descent in Reinforcement Learning with LSTD and Random Features 提出双重下降现象以分析强化学习中的LSTD算法性能 reinforcement learning deep reinforcement learning
8 What do larger image classifiers memorise? 提出对比分析以探讨大规模图像分类器的记忆特性 world model world models distillation
9 When is Agnostic Reinforcement Learning Statistically Tractable? 提出新的复杂度度量以解决无知强化学习的统计可处理性问题 reinforcement learning
10 Knowledge Distillation for Anomaly Detection 提出知识蒸馏方法以提升异常检测模型的部署能力 distillation
11 Imitator Learning: Achieve Out-of-the-Box Imitation Ability in Variable Environments 提出模仿学习以解决多任务适应性问题 reinforcement learning imitation learning

🔬 支柱九:具身大模型 (Embodied Foundation Models) (10 篇)

#题目一句话要点标签🔗
12 Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models 提出Step-Back Prompting以提升大语言模型的推理能力 large language model
13 Transformers and Large Language Models for Chemistry and Drug Discovery 利用变换器和大语言模型解决药物发现中的关键瓶颈 large language model
14 Rethinking Memory and Communication Cost for Efficient Large Language Model Training 提出PaRO策略以平衡大语言模型训练中的内存与通信成本 large language model
15 Foundation Models Meet Visualizations: Challenges and Opportunities 探讨基础模型与可视化的交汇以应对透明性与解释性挑战 foundation model
16 Integration-free Training for Spatio-temporal Multimodal Covariate Deep Kernel Point Processes 提出无积分训练的时空多模态深核点过程模型DKMPP multimodal
17 Scaling Studies for Efficient Parameter Search and Parallelism for Large Language Model Pre-training 提出分布式算法优化以提升大语言模型预训练效率 large language model
18 HyperAttention: Long-context Attention in Near-Linear Time 提出HyperAttention以解决长上下文注意力计算效率问题 large language model
19 Little is Enough: Boosting Privacy by Sharing Only Hard Labels in Federated Semi-Supervised Learning 提出FedCT以通过共享硬标签提升联邦半监督学习的隐私性 large language model
20 Making Scalable Meta Learning Practical 提出SAMA以解决可扩展元学习的实际应用问题 large language model
21 Augmented Embeddings for Custom Retrievals 提出适应性密集检索以解决异构检索问题 large language model

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
22 TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained Models 提出TAIL框架以高效适应新控制任务 manipulation imitation learning language conditioned
23 Memory-Consistent Neural Networks for Imitation Learning 提出记忆一致神经网络以解决模仿学习中的误差累积问题 manipulation imitation learning behavior cloning

🔬 支柱八:物理动画 (Physics-based Animation) (2 篇)

#题目一句话要点标签🔗
24 Automatic Integration for Spatiotemporal Neural Point Processes 提出AutoSTPP以解决时空神经点过程的自动积分问题 spatiotemporal
25 Pre-trained Spatial Priors on Multichannel NMF for Music Source Separation 提出基于预训练空间先验的多通道NMF以解决音乐源分离问题 PULSE

⬅️ 返回 cs.LG 首页 · 🏠 返回主页