cs.LG(2026-07-22)

📊 共 20 篇论文 | 🔗 1 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (9 🔗1) 支柱九:具身大模型 (Embodied Foundation Models) (9) 支柱七:动作重定向 (Motion Retargeting) (1) 支柱四:生成式动作 (Generative Motion) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (9 篇)

#题目一句话要点标签🔗
1 Dreamer-CPC: Message Learning with World Models for Decentralized Multi-agent Reinforcement Learning 提出Dreamer-CPC以解决多智能体强化学习中的信息共享问题 reinforcement learning world model world models
2 Koopman Dreamer: Spectrally Constrained Latent Dynamics for Stable World-Model Imagination 提出Koopman Dreamer以解决长时间序列控制中的稳定性问题 world model world models dreamer
3 The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL 提出梦境排练方法以解决持续模型基础强化学习中的遗忘问题 reinforcement learning world model world models
4 Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design 提出AAMFM模型以解决抗原特异性抗体设计问题 DPO direct preference optimization foundation model
5 User-Centric Modeling of Transactional Sequences with Explainable State Space Models 提出用户中心的交易序列建模方法以解决现有模型不足问题 Mamba SSM state space model
6 Active Inference as a Convex Markov Decision Process 将主动推理框架化为凸马尔可夫决策过程 reinforcement learning world model world models
7 Generalized Kalman filter based temporal difference reinforcement learning 提出基于广义卡尔曼滤波的时间差强化学习框架 reinforcement learning
8 Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence 提出无视时间依赖的强化学习后悔最小化算法 reinforcement learning
9 How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF 提出C++推理引擎以提升RLHF中奖励模型评分速度 RLHF

🔬 支柱九:具身大模型 (Embodied Foundation Models) (9 篇)

#题目一句话要点标签🔗
10 Zero-Shot Heart Rate Variability Forecasting from Consumer Wearables Using Time Series Foundation Models 提出时间序列基础模型以解决可穿戴设备心率变异性预测问题 foundation model
11 Post-Training in Time Series Foundation Models: A Unifying Framework 提出统一框架以优化时间序列基础模型的后训练过程 foundation model
12 Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data 提出假设-精炼学习以解决有机结构解析问题 multimodal
13 AlphaRoute: Large Language Models as Semantic Optimizers for Multi-Objective Routing 提出AlphaRoute以解决VLSI全局路由的多目标优化问题 large language model
14 Expert-Guided Forecast Editing for Time-Series Foundation Models 提出DEFT框架以优化时间序列预测中的专家反馈整合 foundation model
15 The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability 提出PortLLM以解决大语言模型的长期适应性问题 large language model
16 Co-Evolving LLM Evaluators and Policies via DynamicRubric 提出DynamicRubric以解决大语言模型评估与策略优化瓶颈问题 large language model
17 An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies 提出一种保持各向同性的谱帽以优化Muon训练效果 large language model
18 Efficient Clustering with Provable Guardrails for LLM Inference at Scale 提出高效聚类方法以降低大规模LLM推理成本 foundation model

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
19 OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization 提出OPIUM以解决激活引导的外部性和过度拒绝问题 latent optimization large language model

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
20 Analytic Distribution of Classifier-Free Guidance for Schedule Design 提出分布引导的无分类器指导以优化扩散模型调度 classifier-free guidance

⬅️ 返回 cs.LG 首页 · 🏠 返回主页