cs.LG(2026-09-08)

📊 共 25 篇论文 | 🔗 3 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (16 🔗3) 支柱九:具身大模型 (Embodied Foundation Models) (7) 支柱一:机器人控制 (Robot Control) (1) 支柱八:物理动画 (Physics-based Animation) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (16 篇)

#题目一句话要点标签🔗
1 Routing Dense Layouts with History-Aware Offline Reinforcement Learning using LSTM 提出历史感知的离线强化学习以解决密集布局路由问题 reinforcement learning offline RL offline reinforcement learning
2 Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation 提出探针驱动的熵正则化排名掩蔽策略优化以解决代码生成中的测试时强化学习问题 reinforcement learning open-vocabulary open vocabulary
3 MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models 提出MoEMB以高效扩展通用多模态嵌入模型 contrastive learning multimodal
4 Risk-Conditioned Fine-Tuning of Large Language Models 提出风险条件化的强化学习框架以优化大语言模型的风险控制 RLHF large language model
5 Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems 提出行动条件世界模型以解决地球系统模拟的交互性问题 world model world models
6 Online Signature Verification Using Augmented Path Signature and T-Mamba 提出增强路径签名与T-Mamba模型以解决在线签名验证问题 Mamba SSM state space model
7 PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games 提出PlayTrain框架以高效生成适应性JavaScript游戏 reinforcement learning large language model
8 TV-Regulated OPD: Direction Matters in On-Policy Distillation 提出TV-OPD以解决现有OPD方法不稳定问题 distillation large language model
9 Curriculum Learning as Transport: Understanding Curricula with Wasserstein Geodesics 提出Wasserstein课程路径以解构课程学习的设计选择 curriculum learning
10 Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation 提出On-Policy反向蒸馏以解决弱到强泛化问题 distillation
11 Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks 提出环境侧适应以解决长时间任务中的奖励稀疏问题 reinforcement learning large language model
12 Distillation as Probability Transport: Routed On-Policy Distillation 提出RouteOPD以解决现有蒸馏方法的概率重分配问题 distillation
13 Proactive Context-Forecasted Safety Constraints for Nonstationary Reinforcement Learning 提出基于上下文预测的安全约束以解决非平稳强化学习问题 reinforcement learning
14 ZK-Trace: Certified Collusion Tracing with Zero-Knowledge Credentials for Federated GNSS Interference Monitoring 提出ZK-Trace以解决GNSS干扰监测中的泄露问题 distillation feature matching
15 Miles v0.1: Production-Level Post-Training 提出Miles v0.1以实现前沿后训练的生产级系统 reinforcement learning distillation
16 Revisiting Spectral Representations in Generative Diffusion Models 提出自监督光谱表示对齐方法以提升扩散模型训练效果 representation learning distillation

🔬 支柱九:具身大模型 (Embodied Foundation Models) (7 篇)

#题目一句话要点标签🔗
17 NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting 提出NOAH模型以解决多模态患者数据预测问题 multimodal
18 Suan: Rectifying Direct Preference Safety Alignment in Large Language Models 提出Suan算法以解决大语言模型的安全对齐问题 large language model
19 IPM-FM: A Foundation Model with Consensus Feature Selection for Industrial Process Monitoring 提出IPM-FM以解决工业过程监测中的标签效率低下问题 foundation model
20 Do Reasoning Representations Help Humans Evaluate LLM Outputs? 研究推理表示以改善人类对大型语言模型输出的评估 large language model chain-of-thought
21 Chimaera: A Mixture-of-Graph-Experts Architecture for Cross-Task and Cross-Dataset Graph Learning 提出Chimaera以解决图学习中的跨任务与跨数据集问题 large language model foundation model
22 Training-Free Task Vectors for LLM Behavioral Control 提出无训练任务向量以解决后训练模型编辑的高成本问题 large language model
23 Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling 提出变换器作为上下文采样器以解决数据生成问题 large language model

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
24 SUN: Reaching for Novelty in Reinforcement Learning 提出SUN框架以解决强化学习中的探索与目标选择问题 reachability-aware reinforcement learning

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
25 Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families 提出目标无关微干预以预测语言模型训练响应 PULSE

⬅️ 返回 cs.LG 首页 · 🏠 返回主页