cs.LG(2026-07-20)

📊 共 71 篇论文 | 🔗 1 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (37) 支柱九:具身大模型 (Embodied Foundation Models) (28 🔗1) 支柱八:物理动画 (Physics-based Animation) (3) 支柱一:机器人控制 (Robot Control) (2) 支柱三:空间感知与语义 (Perception & Semantics) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (37 篇)

#题目一句话要点标签🔗
1 LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models 提出LLM4EHR以解决临床时间序列与医疗事件序列对齐问题 representation learning large language model foundation model
2 SODA: Semi On-Policy Black-Box Distillation for Large Language Models 提出SODA以解决大语言模型蒸馏中的效率与稳定性问题 distillation large language model
3 Dichotomous Diffusion Policy Optimization 提出DIPOLE以解决扩散策略优化中的不稳定性问题 reinforcement learning diffusion policy vision-language-action
4 CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward Approach 提出情感增强深度强化学习以解决主动交易问题 reinforcement learning deep reinforcement learning PPO
5 Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems 提出PEARL以解决高维动态系统的实时最优控制问题 reinforcement learning policy learning sparse sensors
6 ADS-C: Antidistillation Sampling for Classification 提出ADS-C以解决分类任务中的蒸馏攻击问题 distillation large language model
7 QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides 提出QUADS以解决MoE强化学习中的NVFP4不稳定问题 reinforcement learning large language model
8 SC-JEPA: Stabilizing Latent Predictive Learning for Time-Series Anomaly Prediction 提出SC-JEPA以解决时间序列异常预测中的不稳定性问题 JEPA predictive model distillation
9 The Terminal Representation in Reinforcement Learning 提出终端表示法以改进强化学习中的状态表示 reinforcement learning representation learning reward shaping
10 DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning 提出DADiff以解决强化学习中的跨域策略适应问题 reinforcement learning representation learning
11 A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model 提出经典与量子算法以优化强化学习中的生成模型 reinforcement learning offline reinforcement learning
12 Robust Peak-cost Constrained Reinforcement Learning 提出稳健的峰值成本约束强化学习以解决安全关键问题 reinforcement learning
13 On the Failure of Boundary-Seeking Distillation in Bottlenecked Generative Architectures 提出边界寻求蒸馏方法以解决瓶颈生成架构的知识转移问题 distillation
14 When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis 提出模型合并方法以挑战联合多任务强化学习 reinforcement learning
15 Understanding Reasoning from Pretraining to Post-Training 提出棋类作为控制测试平台以研究预训练与后训练的推理关系 reinforcement learning large language model
16 When Does Muon Help Agentic Reinforcement Learning? 提出Muon优化算法以提升稀疏奖励强化学习性能 reinforcement learning
17 Prediction-Only Distillation in Linear and Logistic Regression 提出预测仅蒸馏方法以解决训练数据缺乏问题 distillation
18 Learning Standard Model structure from LHC data with Riemannian flow matching 提出ShellFlow模型以从LHC数据中学习标准模型结构 flow matching
19 CTC: The Composite Task Challenge for Cooperative Multi-Agent Reinforcement Learning 提出复合任务挑战以解决合作多智能体强化学习中的分工不足问题 reinforcement learning
20 DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone 提出DiffuMamba以提升扩散语言模型的推理效率 Mamba
21 Soft $Q(λ)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces 提出Soft $Q(λ)$以解决多步离策略熵正则化强化学习问题 reinforcement learning
22 Beyond Objective Expressivity: Geometry Preservation in Multimodal Contrastive Learning 提出几何保持编码器以解决多模态对比学习中的优化问题 contrastive learning multimodal
23 FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications 提出FlashRT以优化多模态应用的实时部署 world model world models multimodal
24 A Geometric Perspective on Stabilizing Value Conflict Resolution 提出链式思维以解决大语言模型的价值冲突问题 reinforcement learning RLHF large language model
25 MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models 提出MADA-RL以解决紧凑模型推理效率低的问题 reinforcement learning large language model
26 Information-Based Exploration via Random Features for Reinforcement Learning 提出随机特征信息增益方法以优化强化学习探索策略 reinforcement learning deep reinforcement learning representation learning
27 AGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models 提出JAGG以解决扩散模型GRPO训练的计算瓶颈问题 reinforcement learning flow matching large language model
28 OR Else: A Differentiable Trust Region for Policy Optimization 提出可微信任区域方法以优化策略训练 PPO RLHF large language model
29 Adaptive Mamba Neural Operators 提出自适应Mamba神经算子以解决偏微分方程问题 Mamba SSM
30 Planning with Transformers: Chain of Computation and Structured Context Windows 提出链式计算架构以解决规划问题 world model world models large language model
31 Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal Inference 提出基于RAG的策略学习方法以解决因果推断中的邻近匹配问题 policy learning
32 Enhancing Rubric-based RL via Self-Distillation 提出CriPO以解决Rubric-based RL中的探索不足问题 distillation
33 Theoretical Foundations of $\max$@$k$ Reinforcement Learning 提出$ ext{max}@k$强化学习理论基础以解决评估问题 reinforcement learning
34 A Weisfeiler-Leman Characterization of Global-Attention Graph Transformers for Mixed-Integer Linear Programs 提出图基础模型以解决混合整数线性规划的表达能力问题 linear attention foundation model
35 LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks 提出体验学习方法以提升非可验证任务的学习效果 reinforcement learning distillation
36 Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization 提出基于优势聚合的多智能体策略优化方法 reinforcement learning PPO
37 Distributional Soft Bellman Operator under the Cramér Geometry 提出基于Cramér几何的分布式软贝尔曼算子以优化策略评估 reinforcement learning DRL

🔬 支柱九:具身大模型 (Embodied Foundation Models) (28 篇)

#题目一句话要点标签🔗
38 Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework 提出FedGAMMA以解决多模态图的联邦学习问题 foundation model multimodal
39 DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model Embeddings 提出DELUGE以解决大陆尺度日常降雨洪水损失预测问题 foundation model multimodal
40 AI Trading: Evaluating Large Language Models for Technical Market Analysis 评估大型语言模型在技术市场分析中的应用潜力 large language model
41 Revisiting data-driven dynamic security assessment with a tabular foundation model 提出基于表格基础模型的动态安全评估方法以解决数据需求问题 foundation model
42 MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model 提出MxGPS以解决电网模型的拓扑过拟合问题 foundation model
43 PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning 提出PASs-MoE以解决持续学习中的误对齐漂移问题 large language model multimodal
44 LLM-Guided Transportation Hub Capacity Planning with Textual Business Inputs 提出LLM引导的运输枢纽容量规划以解决定性输入不足问题 large language model chain-of-thought
45 Diffusion models recover accurate mixture weights despite score function insensitivity 提出扩散模型敏感性框架以准确恢复混合权重 multimodal
46 Recursive Harness Self-Improvement 提出递归式工具自我改进以优化模型训练质量 foundation model
47 Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data 提出时间感知评估协议以解决合成序列表格数据的时效性问题 TAMP
48 A Benchmark for Electrical Load Forecasting Across Grid Levels: Time-Series Transformers Outperform Established Methods 提出电力负荷预测基准,Transformer方法显著提升预测精度 foundation model
49 ContinuityBench: A Benchmark and Systems Study of Stateful Failover in Multi-Provider LLM Routing 提出ContinuityBench以解决多提供者LLM路由中的状态故障转移问题 large language model
50 PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization 提出PagedWeight以解决MoE LLM服务中的动态权重量化问题 large language model
51 How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation 提出动态预算分配方法以优化多轮LLM评估 large language model
52 Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values 揭示语言模型的隐性价值泄漏问题及其影响 chain-of-thought
53 Pretrained Event Classification Model for High Energy Physics Analysis 提出基于图神经网络的事件分类模型以解决高能物理分析问题 foundation model
54 Towards Reliable Zero-Shot Crowd Forecasting: Evaluating Time Series Foundation Models for Special Event Pedestrian Forecasting 提出基于预训练时间序列模型的零-shot人群预测方法以解决特殊事件中的人流管理问题 foundation model
55 Lightweight Wrappers for Adapting Time Series Foundation Models to Regional Drought Forecasting 提出轻量级适配框架以解决区域干旱预测问题 foundation model
56 Residual-Guided Multi-Resolution Refinement of Foundation Models: A Case Study in Drought Forecasting 提出RGMR框架以改进气候预测中的时间序列模型 foundation model
57 SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs 提出SelectInfer以解决边缘设备上LLM推理效率问题 large language model
58 Harness Engineering for LLM-Driven GPU Kernel Generation 提出基于测试框架的LLM驱动GPU内核生成优化方法 large language model
59 Topological Signatures of Context-Level Reliability in TabPFN 利用拓扑特征分析TabPFN的上下文级可靠性 foundation model
60 AutoEncoder-Compressed Parallel Split Learning for Pre-trained Model Fine-Tuning 提出AE-PSL以解决边缘设备上大规模模型微调的通信效率问题 foundation model
61 FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches 提出WC2026-Agents基准以评估LLM在未来事件预测中的表现 large language model
62 PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer 提出PoLoRA优化器以提升LoRA模型微调效率 large language model
63 CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning 提出CoCurve以解决结构化LLM剪枝中的依赖性问题 large language model
64 Program Synthesis for Simulation-Based Inference: Joint Model Selection and Parameter Estimation 提出联合模型选择与参数估计的新框架以解决复杂模型推断问题 large language model
65 FailureAtlas: A Taxonomy of Failure Modes in Multi-Provider LLM Serving Infrastructure 提出FailureAtlas以分类多提供者LLM服务基础设施中的故障模式 foundation model

🔬 支柱八:物理动画 (Physics-based Animation) (3 篇)

#题目一句话要点标签🔗
66 Physics-Based Deep Spatiotemporal Hyperlocal Radar Nowcasting with a Multi-Variable U-Net for High-Resolution Precipitation Forecasting 提出基于物理的深度时空超局部雷达短期预报框架以解决降水预测问题 spatiotemporal
67 Boosted Enhanced Quantile Regression Neural Networks with Spatiotemporal Permutation Entropy for Complex System Prognostics 提出集成框架以解决复杂系统的长时段故障预测问题 spatiotemporal
68 Lightweight CNN-Based Anomaly Detection for High Voltage Converter Modulators in the Spallation Neutron Source 提出轻量级CNN模型以解决高压转换器调制器异常检测问题 PULSE

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
69 Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning 提出MPG方法以解决少样本逆强化学习问题 manipulation reinforcement learning inverse reinforcement learning
70 A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control: A Case Study in Additive Manufacturing 提出自适应数字双胞胎框架以解决模型退化问题 model predictive control

🔬 支柱三:空间感知与语义 (Perception & Semantics) (1 篇)

#题目一句话要点标签🔗
71 Rethinking the Global Knowledge of CLIP in Training-Free Open-Vocabulary Semantic Segmentation 提出GCLIP以解决CLIP在无训练开放词汇语义分割中的局限性 open-vocabulary open vocabulary

⬅️ 返回 cs.LG 首页 · 🏠 返回主页