cs.LG(2026-07-09)

📊 共 24 篇论文 | 🔗 2 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (12 🔗1) 支柱二:RL算法与架构 (RL & Architecture) (9) 支柱一:机器人控制 (Robot Control) (2 🔗1) 支柱七:动作重定向 (Motion Retargeting) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (12 篇)

#题目一句话要点标签🔗
1 Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models 提出预算感知的测试时间模型选择以优化大语言模型的响应质量 large language model
2 Eigenvalue Calibration for Semantic Embeddings of Large Language Models 提出新框架以校准大型语言模型的语义嵌入特征 large language model
3 Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization 提出系统感知的KV缓存优化以提升大语言模型服务效率 large language model
4 Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment 提出小型视觉语言模型量化的新框架以优化边缘部署 large language model multimodal
5 Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA 提出自验证LLM风险分析工具以解决安全分析盲点问题 large language model
6 What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents 提出统一的内存压缩框架以优化大型语言模型的上下文管理 large language model
7 Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models 提出敏感度感知阈值和令牌路由以优化大语言模型的激活稀疏化 large language model
8 Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models 提出预算感知的测试时模型选择方法以优化大语言模型的响应质量 large language model
9 NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision 提出NL-PAC框架以解决LLM监督中的规范模糊性问题 large language model
10 TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning 提出TSRouter以解决时间序列推理中的动态模态选择问题 large language model
11 BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving 提出BlockServe以解决扩散大语言模型服务中的收敛异质性问题 large language model
12 Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal 提出激活引导的对抗后缀优化以增强模型安全性 large language model

🔬 支柱二:RL算法与架构 (RL & Architecture) (9 篇)

#题目一句话要点标签🔗
13 Write-Protected Discrete Bottlenecks for Language-Grounded World Models: A Structural Limitation and Sufficient Fix 提出写保护离散瓶颈以解决语言与世界模型交互问题 world model world models JEPA
14 Joint Discrete-Continuous Flow Matching for Open-Vocabulary Inverse Design of Multilayer Optical Coatings 提出IrisFlow框架以解决多层光学涂层的开放词汇逆向设计问题 flow matching open-vocabulary open vocabulary
15 BiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model Compression 提出BiSCo-LLM以解决极低比特大语言模型压缩问题 distillation large language model
16 MatBind: A Shared Embedding Space for Multimodal Materials Characterization 提出MatBind框架以解决多模态材料表征问题 contrastive learning multimodal
17 MPFlow: Learning Budgeted Max-Flow Optimization on the Lightning Network with Deep Graph Reinforcement Learning 提出MPFlow以解决比特币闪电网络的流动性配置问题 reinforcement learning PPO
18 Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning 提出量化分布强化学习以提高统计效率与推断能力 reinforcement learning
19 Self-Adaptive Anomaly Detection with Reinforcement Learning and Human Feedback in Connected Vehicles 提出自适应异常检测框架以解决联网车辆监控问题 reinforcement learning
20 Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning 提出多模态、多环境机器教学以解决鲁棒奖励学习问题 reinforcement learning inverse reinforcement learning
21 Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs 提出视觉政策检查方法以优化强化学习中的开放式课程设计 reinforcement learning VIP

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
22 Prompt-Driven Exploration 提出基于提示驱动的探索方法以解决强化学习中的稀疏奖励问题 manipulation vision-language-action VLA
23 SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions 提出SafeExplorer以解决强化学习中的偏差问题 Unitree reinforcement learning PPO

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
24 Reinforcing the Generation Order of Multimodal Masked Diffusion Models 提出可学习控制模块以优化多模态生成顺序 spatial relationship multimodal

⬅️ 返回 cs.LG 首页 · 🏠 返回主页