| 20 |
MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models |
提出MM-IFEval-Pro以解决多模态指令跟随评估不足问题 |
reinforcement learning multimodal instruction following |
|
|
| 21 |
Wireless Foundation Models: State-of-the-Art and Open Challenges |
系统分析无线基础模型以解决无线数据表示学习问题 |
representation learning foundation model |
|
|
| 22 |
Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens |
通过贝叶斯视角统一ICL、SFT与KL正则化RL |
RLHF distillation large language model |
|
|
| 23 |
Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach |
提出基于强化学习的太阳能光伏政策设计方法以应对不确定性 |
reinforcement learning PPO SAC |
|
|
| 24 |
What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection |
提出基于困难样本选择的蒸馏训练方法以提升数据效率 |
distillation large language model |
|
|
| 25 |
CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution |
提出CoSkill框架以解决技能演化与策略优化的耦合问题 |
reinforcement learning large language model |
✅ |
|
| 26 |
Predicting Spatiotemporal Mobile Sensing-Based PM2.5 Concentrations Using Low-Rank Adapted Spatially Attentive Graph Neural Network |
提出SA-GNN以解决城市PM2.5浓度预测问题 |
MAE spatiotemporal |
|
|
| 27 |
From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments |
提出合理授权框架以解决代理人工智能的局限性 |
world model world models large language model |
|
|
| 28 |
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation |
提出RISE以解决语言模型蒸馏中的教师质量瓶颈问题 |
distillation |
|
|
| 29 |
GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity |
提出GUT方法以量化和优化大型语言模型的推理不确定性 |
reinforcement learning large language model |
|
|
| 30 |
A Unified Physics-Aware Quantum Machine Learning Framework across Power GaN HEMTs and Logic Nanowire FETs: Predicting Unseen Process Splits and Held-Out Geometry Combinations with Lower Error and Tighter Split-to-Split Variability |
提出统一的物理感知量子机器学习框架以提高器件建模精度 |
reinforcement learning PPO MAE |
|
|
| 31 |
MZ-Rain: Moisture-Budget-Guided Zero-Inflated Model for Station-Level Precipitation Nowcasting |
提出MZ-Rain以解决站级降水短期预报中的物理建模与零膨胀问题 |
predictive model MAE |
|
|