Degradation-Aware Pumping Control of Variable-Speed Pumped Storage via Residual Reinforcement Learning

📄 arXiv: 2607.06911v1 📥 PDF

作者: Kyung-bin Kwon, SangWoo Park, Dam Kim

分类: eess.SY

发布日期: 2026-07-08

备注: 11 pages, 9 figures


💡 一句话要点

提出基于残差强化学习的变速抽水蓄能控制方法以降低设备退化

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)

关键词: 变速抽水蓄能 强化学习 控制系统 设备退化 功率调度 水力损失 效率优化

📋 核心要点

  1. 现有的控制方法在同时满足调度承诺和限制设备退化方面存在矛盾,导致性能下降。
  2. 本文提出的两层控制架构通过分离承诺保障与学习过程,采用前馈-PI控制器和残差强化学习策略。
  3. 实验结果显示,该方法在效率和退化控制上均优于固定速度基线,且在高压调度下表现尤为突出。

📝 摘要(中文)

变速抽水蓄能水电(VS-PSH)在满足短期调度承诺的同时,需限制因调节任务加剧而导致的设备退化。现有方法在同时追求这两个目标时,往往导致退化与性能之间的矛盾。本文提出了一种两层控制架构,将承诺保障与学习限制分开。通过确定性前馈-PI门控控制器确保每五分钟区间的平均功率交付,同时使用残差强化学习策略调整转子速度,确保在固定范围内操作,从而限制最坏情况下的指令。该速度策略跟踪需求依赖的最佳效率点,并针对结合了水力损失与功率变化的操作退化指数进行训练。实验结果表明,该策略在正常和高压调度下,最佳效率点跟踪误差降低约96%,总退化减少约56%。

🔬 方法详解

问题定义:本文旨在解决变速抽水蓄能水电在满足调度承诺的同时,如何有效限制设备退化的问题。现有方法在这两者之间存在矛盾,导致设备性能下降。

核心思路:提出一种两层控制架构,分别使用确定性前馈-PI控制器来保障功率交付,并通过残差强化学习策略来调整转子速度,以避免过度退化。

技术框架:整体架构包括两个主要模块:第一层为前馈-PI控制器,负责在每五分钟内确保功率交付;第二层为残差强化学习策略,负责在固定范围内调整转子速度。

关键创新:最重要的创新在于将承诺保障与学习过程分开,使得控制策略在保证性能的同时,显著降低设备退化。与现有方法相比,这种设计有效避免了性能与退化之间的直接冲突。

关键设计:在设计中,速度策略跟踪需求依赖的最佳效率点,并通过一个结合水力损失与功率变化的操作退化指数进行训练,确保控制的物理可解释性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,所提策略在最佳效率点跟踪误差上相较于固定速度基线降低约96%,在最苛刻的调度条件下,总退化减少约56%。该方法的效率与全信息模型优化器相当或略有超越,同时保持了更紧密的调度跟踪。

🎯 应用场景

该研究可广泛应用于变速抽水蓄能电站的控制系统,具有显著的实际价值。通过降低设备退化,能够延长设备使用寿命,提高整体系统的经济性和稳定性,未来可能对可再生能源的调度和管理产生积极影响。

📄 摘要(原文)

Variable-speed pumped storage hydropower (VS-PSH) must honor short-block dispatch commitments while limiting the operational degradation that intensified regulation duty inflicts on its components. When a single controller pursues both aims at once, every tracking gain is paid for in degradation, a conflict that persists even under full model knowledge and look-ahead. This paper proposes a two-layer control architecture that separates the guaranteed commitment from the bounded learning. A deterministic feedforward-PI gate controller, auditable and certifiable for grid-connected operation, secures average power delivery over each five-minute block, while a residual reinforcement learning policy adjusts only the rotor speed within a fixed bound the gate loop can always absorb, so the worst-case command is bounded by construction. The speed policy tracks a demand-dependent best-efficiency-point reference and is trained against an operation-degradation index that combines off-best-efficiency hydraulic loss with power and actuation variation into one physically interpretable signal. Across normal and stressed dispatch, the proposed policy lowers best-efficiency-point tracking error by roughly 96\% relative to a fixed-speed baseline and cuts total degradation by up to about 56\% under the most demanding dispatch. It matches or slightly exceeds a full-information model-based optimizer in efficiency while preserving substantially tighter block tracking.