Back

G-2026-37

Managing short-term hydropower production with a multi-agent reinforcement learning model

, , , and

BibTeX reference

Hydropower generation fulfills 94% of Québec’s electricity needs, making efficient Short-Term Hydropower Scheduling (STHS) critical for daily operations. This paper addresses the STHS problem using a multi-agent Reinforcement Learning (RL) model based on Proximal Policy Optimization (PPO). Multiple agents are trained to select hourly discharge actions defined by efficiency points, pairs of discharge and power output, while a reward function integrates reservoir management. A digital-twin of the hydropower system, based on twelve years of hourly operational data, serves as the environment. Model performance is evaluated over a 96-hour prediction horizon and validated out-of-sample against a Mixed-Integer Linear Programming (MILP) model. Results are analyzed through reservoir volume trajectories and energy production profiles, demonstrating the model’s ability to generate operationally coherent schedules and highlighting the potential of advanced reinforcement learning approaches for improved hydropower decision-making.

, 9 pages

Research Axis

Research application

Document

G2637.pdf (2 MB)