Unit-Commitment-Guided Online Learning for Day-Ahead Electricity Market Bidding
Abstract
This paper proposes a unit-commitment-guided multi-armed bandit method for learning bidding strategies of a generation company participating in a day-ahead electricity market under generator constraints, award uncertainty, and imbalance settlement. The proposed method separates pre-market unit commitment (pre-UC), which extracts operational structure under predicted prices before bidding, from post-market unit commitment (post-UC), which evaluates realized profit after market clearing using realized prices, awarded quantities, and imbalance settlement. By injecting the bid quantity and marginal cost suggested by pre-UC into the exploration score, the proposed method guides ordinary multi-armed bandit (MAB) exploration toward operationally plausible price-quantity regions. An imbalance-aware safety penalty is also introduced to discourage repeated selection of arms that have historically led to large imbalances. The method is therefore positioned as a context-guided MAB rather than a conventional contextual bandit that directly learns a parametric relationship between context and reward. Numerical experiments using Japan Electric Power Exchange (JEPX) spot market data over 10,000 episodes show that the UC-derived context reduces regret and imbalance energy for upper confidence bound (UCB), $\epsilon $ -greedy, and Thompson sampling. In particular, for Thompson sampling, the final average regret decreases from approximately 48,433 kJPY/day without context to 1,688 kJPY/day with context and safety. Additional sensitivity analyses on context weights, safety thresholds, imbalance settlement coefficients, reward scaling, price forecasting, action-space discretization, oracle definitions, and UC computation time clarify the effectiveness, robustness, and limitations of the proposed framework.