Skip to content
Preprint

Structured Differentiable Optimization for Efficient Decision-focused Learning in Power Systems

Aug 2026 · 0 citations · 38 references
Engineering Computer Science

TL;DR

This paper presents DiffAPQP, a solver-flexible framework and open-source Python package for scalable DfL with affine-parametric quadratic programs and presents the first solver-based end-to-end DfL demonstration on the IEEE 118-bus system with a 24-hour coupled economic-dispatch and redispatch horizon.

Abstract

Decision-focused learning (DfL) trains forecasting models to align downstream decision consequences, such as power-system operating costs. However, its application to realistic power networks is limited by the need to repeatedly solve and differentiate large optimization problems during training. This paper presents DiffAPQP, a solver-flexible framework and open-source Python package for scalable DfL with affine-parametric quadratic programs. To accelerate the forward pass, DiffAPQP automatically canonicalizes quadratic power-system models written in CVXPY into a differentiation-ready representation and takes advantage of the repetitive solving structure through solver warm-start and solver-data update during training. For the backward pass acceleration, we establish the equivalence between differentiation through the full KKT system and a reduced system obtained by eliminating inactive inequality constraints. For training losses depending solely on the optimal value, we further derive an envelope-theorem-based gradient that avoids solving an adjoint KKT system, resulting in eligible backward time. To our knowledge, this work presents the first solver-based end-to-end DfL demonstration on the IEEE 118-bus system with a 24-hour coupled economic-dispatch and redispatch horizon. Under matched SCS and Clarabel backends on a Linux machine, DiffAPQP achieves $2.27\times$--$3.58\times$ closed-loop and $3.62\times$--$4.38\times$ counterfactual end-to-end DfL training speedups over CvxpyLayers. The best solver configurations increase these speedups to $3.91\times$ (from 38.65 to 9.55 min/epoch) and $6.39\times$ (from 10.73 to 1.68 min/epoch), respectively. Additionally, DiffAPQP reduces peak memory usage by approximately $50\%$, while keeping similar operating costs as CvxpyLayers.

View source

Similar papers

Preprint Sep 2026

PowerModels-ACOPF-AI: On-the-Fly Machine Learning Approach for Solving AC Optimal Power Flow Integrating Renewable Energy Sources

The increasing complexity of modern power systems, driven by high renewable penetration, load variability, and operational uncertainty, demands fast and reliable solutions to the AC optimal power flow problem (AC-OPF). Traditional optimization methods, though accurate, often struggle with scalability and high computati...

Bhuban Dhamala, J. Tabarez, Anup Pandey · 0 citations
#machine learning Preprint Sep 2026

Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation

Mean-variance portfolio optimization (MVO) is a central framework in data-driven asset management. A widely adopted approach is a two-stage framework that first predicts expected returns and then solves the optimization problem based on these predictions, with the predictive models trained by minimizing prediction erro...

Kensei Nosaka, Shunnosuke Ikeda, Yuichi Takano · 0 citations
Conference

Integrating Machine Learning with Multi-Horizon Stochastic Programming for Scalable Capacity Expansion

We introduce a novel approach to tackling large-scale capacity expansion problems in power systems by integrating machine learning techniques with multi-horizon stochastic programming that combines long-term strategic decision-making under uncertainty with short-term operational decisions and uncertainties. Our propose...

Taehyeon Kwon, A. Subramanyam · 0 citations
Conference Sep 2026

Deep reinforcement learning-enhanced framework for precision investment decision and dynamic optimization in power grid systems

A Deep Reinforcement Learning–Enhanced Dynamic Optimization Framework (DRL-DOF) that integrates uncertainty-aware policy optimization, temporal-attention actor–critic networks, and digital twin simulation for precision investment management is presented.

Yu-Hao Zhou · 0 citations
Preprint Aug 2026

Integrated Learning and Robust Optimization

This work proposes an integrated learning and robust optimization (ILRO) framework, where a robust decision problem is used both to define the training problem (termed the RSPO loss problem), and to produce the deployed decision, which achieves both robustness and learning-decision alignment.

Chengpeng Tan, Yu-Chen Mao, Shu-Ming Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.