Skip to content
Preprint

Counterfactual Optimization of Policy Interventions: Lexical Ordering and Leapfrogging

Aug 2026 · 0 citations · 63 references
Mathematics

Abstract

Most data-driven policy learning methods maximize average outcomes, overlooking the possibility that a policy beneficial on average may still harm a substantial fraction of individuals. Motivated by the ethical principle of"first do no harm", we study how to design a change from a baseline policy that improves overall welfare while keeping the worst-case probability or expectation of individual harm below a specified limit. We establish sufficient conditions under which an optimal policy transition has a lexical leapfrogging structure: groups defined by covariates and current treatment are ranked by a priority score, and any treatment change moves them directly to the conditionally optimal treatment. We derive this score under several models for the dependence among potential outcomes. We demonstrate this harm-aware policy optimization approach in a reanalysis of the I-SPY2 breast cancer platform trial and show how the consideration of counterfactual harm may lead to different conclusions about which treatment-subgroup pairs may warrant deprioritization in further clinical evaluation.

View source

Similar papers

Open access Sep 2026

Offline Policy Evaluation and Learning with Harm Constraints

Offline policy learning aims to optimize individualized decisions using historical data. However, conventional methods primarily focus on maximizing expected rewards while neglecting individual-level counterfactual harm—cases where the assigned treatment leads to worse outcomes than the control. This can result in over...

Jile Chaoge, Qin-Wei Yang, Jing-Yi Li et al. · 0 citations
Review Sep 2026

Orthogonal Policy Learning with Ordinal Outcomes

Policy learning methods based on conditional average treatment effects can obscure subpopulation heterogeneity when applied to ordinal outcomes. We develop a policy learning framework for ordinal outcomes with heterogeneous utilities for individuals who strictly benefit from treatment and those who do not. Since the pr...

Yue Zhang, Shan-Shan Luo, Yan-peng He · 0 citations
Preprint Sep 2026

Risk-Averse Welfare Maximization via Marginal Treatment Effects

This paper studies risk-averse treatment allocation when individuals self-select into treatment based on unobserved characteristics. We develop a framework that combines the marginal treatment effect approach to endogenous selection with a general class of coherent risk measures that capture distributional preferences...

Jarrod Burgh, Emerson Melo · 0 citations
Preprint Aug 2026

A Distribution Mapping Approach to Counterfactually Fair Reinforcement Learning

Reinforcement learning (RL) seeks to optimize sequential decisions to maximize population-level benefits over time. However, when deployed in high-stakes settings such as healthcare, RL decisions might systematically restrict some subpopulation's access to valuable services in a manner contrary to the values and goals...

Jianhan Zhang, Jitao Wang, John D. Piette et al. · 0 citations
#machine learning Preprint Sep 2026

Path-specific harm decomposition: A partial identification framework

A central goal when designing treatment policies is often to"do no harm", that is, to avoid interventions that improve average outcomes while worsening outcomes for some individuals. A widely used notion for harm is the fraction of negatively affected (FNA), defined as the probability that an intervention decreases an...

Rui-Zi Yan, Dennis Frauen, Maresa Schröder et al. · 0 citations
Preprint Aug 2026

Revelation Control

Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equ...

Qin-You Wang · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.