A four-stage audit for frozen proximal policy optimization policies without retraining examines deployment occupancy, matches current information, tests isolated deviations under incumbent continuation, and evaluates repeated deployment of observation-based alternatives.
Abstract
An optimal reference may recommend trading when a learned policy chooses inaction, but the recommendation depends on information and future decisions. We introduce a four-stage audit for frozen proximal policy optimization policies without retraining. It examines deployment occupancy, matches current information, tests isolated deviations under incumbent continuation, and evaluates repeated deployment of observation-based alternatives. In controlled linear-Gaussian simulations, information matching explains part of the disagreement, while continuation changes its interpretation. At unit observation noise, incumbent continuation reverses 99.3% of projected-hard missed-advantage mass; repeated projected-rule deployment improves all 50 policies. These comparisons distinguish isolated action changes from policy replacement. Historical Bitcoin/Tether (BTCUSDT) replay applies this deployment perspective to a hand-specified intervention selected using 2024 data and frozen for 2025. Daily net reward improves by 135.03 basis points, with gains in 46 of 50 policies, primarily through lower turnover costs. The audit clarifies what oracle-flagged inaction implies for deployed decision making.
Independent evaluation can reject harmful policy updates yet also prevent useful continual learning. We argue that update admission must be assessed through both error control and retained learning opportunities at a stated interaction budget. We identify a concrete failure: a range-based confidence gate cannot certify...
Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equ...
Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must jud...
Sang-Su Lee, V. Loganathan, Shishir Dash et al.· 0 citations
Findings establish recoverability as an outcome-grounded decision variable for selective supervision in OPD and show that retaining teacher-correctable prefixes provides the largest individual contribution.
Deng-Du Jiang, Zhengyang Zhang, Ke-Hong Yuan et al.· 0 citations
Empirical comparisons of counterfactual regret minimization (CFR) variants often mix two distinct choices: the update rule used during training and the iterate reported at evaluation time. This note isolates those choices in an exact small-game setting where exploitability is computed by exact best response, so no me...
Behbod Keshavarzi, H. Navidi· Scientific Reports· 0 citations
Successful agent execution need not identify which future product improvement its user would value. We present a decision-specific audit that maps a declared observation channel and product-value contrast to compatible intervals and witness populations. Its foundations are established identification and decision theory...
Shivam Gupta· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.