Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equivalent under declared current information can respond differently to future training and favor different actions. The framework defines decision-sufficient revelation and revelation depth, separates pure information value from productive reuse, embeds static Bayes refinement into state-dependent continuation value, and gives an exact cost-adjusted factorization criterion: an additional shallow coordinate is decision-nonredundant only when states sharing a scalar summary lie on opposite sides of the priced Stop/Continue boundary. We also give a target-independent protocol for model-specific instantiation and prove that bounded stop-flip risk alone cannot certify positive expected utility under unrestricted severity. Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes have positive decision value and productive reuse yields strict equal-compute utility advantages. Qwen additionally provides evidence for a decision-nonredundant shallow revealability regime; in Mistral, a scalar continuation architecture fit only on an independent development panel retains positive familywise-adjusted lower bounds on a disjoint target panel, consistent with scalar decision sufficiency within the tested architecture family and resolution. The evidence supports structural rather than numerical transfer: the decision theory, cost accounting, continuation logic, and evaluation protocol transport, while empirical proxies, coefficients, thresholds, and even the required shallow state dimension may be system-specific.
A four-stage audit for frozen proximal policy optimization policies without retraining examines deployment occupancy, matches current information, tests isolated deviations under incumbent continuation, and evaluates repeated deployment of observation-based alternatives.
Xing-Fei Zeng, Xin Zhong, Nan-Ting Li et al.· 0 citations
The original robust value frontier, support-wise linear-programming algorithm, and binary-action fractional-knapsack specialization are embedded into this implementation framework and embeds the original robust value frontier, support-wise linear-programming algorithm, and binary-action fractional-knapsack specializati...
The classic notion of strategyproofness implicitly assumes that a manipulating agent either possesses complete knowledge of what all other agents are going to report, or is willing to take the risk and act as if they know these reports. To capture the profound uncertainty of real-world voters, recent work introduced \e...
Findings establish recoverability as an outcome-grounded decision variable for selective supervision in OPD and show that retaining teacher-correctable prefixes provides the largest individual contribution.
Deng-Du Jiang, Zhengyang Zhang, Ke-Hong Yuan et al.· 0 citations
This study develops a two-period regulatory model in which precautionary capital, information-producing reporting, provider participation, and supervisory architecture are chosen jointly. Reporting may produce a verified signal before continuation capital is set, but it also entails direct, participation, and fixed set...
Edmund Mallinguh· Journal of Risk and Financia...· 0 citations
Most data-driven policy learning methods maximize average outcomes, overlooking the possibility that a policy beneficial on average may still harm a substantial fraction of individuals. Motivated by the ethical principle of"first do no harm", we study how to design a change from a baseline policy that improves overall...
Martina Scauda, Tobias Freidling, Qingyuan Zhao· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.