Aug 2026· Mathematical programming· 1 citation· 48 references
TL;DR
The new modeling paradigm subsumes several classical SOC/MDP formulations, including risk-averse and distributionally robust SOC/MDPs as well as partially observed and Bayes-adaptive MDPs, and generates so-called preference robust SOC/MDP models.
Abstract
Inspired by Shapiro et al. [74], we consider a stochastic optimal control (SOC) and Markov decision process (MDP) under simultaneous epistemic and aleatoric uncertainties using Bayesian composite risk (BCR) measures. The proposed BCR-SOC/MDP model evaluates the risk of stagewise cost via a two-layer framework: the inner risk measure tackles aleatoric uncertainty conditional on a latent environment parameter, while the outer risk measure deals with the epistemic uncertainty of the inner risk under the Bayesian posterior. The resulting time-varying risk evaluation induced by Bayesian updating enables an information-adaptive risk-sensitive decision framework. Unlike [74], our policies are allowed to depend explicitly on the posterior belief, reflecting that accumulated information about epistemic uncertainty can influence the assessment of future aleatoric uncertainty and, consequently, the decision maker’s actions [79]. The new modeling paradigm subsumes several classical SOC/MDP formulations, including risk-averse and distributionally robust SOC/MDPs as well as partially observed and Bayes-adaptive MDPs, and generates so-called preference robust SOC/MDP models. Moreover, we derive conditions under which the BCR-SOC/MDP model is well-defined, show that finite-horizon BCR-SOC/MDP models can be solved via dynamic programming, and extend the analysis to the infinite-horizon case. Under standard conditions, we establish asymptotic convergence of the optimal values and optimal policies as data accumulate, and provide quantitative error bounds for several representative classes of risk measures. To enhance computational tractability, we develop a hyper-parameter discretization approach for the posterior belief space. Finally, we carry out numerical tests on a spread betting problem and an inventory control problem, demonstrating the effectiveness of the proposed model and numerical schemes.
Model updating under hybrid uncertainty is challenging because aleatory input variability makes the simulator output a probability distribution rather than a scalar, rendering the likelihood analytically intractable. Existing Approximate Bayesian Computation (ABC) methods typically employ nested Monte Carlo sampling, w...
RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty, and proves a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates.
Bayesian contextual optimization (BCO) is proposed, a framework that maintains a Gibbs posterior over the parameter space that weights candidate models by their empirical decision quality rather than statistical fit, thereby avoiding commitment to a potentially misspecified likelihood while encoding the full distributi...
Zhuo-Jun Xie, Adam F. Abdin, Yi-Ping Fang· 0 citations
Reliable uncertainty estimates are critical in safety-sensitive applications. For such estimates to be useful in practice, it is crucial to understand the sources underlying a model's uncertainty, motivating the disentanglement of total uncertainty into epistemic and aleatoric uncertainty. Existing notions of uncertain...
Frieder Wizgall, Georg Tirpitz, Moritz Seiler et al.· 2 citations
Traditional Bayesian optimal experimental design (OED) selects measurements that best inform a model's parameters. However, such measurements can be suboptimal for downstream predictions. Goal-oriented OED targets the prediction directly. However, the existing goal-oriented criteria value all reductions in predictive u...
J. Jakeman, Rebekah White, B. G. van Bloemen Waanders et al.· 0 citations
Coupled-dynamics environments expose the one-step outcomes that would follow from several possible counterfactual actions under a common realization of exogenous randomness. The ordinary Markov decision process formalism allows one to reason about the marginal law of each action but discards dependence across these cou...
Ege C. Kaya, Aliasghar Pourghani, Mahsa Ghasemi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.