A Brownian-bridge self-normalizer yields asymptotically pivotal confidence regions for prespecified state-value contrasts for prespecified state-value contrasts without estimating the long-run covariance or selecting a bandwidth or batch length.
Abstract
Constant-stepsize temporal-difference (TD) learning is attractive for policy evaluation, but inference from a single Markov trajectory must account for serial dependence and a stepsize-dependent stationary target. For fixed-stepsize linear TD, we establish a functional central limit theorem whose covariance retains the multiplicative component induced by the random TD matrix and the stationary iterate error. We then derive a joint functional limit for parallel Richardson--Romberg (RR) recursions driven by the same trajectory. A Brownian-bridge self-normalizer yields asymptotically pivotal confidence regions for prespecified state-value contrasts without estimating the long-run covariance or selecting a bandwidth or batch length. For such a contrast, the procedure admits a one-pass implementation whose memory does not grow with the trajectory length. At a fixed stepsize, the inferential center is the RR stationary target. We also study horizon-indexed designs in which the stepsize remains constant within each run and decreases across longer horizons. Under an explicit RR-dependent rate window, the residual RR target shift, multiplicative remainder, and initialization effect are negligible at the root-$n$ scale, yielding inference for the projected Bellman solution. Experiments on FrozenLake and Garnet illustrate stationary-target coverage, RR target correction, and the finite-sample behavior of the horizon-indexed design.
This work introduces regularized emphatic TD (RETD), a normalized first-order post-shock repair that leaves the trace and importance ratios unchanged, stores the emphatic TD signal in a leaky scalar state, and releases a delayed correction.
Xing-Guo Chen, Zhao-Hui Wu, Ji Ye et al.· 0 citations
A global high-probability last-iterate guarantee for synchronous tabular QTD under general positive, nonincreasing step-size sequences and arbitrary initialization in the natural parameter range is established.
Zijie Cheng, Xiang Li, Yang Peng et al.· 0 citations
It is shown that, when $\delta\leq1/2$ and $T/\delta$ is sufficiently large, any deterministic anytime boundary valid uniformly over the stated conditionally sub-Gaussian model class must be at least of order at some time by horizon $T$; for nondecreasing boundaries, this order is required at time $T$.
We develop closed- and open-end procedures for monitoring changes in the marginal distribution of object-valued time series. The method combines two distance-based Hilbert-space embeddings, a monitoring-time-dependent projection, and self-normalization. It is computable entirely from pairwise distances, does not requir...
We study online statistical inference for functionals of the return distribution under a fixed policy. The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Markov trajectory. For the Polyak--Ruppert averaged estimator, we prove that its root-$T$ error converges...
We study a data-driven reflection control problem for a Brownian model with unknown drift and volatility. We first propose a learn-then-optimize (LTO) algorithm: it estimates the policy-relevant parameter during exploration, plugs the estimate into the optimality equation, and exploits the resulting policy---achieving...
Guo-Dong Pang, Da-Cheng Yao, Hao Yin· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.