We consider the estimation of partial derivatives of multivariate regression-type functionals from incomplete observations generated by a discrete-time strictly stationary ergodic process. The response variable is subject to a missing-at-random (MAR) mechanism, whereas the covariates are fully observed. Building upon the complete-data wavelet methodology developed in Didi and Bouzebda (2025), we construct inverse-probability-weighted empirical wavelet estimators that compensate for the selection bias induced by missing responses. When the propensity score is unknown, a feasible estimator is obtained by replacing the oracle weights with a nonparametric Nadaraya–Watson estimator. The analysis is carried out under stationary ergodicity without imposing mixing assumptions. The estimation error is decomposed into three analytically distinct components: the deterministic multiresolution approximation error, the stochastic fluctuation of the oracle inverse-probability-weighted estimator, and the additional error arising from propensity score estimation. This decomposition makes it possible to isolate the respective effects of approximation, dependence, and missingness within a unified asymptotic framework. Under explicit assumptions on the multiresolution approximation, missingness mechanism, conditional density stabilization, moment conditions, and accuracy of the propensity estimator, we establish non-asymptotic integrated mean squared error bounds together with their asymptotic rates. We further prove almost-sure uniform consistency over compact subsets of the interior of the support and derive a pointwise central limit theorem for both the oracle and feasible estimators. The limiting variance explicitly reflects the information loss induced by inverse probability weighting, and for general orthogonal projection kernels is formulated under the corresponding dyadic-phase condition. The general methodology is specialized to the estimation of first- and second-order derivatives of ordinary regression functions. A finite-sample simulation study investigates the empirical behavior of the proposed estimators under stationary ergodic dependence and MAR missingness, examines the influence of both the wavelet resolution level and the propensity-score bandwidth, evaluates the finite-sample performance of the asymptotic confidence intervals, and compares the proposed procedure with oracle, complete-case, and competing nonparametric estimators. The numerical results are consistent with the theoretical analysis and illustrate the respective contributions of wavelet approximation, inverse probability weighting, and propensity score estimation to the overall estimation error. When the propensity score is identically equal to one, the proposed methodology reduces to the corresponding complete-data wavelet estimator.
This paper develops asymptotic theory for kernel estimation of density-weighted conditional functionals and regression derivatives when responses are missing at random (MAR) and the observations form a strictly stationary ergodic process. Sequential MAR and positivity identify the complete-data conditional target through an inverse-probability-weighted pseudo-response, while the fully observed covariate density and its derivatives are estimated without unnecessary response weighting. A martingale-predictable decomposition yields uniform almost-sure rates, pointwise Gaussian limits, variance expansions, studentization, and AMISE results under explicit projective/maximal, conditional-moment, conditional-density, and variance-stabilization conditions. These quantitative assumptions are additional to stationarity and ergodicity: the results are not asserted for arbitrary stationary ergodic sequences. Exact-quotient and multi-index identities transfer the primitive-estimator theory to regression derivatives, and feasible propensity estimation contributes an explicit additional remainder. Monte Carlo experiments show that stronger dependence, weak response probabilities, higher derivative order, propensity misspecification, and smoothing bias can materially degrade finite-sample performance; undersmoothing improves centring but need not eliminate coverage distortion at moderate sample sizes.
We develop a design-conditional limit theory for kernel estimators of conditional U-functionals based on locally stationary functional random fields observed at irregular random locations and under incomplete response observation. The covariates take values in a separable Hilbert space, the responses are allowed to take values in a general Polish space, and the target is indexed by a class of symmetric kernels of a fixed order. Functional localization is induced by single-index semi-metrics, while spatial localization is performed on the rescaled observation domain. Missing responses are incorporated through a complete-case construction under a Missing At Random condition and a uniform-positivity assumption. The resulting estimator is a ratio of spatially weighted U-statistics with random tuplewise observation indicators. The asymptotic analysis must account simultaneously for four sources of complexity: dependence within the spatial field, nonstationarity across an expanding domain, concentration in an infinite-dimensional covariate space, and the random thinning generated by missing responses. Conditioning on the sampling locations removes the randomness of the spatial design weights but does not eliminate dependence among the observations. We therefore derive a design-conditional projection decomposition adapted to the triangular-array structure of the model. The leading component is represented by a spatially dependent complete-case empirical process, whereas the higher-order canonical terms are controlled uniformly over the response kernels, functional-target points, single-index directions, and rescaled spatial locations. The proofs combine stationary tangent-field approximations for locally stationary random fields, large-block–small-block decompositions, coupling arguments under spatial absolute regularity, small-ball probability estimates, and entropy bounds for the joint indexing class. These arguments yield a uniform stochastic expansion in which the empirical fluctuation, the spatial–functional smoothing bias, and the local-stationarity approximation error appear as distinct contributions. In particular, the local-stationarity remainder has no counterpart in the strictly stationary theory and quantifies the cost of replacing the observed nonstationary field with its stationary tangent approximation. Under the MAR and positivity conditions, complete-case sampling reduces the effective local information and modifies the covariance structure, but it does not change the formal order of the uniform-convergence rate. Under strengthened moment, mixing, entropy, and negligibility conditions, we establish weak convergence of the normalized conditional U-process in the corresponding supremum-norm function space to a tight centered Gaussian process. The limiting covariance is determined by the complete-case first-order projection and consequently retains the effect of the observation propensity and the spatial dependence structure. We also introduce a complete-case leave-tuple-out spatial prediction criterion for bandwidth selection and prove oracle optimality over admissible bandwidth families. The general theory applies to conditional rank association, discrimination probabilities, set-indexed conditional distribution functionals, and related pairwise statistical-learning criteria. Simulation experiments and applications to spatial environmental and epidemiological data illustrate the finite-sample implications of the theory and the stabilizing role of single-index localization. Viewed through the lens of data-driven science, the framework addresses a fundamental asymmetry between the information carried by irregular, locally heterogeneous functional covariates and the selectively observed response tuples. By combining design conditioning, complete-case normalization, tangent-field localization, and single-index dimension reduction, the proposed approach resolves this inferential asymmetry at the level of the model by matching estimation and uncertainty quantification to the information actually available locally, without imposing artificial stationarity or complete-data symmetry.
This paper develops a pointwise distributional theory for linear wavelet density and regression estimation from randomly right-censored observations exhibiting stationary ergodic dependence. In contrast to the prevailing literature, which typically relies on quantitative mixing conditions, our analysis is conducted under ergodicity alone, thereby encompassing substantially broader classes of dependent processes. We establish asymptotic normality for an oracle inverse-probability-weighted estimator based on the true censoring distribution and for its feasible counterpart obtained through Kaplan–Meier substitution. A central result shows that estimating the censoring distribution has no first-order effect on the limiting law, so that the feasible and oracle procedures are asymptotically equivalent. The proof strategy departs from conventional covariance inequalities and blocking arguments and instead combines a martingale-predictable decomposition with martingale central limit theory and ergodic convergence of conditional moments. The framework is further extended to a broad family of wavelet regression functionals involving transformed responses. To render the asymptotic theory directly usable for statistical inference, we introduce a randomly weighted procedure that consistently reproduces the limiting distribution of the feasible estimator. This yields asymptotically valid pointwise confidence intervals without requiring explicit estimation of the unknown asymptotic variance or the introduction of additional smoothing parameters. The scope of the theory includes several important non-mixing and long-range dependent models, while an extensive simulation study demonstrates the finite-sample accuracy and robustness of the proposed inferential methodology.
We establish an asymptotic theory for the Jones inverse-weighted kernel density estimator when length-biased observations form a strictly stationary short-range dependent sequence. The statistical difficulty is intrinsically composite: reciprocal weighting is singular at the origin, the normalizing mean is estimated from the same dependent sample, kernel localization shrinks with the bandwidth, and the centered summands form a row-wise stationary triangular array whose envelope diverges at rate hn−1. Under a non-negative compactly supported Lipschitz kernel, an inverse-moment condition, geometric α-mixing, local regularity of the target density, and uniform local bounds on lagged bivariate densities, we prove strong uniform consistency on compact subsets of (0,∞) and, separately, the uniform stochastic bound OP{hn2+(logn/(nhn))1/2}. A covariance-localization argument shows that the scaled serial-covariance contribution is O{hnlog(1/hn)}=o(1), so the first-order pointwise variance coincides with that of the corresponding independent length-biased estimator. Pointwise and finite-dimensional Gaussian limits are obtained by an explicit big-block/small-block argument with off-diagonal covariance control. The ratio normalization is treated directly: its variance contribution, its product with the localized fluctuation, and its cross-covariance with that fluctuation are all negligible at the nhn scale. We further derive first-order AMSE and AMISE criteria, their oracle bandwidths, and feasible pointwise studentization under undersmoothing. The numerical study separates oracle from data-driven bandwidth selection, evaluates full-ratio HAC and moving-block corrections, examines a Frank-copula Markov robustness design, and benchmarks the Jones estimator against an alternative length-biased estimator. The simulations support the first-order theory while demonstrating that persistent short-range dependence can remain consequential for finite-sample uncertainty.
This paper develops a unified asymptotic theory for inverse-probability-weighted conditional U-statistics of arbitrary fixed order in the presence of missing-at-random responses and infinite-dimensional functional covariates. The target is a conditional higher-order functional generated by a measurable response kernel and evaluated locally on a separable Banach space. Localization is formulated through delta sequences, providing a common framework for kernel, partition, regressogram, orthogonal series, and related smoothing procedures without recourse to finite-dimensional density arguments. For bounded kernels, we establish uniform almost-complete convergence over pseudo-compact functional domains and obtain a sharp decomposition into deterministic localization bias and stochastic fluctuation. The latter is governed by the localized-kernel variance, the envelope of the delta sequence, the metric complexity of the indexing domain, and the small-ball concentration of the functional covariate. Unbounded kernels are treated under explicit weighted moment, truncation, and summability conditions. The feasible theory quantifies the additional perturbation induced by estimating the propensity score and identifies conditions under which this first-stage uncertainty is asymptotically negligible. Pointwise distributional theory is derived through a denominator linearization combined with the Hoeffding decomposition of the centered localized kernel. The Gaussian limit is driven by the first projection, while the higher-order canonical components are shown to be negligible under explicit local-mass, moment, and noncancellation assumptions. This yields oracle-equivalent feasible inference, a consistent first-projection variance estimator, and asymptotically valid studentized confidence intervals. A finite-grid adaptive comparison principle is also developed for data-driven resolution selection. The scope of the theory is illustrated through conditional rank functionals, discrimination with incomplete labels, metric-learning criteria, and functional prediction. Synthetic and semi-synthetic studies based on functional classification, phoneme log-periodograms, and growth trajectories document the finite-sample interaction between covariate-dependent label observation, local information loss, propensity estimation, and inverse-weighting variance.
Salim Bouzebda· Symmetry· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.