Aug 2026· 1 citation· ⚡ 1 influential· 22 references
MathematicsComputer Science
TL;DR
This work decomposes label variance into an $O(1)$ trait component and an $O(T^{-1})$ state component, explaining why a snapshot can retain cross-sectional predictability while poorly tracking within-person change, and derives task-dependent effective temporal spans.
Abstract
Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows. Their apparent performance ceiling may therefore be an acquisition-protocol ceiling rather than a model-capacity ceiling. We study labels of the form $\Theta_{g,T}=T^{-1}\int_0^T g\{Z(t)\}\,\mathrm{d}t$ when the latent Gaussian process contains both a stable individual trait and a correlated within-individual state. An exact protocol-conditioned Bayes-risk identity provides a common tool. First, we decompose label variance into an $O(1)$ trait component and an $O(T^{-1})$ state component, explaining why a snapshot can retain cross-sectional predictability while poorly tracking within-person change. Second, we derive task-dependent effective temporal spans: mean labels depend on the ordinary correlation time, whereas occupation-time labels depend on an entire spectrum of higher-order correlation times. Third, state-driven occupation-label variance is maximal when the stable trait lies at the threshold; window efficiency decays much more slowly away from that boundary. Under an equal segment budget, exact risks and Monte Carlo experiments show that repeated segments at one time rapidly saturate, whereas temporally dispersed observations continue to increase state explainability. The trait ceiling uses quantities available from ordinary test-retest data; only the state ceiling requires short-lag temporal calibration. The results distinguish architectural limits from protocol limits and show that the label, rather than duration or segment count alone, defines the relevant timescale.
The \emph{conditional forecast-revision scale} $\It=\{\Var(\E[X_{t+1}\mid\F_t]\mid\F_{t-1})\}^{1/2}$ measures the history-specific size of the forecast update induced by observing $X_t$. Because it is a conditional second moment built from two unknown conditional means, it is not directly observed. We study which estim...
This work proves that on generated data, the state channel cuts regret by up to 79 percent against action-level weighting and, on the funnel family, by 32 to 68 percent against a tuned minimax-optimal method.
Latent trait models differ less in what they measure than in what they fix. Item response theory frees the trait from the items administered but, in doing so, surrenders the origin and unit of the scale to convention. The Cognitive Trait Model (CTM; Choi, 2022) restores both by bounding the trait on [0, 1], where 0 den...
A (computationally inefficient) adaptive estimator that, so long as $p$ is a mixture of $k$ symmetric log-concave densities, achieves error comparable with the optimal estimator that knows $p$ and has $\tilde\Theta(n/k)$ samples.
This work composition the observables'likelihoods in per-task free-routed last-layer beliefs on a shared backbone absorbs unit-dependent loss scaling into likelihood parameters learned in the same gradient pass, and results land where theory puts them.
This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models, together with numerically stable positive-decay renormalization, to remain competitive in language modeling and improve length extrapolation on variable-digit addition.
Yi-Fan Zhang, Steve Ta, Jasper Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.