Skip to content

One Patch, Three Roles: What Is Actually Coupled in Autoregressive Time-Series Forecasting?

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

The main finding is that a frozen parent's recursive trajectory is easier to fit than the observed future with lightweight parallel exits, and a correctable residual projection along a train-selected periodic history direction is found.

Abstract

Patch-based autoregressive time-series forecasting often ties input representation, learned transitions, and recursive execution to one patch length. We ask which of these roles can be adjusted separately. A supporting atomic-encoding study finds greater sensitivity to model width than to atom grouping on the evaluated grid. Our main finding is that a frozen parent's recursive trajectory is easier to fit than the observed future with lightweight parallel exits. Autoregressive Trajectory Distillation (ATD) turns this into selectable ATD-1/2/4/8 execution, with ATD-1 exactly recovering the parent. On a paired four-data-set comparison, ATD-8 reaches $5.54\times$ end-to-end speedup with stable quality across widths. Fewer calls do not automatically remove the parent's existing forecast error: ATD improves trajectory fidelity in all 21 seed runs but forecast accuracy in only 15 against matched clean-future supervision. We further find a correctable residual projection along a train-selected periodic history direction. Spectrum Tangent applies this correction without adding neural parameters or Transformer calls. At horizon 720, it reduces mean squared error (MSE) and mean absolute error (MAE) by 2.54% and 2.33% over seven data sets and two output widths, while remaining $3.24\times$ faster than recursive inference. Level and shape projections sometimes disagree. Trajectory compressibility, the fidelity-accuracy mismatch, and the correction recur across three public AR parents. Together these results separate representation, transition, and execution as AR design axes. Code is available at https://github.com/RowanFFF/ATD-Spectrum-Tangent.

View source

Similar papers

#machine learning Preprint Sep 2026

Aurora-X: Built for Extreme Time Series Forecasting

A novel pattern-guided mixture-of-experts that expands model capacity through sparse activation and uses shallow patch similarities to constrain deep-layer routing, guiding expert specialization across heterogeneous time series and an implicit quantile network head that predicts arbitrary quantiles to characterize pred...

Xingjian Wu, Chen-Juan Guo, Xiangfei Qiu et al. · 0 citations
#machine learning Preprint Sep 2026

Tabby: An Open Pretraining Recipe for Time Series Foundation Models

Tabby, a long context probabilistic time series foundation model, is released together with a complete and open recipe of how it was built, which achieves competitive zero-shot forecasting performance on GIFT-Eval and the out-of-distribution TIME benchmark.

Shi-Feng Xie, Bahaeddine Abdessalem, Ze-Hao Xiao et al. · 1 citation
#artificial intelligence Preprint Sep 2026

RATL: Learning from Retrieved Residuals for Robust Multivariate Time-Series Forecasting

Overall, RATL shifts the retrieved object from historical target values to base-model-specific historical forecast errors, providing a plug-in, residual-memory-based paradigm for learned feedback correction in continuous-output forecasting.

Yu-Chen He, Yueyang Cang, Zhi-Yuan Ning et al. · 0 citations
Open access Aug 2026

CASCADED GLOBAL–LOCAL REPRESENTATION LEARNING FOR FINANCIAL TIME-SERIES FORECASTING

The findings indicate that passing attention-derived context into a bidirectional memory module offers a practical means of combining long-horizon structure with local temporal variation, although computational cost remains relevant for latency-sensitive trading applications.

Hao Wu · 0 citations
#machine learning Preprint Sep 2026

When, Not How Much: Evaluating Time-Series Foundation Models on Sparse Events

Pretrained time-series foundation models (TSFMs) are evaluated as forecasters of future values, yet for sparse series many decisions depend only on which future periods contain activity. Standard benchmarks do not assess this. On five sparse datasets, we rank positions within forecast windows that contain both events a...

Daniel Schoess, F. von Wangenheim · 0 citations
#artificial intelligence Preprint Sep 2026

DualCast: A Dual-Path Language Model for Bimodal Financial Time-Series Forecasting

Financial time-series forecasting must capture price dynamics across heterogeneous assets while incorporating news available at prediction time. We introduce DualCast, a dual-path framework that extends a frozen language model with a discrete financial vocabulary. Each log-return patch is represented by a learned summa...

Wen-Tao Zhao, Hong-Qiang Wu, Shang-Hang Liu et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.