Skip to content

Persistent Structure Meets Dynamic Attention: Cross-Variable Priors for Multivariate Time Series Forecasting

· 0 citations · 17 references

TL;DR

A Params-Per-Pair diagnostic is introduced that predicts from dataset properties alone whether structural priors will help and reveals a horizon-dependent complementarity: the structural prior contributes 33% of the gain at short horizons but 88% at long horizons, confirming that time-invariant knowledge compensates as temporal signal fades.

View source

Similar papers

Open access Sep 2026

Enhancing TimeXer with TCN and differential attention for multivariate stock forecasting

Multivariate stock forecasting requires models that can preserve short-lived price dynamics while using auxiliary market variables without amplifying unstable cross-variable relations. TimeXer provides an Endo/Exo dual-branch architecture that separates target-sequence modeling from auxiliary-variable interaction, but its patch-level tokenization may underrepresent fine-grained local transitions, while standard Exo-branch attention may be sensitive to redundant or weakly informative variables. This paper proposes DAT-TimeXer, a structure-aware adaptation of TimeXer for closing-price forecasting. The model introduces a temporal convolutional network before tokenization to encode causal and dilated local temporal patterns at the original resolution, and applies multi-head differential attention exclusively to the Exo branch. By contrasting two independently learned attention maps, the Exo-side module refines auxiliary-variable dependencies before Endo–Exo interaction while preserving target-sequence dynamics in the Endo branch. Experiments are conducted on three Chinese A-share series and nine U.S.-listed stocks using chronological splits, training-only feature screening, and one-step and multi-step forecasting settings. DAT-TimeXer achieves the lowest mean forecasting errors among the compared models in the one-step evaluations and maintains lower errors at horizons of 1, 3, 5, and 10 in the evaluated multi-step tasks. Cross-asset ablation, chronological subperiod, statistical, and attention analyses provide further evidence for the complementary contributions of pre-tokenization TCN enhancement and Exo-specific differential attention. The added components introduce moderate computational overhead relative to TimeXer.

Si-Xing Liu, Quan-Xiang Lan, Jing Zhang et al. · 0 citations
Conference Open access Sep 2026

ConDyGNet: Constraint-Guided Dynamic Graph Networks for Multivariate Time Series Forecasting

A Constraint-Guided Dynamic Graph Network (ConDyGNet), whose core idea is “global basis, dynamic weights”, which learns a low-rank global basis as a shared structural constraint and generates patch-wise basis mixing weights to construct dynamic propagation graphs.

Zhen-Zhou Li, Xiang Li, Zhibin Niu · 0 citations
Open access Sep 2026

FCA-Transformer: A Feature Pyramid Time Series Forecasting Model Driven by Cross-Attention Mechanism

Multivariate time series forecasting requires modeling both hierarchical temporal dynamics and complex inter-variable dependencies, a dual requirement that often degrades predictive performance and incurs high computational costs in standard Transformer architectures. Unlike current channel-independent models that ignore vital cross-variable synergies, or dense-attention frameworks that suffer from quadratic computational noise, our approach extracts structurally sparse dependencies. To address these specific limitations, this study introduces the FCA-Transformer. The proposed framework integrates a Feature Pyramid Network (FPN) to isolate macroscopic trends from high-frequency localized fluctuations via hierarchical downsampling. Concurrently, a structured Transformer-based Cross-Attention (TCA) mechanism employs Dimensional Segmentation with Weighting (DSW) and a Two-Stage Attention (TSA) layer to map topological variable interactions, effectively extracting robust cross-variable pathways and mitigating distributional noise. Extensive empirical evaluations across three real-world multivariate benchmarks (ETTh1, Electricity, and Exchange Rate) demonstrate that the FCA-Transformer achieves an average reduction of up to 4.39% in MSE and 5.11% in MAE compared to leading baselines. These findings indicate that the proposed architecture successfully reconciles multi-scale feature extraction with lightweight dependency modeling, enhancing structural generalization and providing a scalable framework for real-time temporal analysis in complex industrial environments.

Lin-Li Wu, Ji-Yong Zhang, Zhi-Ming Zhang et al. · 0 citations
Preprint Aug 2026

Multivariate Time Series Forecasting needs Cross Variable Loss

This work proposes CvLoss, a plug-in structural regularizer that constrains forecast residuals on a cross-variable graph and shows that CvLoss consistently improves competitive forecasting models, outperforms representative learning objectives, and is compatible with a variety of forecasting backbones.

Kuiye Ding, Yifan Hu, Hanchen Wang et al. · 0 citations
#machine learning Preprint Sep 2026

CoRe: Coherence and Relational Alignment for Multivariate Time Series Forecasting

Direct forecasting has become a standard paradigm for multivariate time-series forecasting because it predicts the full future horizon in a single pass. However, its training objective is often still decomposed into pointwise errors such as MSE. Such objectives provide stable supervision, but they do not explicitly preserve the structure of the future trajectory: temporal coherence within each variable and relational consistency across variables can both be weakened. We propose CoRe, a model-agnostic learning objective for direct multivariate forecasting. CoRe replaces pointwise supervision with two output-space constraints: a frequency coherence loss that aligns predicted and target spectra, and a low-rank relational graph loss that matches sampled pairwise differences in a target-derived PCA subspace. The resulting objective introduces no trainable parameters and can be applied to existing forecasting backbones by changing only the loss. Experiments on standard benchmarks show that CoRe improves strong baselines, compares favorably with recent forecasting objectives, and remains effective across different backbones, datasets, and hyperparameter settings overall consistently.

Xiao-Yu Lin, Hui-Ran Duan, Yi-Ning Liu et al. · 0 citations
Open access Aug 2026

CASCADED GLOBAL–LOCAL REPRESENTATION LEARNING FOR FINANCIAL TIME-SERIES FORECASTING

The findings indicate that passing attention-derived context into a bidirectional memory module offers a practical means of combining long-horizon structure with local temporal variation, although computational cost remains relevant for latency-sensitive trading applications.

Hao Wu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.