Interpretability analysis shows that spectral attention patterns are associated with financial characteristics such as momentum, volatility, and liquidity that are associated with financial characteristics such as momentum, volatility, and liquidity in stock forecasting.
Abstract
Forecasting stock prices is challenging due to the non-stationarity and volatility of financial time series. We propose the Generative Adaptive Decomposition Hierarchical Transformer (GADHT), a hybrid framework that combines adaptive decomposition, masked self-supervised pretraining, and hierarchical attention. GADHT applies Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN) to decompose financial signals into intrinsic mode functions (IMFs) that capture multi-scale temporal dynamics. A masked self-supervised pretraining task based on IMF reconstruction is used to learn spectral–temporal representations without labeled data, while a hierarchical transformer with energy-weighted attention emphasizes informative IMFs during forecasting. Experiments on large-cap equities across multiple forecasting horizons show that GADHT achieves competitive and stable forecasting performance. The model maintains stable predictive behavior during stress periods such as the 2008 financial crisis and the 2020 COVID-19 crash, and shows positive economic performance under the adopted backtesting assumptions. Zero-shot and cross-market experiments further suggest that the learned representations can transfer to unseen large-cap equities and selected international markets. Interpretability analysis shows that spectral attention patterns are associated with financial characteristics such as momentum, volatility, and liquidity. Overall, the results suggest that GADHT provides a coherent and interpretable framework for multi-horizon stock forecasting.
A CNN–Transformer dual-channel architecture equipped with a dynamic attention fusion module for stock price forecasting that reduces mean absolute error and root mean square error and remains effective across markets with differing volatility profiles is introduced.
Experiments on eight representative Chinese A-share and U.S. stock datasets show that FI-CAECNet achieves the best overall forecasting performance among all compared models and reduces the average MSE and MAE and improves the hit ratio.
ResDIF, a Residual Disentanglement framework for Interpretable financial time series Forecasting inspired by Asset Pricing Theory, is proposed and demonstrates that ResDIF outperforms existing methods while providing actionable interpretability support for refined portfolio risk management.
Chengwei Fu, Gang Xiao, Yu-Chao Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
The findings indicate that passing attention-derived context into a bidirectional memory module offers a practical means of combining long-horizon structure with local temporal variation, although computational cost remains relevant for latency-sensitive trading applications.
The findings indicate that the proposed architecture successfully reconciles multi-scale feature extraction with lightweight dependency modeling, enhancing structural generalization and providing a scalable framework for real-time temporal analysis in complex industrial environments.
The revised evidence supports lower price-level errors, while directional and significance results are mixed across markets, and the findings establish cross-market consistency rather than transfer learning.