DTFNet: Dual-stage LSTM-Transformer fusion architecture for time series prediction
Abstract
Time series prediction plays a vital role across numerous real-world domains. However, existing methods struggle to simultaneously capture short-term temporal dynamics and long-term dependencies. To overcome this limitation, this paper proposes a Dual-stage Temporal Fusion Network, termed DTFNet, which realizes multi-scale timing modeling by integrating the architectural advantages of LSTM and Transformer. The model first employs LSTM with its inherent gating mechanism to extract local dynamic features from the sequence. Then, a lightweight cross-stage feature calibration module (CFC) is introduced to adaptively refine and recalibrate the feature distribution, making it more suitable for the input requirements of the Transformer. Finally, the hierarchical multi-headed self-attention module (HMSA) is used to efficiently capture the global contextual dependencies of the sequence in a windowed manner. Experimental results on the four benchmark datasets (ETTh1, ETTm1, Weather and ECL) show that DTFNet reduces the mean square error (MSE) by 2.2% to 3.7% compared with the state-of-the-art methods. Especially in the short sequence prediction task, compared to TimeMixer, the MSE is reduced by 3.8%, which demonstrates the effectiveness of the proposed dual-stage architecture in balancing short-term and long-term forecasting capabilities.