Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· 0 citations· 29 references
TL;DR
This work defines cross-modal attention as a selection process, where semantic prompts are employed to induce a sparse temporal attribution distribution over temporal positions and employs low-entropy regularization alongside global cross-modal consistency constraints to regulate the selection behavior.
Abstract
Recent advances in time series forecasting (TSF) leverage large language models (LLMs) to provide semantic priors, enabling more robust forecasting under limited training data. However, directly fusing semantic signals into temporal features often induces modality entanglement and obscures local temporal structures when cross-modal correlations are weak, noisy, or inconsistent. To address these issues, we propose Attention as Selection (AAS), a dual-branch framework that decouples temporal and semantic representations. Specifically, we define cross-modal attention as a selection process, where semantic prompts are employed to induce a sparse temporal attribution distribution over temporal positions. This guides the model to focus on critical time steps without interfering with the construction of temporal representations. Furthermore, low-entropy regularization is employed alongside global cross-modal consistency constraints to regulate the selection behavior, ensuring that semantic guidance remains sparse, stable, and aligned with temporal dynamics. Extensive experiments on six real-world datasets demonstrate that AAS outperforms existing methods across various forecasting scenarios. Code is available at https://github.com/VIMLab-hfut/Attention-as-Selection.
This work proposes SAGE (Seeing and Augmenting with Grounded Encoding), an end-to-end CLIP-based framework that jointly models temporal, cross-variable, textual, and visual information and achieves state-of-the-art accuracy.
Time series forecasting (TSF) is critical in numerous real-world applications, yet its sequential scalar presentation limits semantic richness and the capture of complex temporal patterns. Recent advances leveraging patchwise modeling and pretrained large language models (LLMs) have achieved notable progress. However, existing methods largely focus on raw sequential patterns while overlooking intrapatch semantics, which limits the richness of time series representations, and they fail to effectively exploit complementary information from multiple views. To tackle these challenges, we propose a self-driven multiview architecture for TSF (SMArT). SMArT enriches intrapatch semantics through self-supervised multiview fusion. It jointly captures fine-grained temporal 1-D dependencies via pointwise self-attention and global temporal 2-D structures via Gramian angular fields (GAFs), all without requiring external supervision. To further bridge the gap between time series data and LLMs, SMArT pioneers a dual-prompt strategy, combining static, dataset-level guidance with dynamic, input-specific prompts derived from frequency-domain decomposition, significantly enhancing LLMs' adaptability and generalization. Extensive experiments across diverse TSF tasks validate SMArT's robustness, achieving state-of-the-art performance in long-term forecasting and excelling in few-shot and zero-shot scenarios. By integrating self-driven multiview learning with LLMs' reasoning power, SMArT establishes an effective framework for TSF. Our code and Supplementary Materials are available at https://github.com/BMRETURN/SMArT.
Wen-Bin Xing, Meng-Ran Li, Bo-Yu Zhang et al.· IEEE Transactions on Neural...· 0 citations
The extreme sparsity, zero inflation, and weak temporal regularity inherent to intermittent demand forecasting make it difficult to use traditional deep learning models. This paper presents a novel framework called SemTFT, a conditional semantic-temporal fusion transformer, which incorporates semantic information in a controlled, data-driven fashion. In contrast to the previous methods, where embeddings are added randomly, our method is based on a framework of semantic consistency that measures how the similarity of embeddings is related to demand behavior. Semantic features are then selectively activated by a gating mechanism when informative, effectively suppressing semantic noise. To explicitly deal with zero-inflated demand distributions, the model is trained with a Tweedie loss function. This experiment on real-world e-commerce data shows that naive semantic integration may actually decrease the performance, whereas the proposed conditional approach enhances forecasting accuracy, attaining a decrease in MASE and RMSSE compared to the baseline and standard TFT models. Noise and covariate shift robustness tests also substantiate the fact that our approach provides stable performance, emphasizing the role of selective semantic use in intermittent demand forecasting.
Halimah Alkhorasani, M. Alsamhi, A. V. Shvetsov· 2026 6th International Confe...· 0 citations
Direct forecasting has become a standard paradigm for multivariate time-series forecasting because it predicts the full future horizon in a single pass. However, its training objective is often still decomposed into pointwise errors such as MSE. Such objectives provide stable supervision, but they do not explicitly preserve the structure of the future trajectory: temporal coherence within each variable and relational consistency across variables can both be weakened. We propose CoRe, a model-agnostic learning objective for direct multivariate forecasting. CoRe replaces pointwise supervision with two output-space constraints: a frequency coherence loss that aligns predicted and target spectra, and a low-rank relational graph loss that matches sampled pairwise differences in a target-derived PCA subspace. The resulting objective introduces no trainable parameters and can be applied to existing forecasting backbones by changing only the loss. Experiments on standard benchmarks show that CoRe improves strong baselines, compares favorably with recent forecasting objectives, and remains effective across different backbones, datasets, and hyperparameter settings overall consistently.
Xiao-Yu Lin, Hui-Ran Duan, Yi-Ning Liu et al.· 0 citations
Probabilistic forecasting models are widely used for time series forecasting in domains such as energy systems, finance, medicine, and transportation. In recent years, deep generative models have shown strong results on probabilistic forecasting, yet many conventional approaches struggle to capture internal temporal dependencies, leading to latent representations with limited expressive power. To address this limitation, we propose \textit{CLaST}, a VAE framework for probabilistic multivariate time series forecasting. Unlike existing generative models, CLaST learns embeddings that preserve contextual similarity between observations through our contrastive loss function. Experiments across nine widely adopted benchmarks demonstrate that CLaST consistently surpasses strong baseline methods. In short-term forecasting tasks, our approach achieves improvements of up to $16.4\%$ in CRPS and $14.4\%$ in NMAE over the second-best method. Furthermore, in long-term prediction CLaST attains superior overall performance, exceeding the second-best method by up to $48.6\%$ and $25.1\%$ in CRPS and NMAE, respectively.
A. Marusov, D. Anikin, P. Sokerin et al.· 0 citations
It is concluded that, on this benchmark and within this family of frozen-encoder architectures, text content is not the operative signal behind the reported gains and the perturbation protocol and evaluation harness are released as a reusable diagnostic toolkit.
K. Sridhar, Atharva Gupta, Nishant Pradhan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.