Skip to content

GAttNHP: Group Attention Neural Hawkes Process for Extrapolation Reasoning in Temporal Knowledge Graphs

Jul 2026 · arXiv.org · Vol abs/2607.14733 · 0 citations · 32 references
Computer Science Mathematics

TL;DR

GAttNHP improves over state-of-the-art baselines on both entity prediction and time prediction, and ablations confirm that its largest gains arise on the long-tail event chains where existing models fail most severely.

Abstract

Temporal Knowledge Graphs (TKGs) record how facts evolve over time, but forecasting future events on a TKG remains difficult for three reasons: (i) long-range temporal dependencies are hard to encode; (ii) events on different chains mutually excite or inhibit one another in ways that snapshot-level models cannot express; and (iii) inter-arrival times are heavy-tailed and statistically sparse, so deterministic time predictors are unreliable. We address these three issues with a single framework, the \textbf{Group Attention Neural Hawkes Process (GAttNHP)}, built around three matched components. First, a self-attention encoder casts each subject--relation chain as a continuous-time point process and captures the lingering excitation of distant history. Second, a semantic soft-grouping module turns globally learnable Hawkes priors into an analytical cross-attention mask, so chains share excitation patterns through their latent group memberships rather than through exhaustive pairwise computation. Third, a Non-Crossing Quantile (NCQ) regression head replaces mean-based time prediction, providing calibrated, monotonically ordered quantile estimates that remain stable under heavy-tailed inter-arrival distributions. On six benchmark TKG datasets, GAttNHP improves over state-of-the-art baselines on both entity prediction and time prediction, and ablations confirm that its largest gains arise on the long-tail event chains where existing models fail most severely.

View source

Similar papers

Conference Open access Sep 2026

History Doesn’t Repeat, but Its Patterns Echo: A Parallel Pairwise Negative-Sampling Framework for Temporal Link Prediction

Temporal link prediction with temporal graph neural networks (TGNNs) is increasingly used to model spatio-temporal dependencies in temporal graphs and to forecast future interactions among entities. Existing sampling-based training methods typically rely on random negative sampling and pointwise loss formulations, which often lead to suboptimal convergence and limited generalization due to low-quality negative samples. We propose ATNSF, a temporal graph learning framework with a hybrid negative sampling strategy that uses a portion of historical edges as hard negatives. For efficiency, we design an asynchronous parallel training pipeline for scalable optimization and introduce a pairwise sampled softmax loss that contrasts each positive instance with a batch of negatives to learn more discriminative representations. Finally, we theoretically show that jointly designing the loss function and negative sampling strategy is crucial for improving performance and generalization. Extensive experiments across six temporal graph datasets demonstrate that ATNSF improves the average AP from 0.724 to 0.827 (+0.103). Remarkably, it also accelerates training by 1.37× to 6.85×, achieving a 2.57× geometric mean speedup. The source code of this paper can be found at https://github.com/yongqiu-star/ATNSF.

Yong-Chun Jiang, Heng Zhang, Jian Gao et al. · 0 citations
Preprint Aug 2026

SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting

This work proposes SAGE (Seeing and Augmenting with Grounded Encoding), an end-to-end CLIP-based framework that jointly models temporal, cross-variable, textual, and visual information and achieves state-of-the-art accuracy.

Hai-Zhao Fan, Xinh Le · 0 citations

Persistent Structure Meets Dynamic Attention: Cross-Variable Priors for Multivariate Time Series Forecasting

A Params-Per-Pair diagnostic is introduced that predicts from dataset properties alone whether structural priors will help and reveals a horizon-dependent complementarity: the structural prior contributes 33% of the gain at short horizons but 88% at long horizons, confirming that time-invariant knowledge compensates as temporal signal fades.

C. Mohapatra, Rohit Malshe, J. Pachón · 0 citations
Preprint Aug 2026

TIEM: Temporal Integration of Hypergraph Evidence and Skill Memory for Event-Driven Financial Forecasting

TIEM, a timestamp-gated framework with three coordinated components: an Event-Evidence Hypergraph (EEH) for timestamp-filtered multi-tier retrieval; a Case-based Skill Memory (CSM) for source-tagged temporal skills; and Heterogeneous Evidence-Experience Fusion Reasoning (HEFR) for evidence-experience fusion and prediction.

Wenjin Liu, Shengjie Pang, Chen-Xi Wang et al. · 1 citation
Preprint Aug 2026

FITTER: Vocabulary-Agnostic Cross-Domain Inference on Temporal Knowledge Graphs

FITTER consistently outperforms inductive baselines without retraining, indicating that vocabulary-agnostic structural learning is a viable foundation for inference over the heterogeneous knowledge graphs of the Semantic Web.

Jia-Xin Pan, M. Nayyeri, Osama Mohammed et al. · 0 citations
Jul 2026

What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation

Mass-Aware Attention is proposed, which generalizes standard L1 normalization to an Lp family and is positioned as a general normalization principle for improving predictor-facing representation informativeness by controlling repetition invariance in standard attention.

Min-Woo Yu, Young-Guk Ha · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.