CONGA: Continual Neural Gated Architecture for Long-History Sequential Recommendation
Abstract
Existing sequential recommendation models rely on absolute positional encodings and fixed context windows, producing two structural failure modes: out-of-distribution degradation on histories longer than the training window, and hard context limits that discard long-range interactions. We present CONGA (COntinual Neural Gated Architecture), which addresses these limitations through three contributions: (1) Rotary Positional Embeddings (RoPE) with a norm-preserving property (\(\Vert \mathcal {R}_m\Vert _F = \sqrt {d}\) for any sequence length), accelerated by custom CUDA kernels; (2) KromHC multi-stream fusion with exact doubly-stochastic mixing via Kronecker-product parametrization, where ablation confirms the expressivity gain arises from balanced gradient flow rather than additional parameters, together with a data-adaptive stream selection mechanism that prevents overfitting on sparse corpora; and (3) TITANS neural associative memory adapted to discrete-item recommendation – the first such proof-of-concept – via a two-phase training protocol with a structural forgetting-prevention property: the base encoder is frozen, preserving its short-sequence predictions, while Phase 2 only adds a learned memory term. Evaluated under a rigorous full-ranking protocol across four benchmarks (ML-1M, Beauty, Yelp, Steam), CONGA achieves state-of-the-art performance on long-history benchmarks, with gains up to \(+18.7\%\) HR@5 on ML-1M – the densest, longest-history dataset where all three failure modes are simultaneously active – and consistent improvements on long-history benchmarks, with competitive results on short-history datasets where the RoPE backbone is the primary driver.