Skip to content

Delayed Supervision for Test-Time Language Models

Sep 2026 · 0 citations · 12 references
Computer Science

TL;DR

Delayed semantic supervision as a practical outer training objective for usable test-time memory is supported, while leaving open how much of the benefit derives specifically from delay rather than general question-answering and answer-termination supervision.

Abstract

Test-time language models adapt a compact memory while processing the input sequence. This perspective encompasses nonlinear fast-weight learning in LaCT, associative delta-rule updates in DeltaNet, and generalized delta-rule state updates in RWKV-7. Training these models to predict the next token does not explicitly require a fact to remain accessible after many subsequent memory updates. We study delayed supervision for this test-time memory: during post-training, ask a simulator-grounded question only after a long interval of unrelated events, and supervise its answer alongside ordinary next-token prediction. Questions are evaluated on disposable branches, so their answers never enter the continuing event stream. The construction distinguishes retention from revision: a retained fact must remain valid throughout the delay, whereas a revised fact must be answered with its latest value. We evaluate this approach on LaCT-760M and plain DeltaNet-1.3B using TextWorld training trajectories and shared BABILong and RULER evaluation panels, and include a separately reported RWKV-7 comparison. Relative to event-only training, delayed QA improves BABILong by 5.48 percentage points for LaCT and 1.32 points for DeltaNet, and single-needle RULER by 1.45 and 3.27 points, respectively. The RWKV-7 comparison reports gains of 4.60 and 7.00 points on its own panels. These results support delayed semantic supervision as a practical outer training objective for usable test-time memory, while leaving open how much of the benefit derives specifically from delay rather than general question-answering and answer-termination supervision.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

Decoupling Memory from Context: Structured Memory for Token-Efficient Test-Time Continual Learning

Large language models (LLMs) are increasingly deployed in enterprise, scientific, and medical applications, where agents must incorporate domain-specific knowledge and adapt from experience. Context engineering offers a practical alternative to weight updates by improving model behavior through instructions, strategies...

Yehya Farhat, Michael Desmond, Anastasios Kyrillidis · 0 citations
Preprint Aug 2026

Rethinking Expressivity and Efficiency in Test-Time Training

Under the standard approximation of taking gradients at the chunk-start weights, a closed-form state transition is derived that exactly reproduces the chunk-end fast-weight and momentum states of the per-token recurrence of Test-Time Training.

Zeyun Zhong, Joya Chen, Manuel Martín et al. · 2 citations
#artificial intelligence Preprint Oct 2026

Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation

Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-wr...

Cheng-Zhi Luo, Bing Li, Bernard Ghanem · 0 citations
#natural language process... Preprint Sep 2026

What Attention Recalls and Recurrence Controls in Hybrid Language Models

Hybrid language models combine attention with a fixed-size recurrent state, but the role of each channel remains unclear. We introduce two cache-level interventions. Split-prefill keeps only the KV cache or only the recurrent state from a prefilled context, then generates an answer. State-swap pairs the KV cache from o...

Kirill Afendulev, Alexey Dontsov, Elena Tutubalina et al. · 2 citations · ⚡1

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.