Jun 2026· arXiv.org· Vol abs/2606.21803· 0 citations· 41 references
Computer Science
TL;DR
Test-Time Training with Next-Token Prediction (TTT-NTP), a drop-in fast-weight adaptation method for pretrained LLMs that instead supervises updates using the model's own next contextual hidden state, while preserving commonsense and knowledge performance.
Abstract
Next-token prediction is the self-supervised signal that trains language models, and every observed prompt token provides the same signal at test time. We study whether this signal can define the inner-loop objective for test-time training (TTT) in pretrained long-context language models. Many TTT architectures require models to be trained with test-time adaptation in mind, limiting their direct applicability to released LLM checkpoints. While recent in-place TTT methods make fast-weight adaptation possible for pretrained LLMs without redesigning the backbone, they leave a central question unresolved: what should each test-time write store? Existing recipes train the fast weight to match a learned local value proxy but they are not directly tied to the self-supervised next-token prediction signal. We introduce Test-Time Training with Next-Token Prediction (TTT-NTP), a drop-in fast-weight adaptation method for pretrained LLMs that instead supervises updates using the model's own next contextual hidden state. This makes each local write follow the same causal computation that supports next-token prediction: the value target is a pointwise linear projection of a single next-position contextual state. On RULER Full-13, averaged over 4k to 32k contexts, TTT-NTP is the only method that consistently improves the released backbone across four models spanning three families and a 0.6-8B size range, by 3.9 points on Llama-3.1-8B, 3.0 on Mistral-7B-v0.3, 4.1 on Qwen3-4B, and 2.9 on Qwen3-0.6B. On the real-world LongBench-v2 long-document QA benchmark, TTT-NTP improves over the base model by 5.6 points on Llama-3.1-8B and 3.7 on Mistral-7B-v0.3, while preserving commonsense and knowledge performance. Our code is publicly available at https://github.com/yancyou/TTT-NTP.
The experimental results show that even 130M-parameter models benefit from including the MTP task in the pre-training objective, and hold even under severe data constraints, as demonstrated on both zero-shot benchmarks and downstream tasks.
Delayed semantic supervision as a practical outer training objective for usable test-time memory is supported, while leaving open how much of the benefit derives specifically from delay rather than general question-answering and answer-termination supervision.
Under the standard approximation of taking gradients at the chunk-start weights, a closed-form state transition is derived that exactly reproduces the chunk-end fast-weight and momentum states of the per-token recurrence of Test-Time Training.
Zeyun Zhong, Joya Chen, Manuel Martín et al.· 2 citations
Answer-Convergence Stopping (ACS) is introduced, a training-free stopping rule that measures rather than asks, and reveals that by properly utilizing the output signals of frozen models, it can achieve favorable behaviors like adaptive stopping without the need for additional training.
Muath Alyobi, M. Eltahir, Almoayyad Abuljdail et al.· 0 citations
It is shown that the mismatch degrades both the self-conditioning input and the solver update, and derive a correction for each from the model's own structure, and the resulting sampler, Untied Self-Conditioning, requires no retraining and uses one evaluation per step.
This paper takes state--prediction separation to its limit with a free pause token: a prediction stream that writes no keys or values at all and so rides the sequence's existing positions, and reduces the raw flops required at inference time.
John Langford, Nathan Godey, Giovanni Monea et al.· 1 citation
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.