Skip to content

Test-Time Training with Next-Token Prediction

Jun 2026 · arXiv.org · Vol abs/2606.21803 · 0 citations · 41 references
Computer Science

TL;DR

Test-Time Training with Next-Token Prediction (TTT-NTP), a drop-in fast-weight adaptation method for pretrained LLMs that instead supervises updates using the model's own next contextual hidden state, while preserving commonsense and knowledge performance.

Abstract

Next-token prediction is the self-supervised signal that trains language models, and every observed prompt token provides the same signal at test time. We study whether this signal can define the inner-loop objective for test-time training (TTT) in pretrained long-context language models. Many TTT architectures require models to be trained with test-time adaptation in mind, limiting their direct applicability to released LLM checkpoints. While recent in-place TTT methods make fast-weight adaptation possible for pretrained LLMs without redesigning the backbone, they leave a central question unresolved: what should each test-time write store? Existing recipes train the fast weight to match a learned local value proxy but they are not directly tied to the self-supervised next-token prediction signal. We introduce Test-Time Training with Next-Token Prediction (TTT-NTP), a drop-in fast-weight adaptation method for pretrained LLMs that instead supervises updates using the model's own next contextual hidden state. This makes each local write follow the same causal computation that supports next-token prediction: the value target is a pointwise linear projection of a single next-position contextual state. On RULER Full-13, averaged over 4k to 32k contexts, TTT-NTP is the only method that consistently improves the released backbone across four models spanning three families and a 0.6-8B size range, by 3.9 points on Llama-3.1-8B, 3.0 on Mistral-7B-v0.3, 4.1 on Qwen3-4B, and 2.9 on Qwen3-0.6B. On the real-world LongBench-v2 long-document QA benchmark, TTT-NTP improves over the base model by 5.6 points on Llama-3.1-8B and 3.7 on Mistral-7B-v0.3, while preserving commonsense and knowledge performance. Our code is publicly available at https://github.com/yancyou/TTT-NTP.

View source

Similar papers

Babies Learn to Look Ahead: Multi-Token Prediction in Small LMs

The experimental results show that even 130M-parameter models benefit from including the MTP task in the pre-training objective, and hold even under severe data constraints, as demonstrated on both zero-shot benchmarks and downstream tasks.

Ansar Aynetdinov, Alan Akbik · 1 citation
#artificial intelligence Preprint Sep 2026

Delayed Supervision for Test-Time Language Models

Delayed semantic supervision as a practical outer training objective for usable test-time memory is supported, while leaving open how much of the benefit derives specifically from delay rather than general question-answering and answer-termination supervision.

Jinha Kim, Taksh Kothari · 0 citations
Preprint Aug 2026

Rethinking Expressivity and Efficiency in Test-Time Training

Under the standard approximation of taking gradients at the chunk-start weights, a closed-form state transition is derived that exactly reproduces the chunk-end fast-weight and momentum states of the per-token recurrence of Test-Time Training.

Zeyun Zhong, Joya Chen, Manuel Martín et al. · 2 citations
#machine learning Preprint Sep 2026

The Model Knows When to Stop: Training-Free Early Stopping for Long-Context Reading

Answer-Convergence Stopping (ACS) is introduced, a training-free stopping rule that measures rather than asks, and reveals that by properly utilizing the output signals of frozen models, it can achieve favorable behaviors like adaptive stopping without the need for additional training.

Muath Alyobi, M. Eltahir, Almoayyad Abuljdail et al. · 0 citations
Preprint Aug 2026

Improving Few-Step Language Flows with Untied Self-Conditioning

It is shown that the mismatch degrades both the self-conditioning input and the solver update, and derive a correction for each from the model's own structure, and the resulting sampler, Untied Self-Conditioning, requires no retraining and uses one evaluation per step.

Bocheng Li, Linli Xu · 1 citation
#artificial intelligence Preprint Sep 2026

Almost Free State Prediction Separation

This paper takes state--prediction separation to its limit with a free pause token: a prediction stream that writes no keys or values at all and so rides the sequence's existing positions, and reduces the raw flops required at inference time.

John Langford, Nathan Godey, Giovanni Monea et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.