Skip to content

When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models

Aug 2026 · 0 citations · 47 references
Computer Science

TL;DR

A formal account of the jump is developed in four steps and measured, proving that jump instances are well-posed and establish a family theorem that certifies instances of unbounded difficulty without enumeration and further formalize when a jump is correct and how successive jumps compound.

Abstract

Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are structurally incapable of such jumps, while recent studies challenge both its mechanism and its evidence. However, the debate remains difficult to settle, since the field still lacks a formal definition of the jump and a measure to test either side. In this paper, we develop a formal account of the jump in four steps and measure the second. The steps ask what the default completion of partial data is, when abandoning it is forced, when the abandonment is correct, and how successive jumps compound. Specifically, we define a jump instance as a finite extension problem with a machine-checked certificate that a correct completion exists, is unique up to renaming, and differs from the canonical completion of the data. The canonical completion is given by the left and right Kan extensions and is also what models produce without constraints, so it serves as the default. We prove that jump instances are well-posed and establish a family theorem that certifies instances of unbounded difficulty without enumeration. We further formalize when a jump is correct and how successive jumps compound. Finally, we run the measurement on nine certified instances and four frontier models. The Kan-default rate is zero in all 248 constrained trials, so the models do jump at this step and abandon the excluded default every time. Failures at higher difficulty stem from exhausted reasoning budgets or constraint errors, never from reverting to the default. These results indicate that the second step is not the bottleneck. If the disputed incapacity is real, it lies in generating the constraints or inventing the framework. Code can be found at: https://github.com/EEthanShi/kan-jump-test.

View source

Similar papers

Preprint Aug 2026

When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models

Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are structurally incapable of such jumps, while recent studies challenge both its mechanism and em...

Dai Shi, Xiao-Yu Li, José Miguel Hernández-Lobato · 0 citations
2026

ProofTeller: Exposing Recency Bias in LLM Reasoning and Its Side Effects on Communication (Extended Abstract)

ProofTeller, a benchmark that evaluates large language models' ability to faithfully interpret its reasoning, finds a consistent near-conclusion bias: LLMs tend to focus on steps closest to the final proof conclusion rather than on the most informative ones.

Mayank Jobanputra, A. Kovtunova, Brisca Balthes et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification

Evaluating state-of-the-art open LLMs reveals a significant robustness gap, and shows that reliable process-level verification remains challenging, and evaluator robustness should be measured separately from solver accuracy, even in a simple algebraic domain with exact ground truth.

Fatemeh Mazdarani, Carlos Toxtli · 1 citation
#artificial intelligence Preprint Aug 2026

Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free

It is argued that machine checking produces verification abundance while leaving adjudication scarce, and proposes a six-category taxonomy of representational mismatch, a disclosure schema for machine-generated mathematical claims, and implications for software, cryptography, and regulated decision systems.

Maher Kallel, Mohamed El Louadi · 0 citations
Preprint Aug 2026

Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction

This work introduces Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k), with every constraint scored by a deterministic, rule-based verifier and zero LLM-judge involvement.

Mariya I. Vasileva · 0 citations
Preprint Aug 2026

The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure

The results show that models often do not affirm culmination but nevertheless accept the corresponding simple-past hypothesis, a pattern the authors characterize as Sufficiency Bias, and show that prompting interventions produce a Decision Shift among labels without reliably improving the underlying semantic understand...

Kaiqiao Han, Yi-Zhou Sun · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.