Skip to content

Training Continuous Chain of Thought Models: A Tale of Two Regimes

Jul 2026 · arXiv.org · Vol abs/2607.16972 · 1 citation · 50 references
Computer Science

TL;DR

C-MTP is introduced, a simpler, faster direct supervision approach that models each latent as an average of the embeddings in the CoT traces to be compressed, and outperforms a prior direct supervision method that approximates the distribution of compressed tokens.

Abstract

Continuous Chain-of-Thought methods replace verbose reasoning traces with a short sequence of dense latent representations. Earlier continuous CoT methods indirectly supervise the latent representations such that its final state match that of verbose reasoning traces, requiring autoregressive, slow generation during training. We introduce C-MTP, a simpler, faster direct supervision approach that models each latent as an average of the embeddings in the CoT traces to be compressed. Our approach outperforms a prior direct supervision method that approximates the distribution of compressed tokens, and performs competitively to slower indirect supervision approaches in existing evaluation setup with simplified CoT traces (less than 100 tokens). Lastly, we extend the evaluation of Continuous CoT methods to complex tasks with longer reasoning traces ($\ge$ few hundreds reasoning tokens). We find both direct and indirect supervision training methods perform poorly (roughly 65\% performance drop) in this setting, revealing the limitations of current continuous CoT methods. The code and checkpoints are released at https://github.com/Varun221/cmtp_research

View source

Similar papers

#artificial intelligence Preprint Sep 2026

A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM

Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics...

Xiao-An Xu, Si-Yuan Liu, Shuo Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Structural Process Supervision for Latent Chain-of-Thought Reasoning

Latent reasoning approaches enhance token-level efficiency and robustness by replacing verbose, explicit chain-of-thought (CoT) tokens with compact continuous-space embeddings. However, existing methods lack efficient process supervision over these embeddings, which often leads to representation collapse and uneven inf...

Yi-Qi Li, Xu Chen, Chen Ju et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs

Latent Recurrent Thoughts substantially outperforms prior frozen-decoder continuous-space reasoning methods under an identical decoder, prompt, data, and training budget, and outperforms non-thinking-mode chain-of-thought prompting on the same backbone at a small fraction of its inference compute.

Zhaoxing Chen, Jie Fu · 0 citations
#machine learning Preprint Sep 2026

The Dynamics of Continuous Mixture Collapse in Language Models

This work studies why pretrained language models often fail to preserve mixtures of many components and shows that exact preservation generally requires context-dependent correction, whose required dimensionality can grow with the number of components.

Ali Backour · 0 citations
#computer vision Preprint Aug 2026

Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning

Latent-OPD is proposed, which augments OPD with trajectory-level latent distillation and introduces a progressive teacher-lookahead strategy, which aligns middle-to-late student layers with increasingly deeper teacher layers, establishing Latent-OPD as a highly effective approach to frame-efficient video reasoning.

Aoni Shen, Yongheng Zhang, Ying-Hui Li et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.