Skip to content

Complex Problem Solving in Large Language Models: A Statistical Control Survey and Diagnostic Framework

Sep 2026 · 0 citations
Mathematics Computer Science

TL;DR

This framework organizes existing methods around five components: explicit state representation, transition structuring, validation and constraint enforcement, search and rollback, and uncertainty management, and yields a diagnostic hypothesis: interventions should be most effective when they target the error or uncertainty component implicated by an observed failure.

Abstract

Complex problem solving (CPS) with large language models (LLMs) is often framed as a matter of stronger reasoning or longer generation. Yet early-step error amplification, prompt brittleness, and failures to revise incorrect commitments are difficult to explain by missing knowledge or expressive capacity alone. This survey interprets CPS as a sequential estimation-and-decision problem over a latent solution state. A controller maintains a belief about an unobserved solution trajectory, updates it as noisy intermediate evidence arrives, and decides whether to commit, verify, branch, roll back, or abstain to minimize expected loss. Reasoning supplies candidate transitions and interpretations, whereas process control shapes and evaluates those proposals and regulates subsequent transitions and observations. Within this framework, we organize existing methods around five components: explicit state representation, transition structuring, validation and constraint enforcement, search and rollback, and uncertainty management. We also interpret evaluation metrics according to the statistical quantities they estimate. The framework further yields a diagnostic hypothesis: interventions should be most effective when they target the error or uncertainty component implicated by an observed failure. We distinguish systematic, stochastic, and irreducible error together with epistemic and aleatoric uncertainty, and call this alignment problem-control fit and its failure control mismatch. For example, additional sampling may reduce sampling variability while leaving a shared systematic error unchanged. This perspective clarifies what current methods estimate and control, what remains uncontrolled, and why reliable validation, targeted recovery, calibrated uncertainty, and matched-budget evaluation are central open problems.

View source

Similar papers

Preprint Aug 2026

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

This empirical study covers broad knowledge, symbolic reasoning, and competition mathematics, and it introduces an evaluation profile whose coordinates and simple functionals recover or bound common repeated-sampling metrics, and require compute accounting and uncertainty estimates that match the protocol.

Mohsen Hariri, Weicong Chen, Nahal Shahini et al. · 5 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability

It is proved that the LLM paired with the BSE is a sound Markov policy on the belief MDP induced by the underlying POMDP, and inherits the Bellman optimality guarantees of classical POMDP theory, provided the LLM is never exposed to the raw history.

Arnab Chattopadhayay, Debdipta Halder · 0 citations
#machine learning Preprint Sep 2026

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

LLM agents deployed for software engineering fail expensively: they act confidently wrong, and bad actions are recognized only after costly execution and retry. We present Speculative Uncertainty (SU), a method that recovers a predictive failure signal for a black-box agent from its output tokens alone, with no access...

Konstantin Grotov, Valentin Malykh · 1 citation
#artificial intelligence Preprint Aug 2026

Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy

Large language model (LLM) agents may perform well on isolated tasks yet drift into inconsistency over extended interaction. We evaluate temporal consistency in a controlled 20-step multi-agent setting inspired by delayed-gratification studies. At each step, an agent chooses between continuing to delay a reward or clai...

Igor Bogdanov, Olga Manakina, Chung-Horng Lung · 1 citation
#machine learning Preprint Sep 2026

Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

A central question in LLM reasoning is whether reinforcement learning (RL) instills genuinely new capabilities or merely reshapes how existing knowledge is expressed during inference. Building on the distribution-sharpening hypothesis, which holds that RL reallocates probability mass toward high-reward trajectories alr...

Zhen-Dong Mi, Shao-Yi Huang · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.