Skip to content

LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos

Jul 2026 · arXiv.org · Vol abs/2607.12733 · 0 citations · 71 references
Computer Science

TL;DR

Elenchos is introduced, a generative evaluation framework that measures abductive reasoning as a structural inverse problem and suggests diminishing returns from increased inference-time reasoning, with only modest improvements under larger reasoning budgets.

Abstract

Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses that explain observed behavior - remains poorly understood. Here, we introduce Elenchos (named after the Socratic method of cross-examination), a generative evaluation framework that measures abductive reasoning as a structural inverse problem. Given a reference formal system, such as the lambda-calculus, and a potentially mutated counterpart, agents must determine whether a mutation has occurred and infer the rule modifications responsible for the resulting behavioral differences. Evaluating frontier and mid-tier LLMs reveals a consistent detection-attribution dissociation: models often recognize that a system has been altered but struggle to identify the latent mutations causing the observed discrepancies. Performance degrades substantially under interacting mutations, where models frequently recover only a subset of the underlying mutations. Preliminary evidence also suggests diminishing returns from increased inference-time reasoning, with only modest improvements under larger reasoning budgets, though this finding requires further validation.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark

DeReLab is introduced, a generative framework that produces multi-turn belief-updating conversations from parameterized graph structures across default and inheritance reasoning, with formally verified ground truth at every turn, enabling controlled measurement of how models respond to confirming and disconfirming evid...

Jayanta Sadhu, S. Shahad, Kenneth Marino · 1 citation
Book Open access Aug 2026

SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification

Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when final answers are correct. Most existing ''verification'' signals are not diagnostic: answer matching observes only the outcome, LLM-as-judge provides subjective and non-verifiable cri...

Wenyao Cui, Hua-Ping Zhang, Yongyi Huang et al. · 0 citations
#natural language process... Preprint Aug 2026

Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents

This work proposes a refined security principle: only agents capable of deliberative System 2 reasoning may access untrusted documents, and empirically compares state-of-the-art reasoning language models with standard language models to show that reasoning-capable models are substantially more robust to corrupted evide...

Mehrdad Ghassabi · 0 citations
Preprint Aug 2026

The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics

Results show that capability failures can manifest as distributed, task-dependent changes in the structure of visible reasoning, and that CoT dynamics agnostic to whether the verbalized trace reflects the model's internal computations can help diagnose and correct failures.

Shashwat Sourav, Aishwarya H. Balwani · 1 citation
#artificial intelligence Preprint Sep 2026

State of Thought Enables Endogenous Reasoning

Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reasoning programs or through costly expansion in constrained search spaces, limiting both gen...

Z. Gong, Yi-Kun Hou, Zi-Hao Zeng et al. · 0 citations
Preprint Aug 2026

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

The nature of test-time exploration in RLVR-trained LLMs is investigated by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence to delineate between entropy arising from stylistic variations and genuine inferential...

Soumadeep Saha, Krish Sharma, Akshay Chaturvedi et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.