Skip to content
Open access

Can LLMs Self-Correct Table Reasoning Errors?

2026 · Proceedings of the First Workshop on Structured Understanding, Retrieval, and Generation in the LLM Era (SURGeLLM 2026) · pp. 298-312 · 1 citation · 26 references

TL;DR

Structured Self-Correction (SSC) is proposed, a table-specific verification chain that guides models through cell verification, computation checking, logic validation, and completeness assessment, indicating that mode robustness itself is model-dependent.

Abstract

Self-correction—the ability of LLMs to detect and fix their own errors—has been studied extensively for mathematical and code reasoning, with limited prior work on table reasoning (primarily multi-agent pipelines such as Table-Critic, ACL 2025, rather than single-model structured prompting). Tables present unique challenges: errors arise from wrong cell retrieval, incorrect computation, flawed logic, and hallucination of values not present in the data. We conduct the first cross-provider single-model self-correction analysis for table reasoning across five providers (Google, Moonshot AI, Zhipu, Alibaba, Min-iMax), testing five models (Gemini 3.1 Pro, Kimi K2.5, GLM 5, Qwen 3.5+, MiniMax M2.5) on WikiTableQuestions and TabFact with a multi-seed paired protocol. We pro-pose Structured Self-Correction (SSC) , a table-specific verification chain that guides models through cell verification, computation checking, logic validation, and completeness assessment. We confirm that the Accuracy-Correction Paradox (terminology from Li (2025)) previously observed in math extends to tables: models with base accuracy in the mid-60s–mid-70s region benefit modestly from self-correction (multi-seed mean SCG up to + 1.3% with within-seed point estimates as high as +3.4%), while stronger models above this region are systematically harmed by over-correction (multi-seed mean SCG down to − 1.3%, with 95% bootstrap CIs significantly below zero). SSC reduces over-correction rates in 9 of 10 conditions, with reductions of 38–69% on TabFact. An inference-mode-controlled probe shows that SSC’s qualitative direction is robust for Qwen 3.5+ across reasoning-ON and reasoning-OFF settings, while GLM 5 exhibits a substantial mode-dependent shift, indicating that mode robustness itself is model-dependent. Stronger

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Available but Unclaimed: An Empirical Study of Human-AI Synergy

People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-L...

Robin Welsch, Michelle Rausch, Pascal Knierim et al. · 0 citations
Review Open access Sep 2026

Why Retrieval Doesn't Cure Everything: A Review of Hallucination in Retrieval-Augmented Generation

responses depending on domain, retriever quality, and model family. This paper reviews the literature on why RAG systems continue to hallucinate even when correct evidence is available in context, organizes the reported causes into a five-part taxonomy (retrieval failure, conflicting evidence, unfaithful generation, ov...

Sanchita H., Skandamahima V. M., Akshitha Katkeri · 0 citations
Preprint Aug 2026

Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary

Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to experts and non-experts alike. Large language models could, in principle, bridge this gap, but...

S. A. Hebbar, Peiyao Sheng, Sewoong Oh et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Self-Play Search Distillation for Large Language Model Reasoning

Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a fra...

Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni et al. · 0 citations
2026

A Draft-Ground-Verify-Revise Framework for Reducing Hallucination in Large Language Models

Large language models (LLMs) often generate fluent but factually unsupported or logically invalid text: failure modes broadly referred to as hallucination. A variety of approaches have been proposed to mitigate hallucination, ranging from careful benchmark design, preference modeling, and fine-tuning (e.g., RLHF) to po...

Syed Mubashir · 0 citations
Preprint Aug 2026

Execution-Anchored Hallucination Calibration Reranking for Verilog Code Generation

EAHC is proposed, an Execution-Anchored Hallucination Calibration reranking framework that anchors reasoning judgments to execution behavior so that execution-equivalent candidates receive consistent scores, which implements a dual-channel architecture.

Guang Yang, Xing Hu, Xiang Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.