Skip to content

Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination

Jul 2026 · arXiv.org · Vol abs/2607.08403 · 0 citations · 35 references
Computer Science

TL;DR

G-Frame, an adaptive multi-agent framework integrating Bayesian and team game principles, establishes an automated closed-loop for high-quality data synthesis and model training and synthesizes a specialized corpus of 363,045 chains-of-thought and 199,589 question-answer pairs.

Abstract

The application of lightweight Large Language Models in rule-based scientific domains remains severely limited by their tendency to mimic linguistic patterns rather than reproduce axiomatic reasoning, causing frequent hallucinations. Here, we show that G-Frame, an adaptive multi-agent framework integrating Bayesian and team game principles, establishes an automated closed-loop for high-quality data synthesis and model training. By forcing the internalization of domain constraints through structured reasoning, we synthesized a specialized corpus of 363,045 chains-of-thought and 199,589 question-answer pairs. The resulting 7B model OmniChem achieves performance parity with GPT 4o mini on custom benchmarks and ChemBench while exhibiting a 79.46% reduction in hallucinations relative to its base architecture. We further demonstrate the advanced capabilities of OmniChem in molecular design and synthesis planning. This work establishes a scalable paradigm utilizing adaptive multi-agents to overcome inherent reasoning deficiencies, offering a feasible pathway for accelerating knowledge discovery in specialized scientific fields.

View source

Similar papers

Conference Aug 2026

Quorum-Inspired Multi-Agent Consensus for Detecting Reasoning Hallucinations in Retrieval-Augmented Generation: A Cautionary Study

Collective decision-making in nature, such as bacterial quorum sensing and swarm consensus, has inspired the intuition that a panel of large language model (LLM) verifiers can outperform a single model through aggregated voting. In this work, we rigorously evaluate this hypothesis for detecting reasoning hallucinations...

Anish Chaulagain, Dipika Acharya, Lucky Shrestha et al. · 0 citations
Preprint Aug 2026

Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations

SCHEMA reveals that hallucinations concentrate at a small set of highly connected knowledge hubs, and that final-answer accuracy decouples from trajectory honesty; models often reach correct conclusions through structurally flawed reasoning.

Xinshun Feng, Ziqi Miao, Li-Jun Li et al. · 0 citations
#small language model Book Open access Sep 2026

Evaluating LLM Social Cognition Through Multi-Agentic Strategic Games

Large language models (LLMs) now power the reasoning core of intelligent virtual agents deployed across an expanding range of social settings, from tutoring students and supporting patients in healthcare, to mediating group discussions and representing humans in various social settings. Effective deployment demands soc...

Kevin Kurian, Kevin Scroggins, Emmanuel Dorley et al. · 0 citations
#natural language process... Preprint Sep 2026

MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance

Recent advances in Large Language Models (LLMs) have revolutionized artificial intelligence and how human interact with AIs. Despite impressive advancements, LLMs struggle with complex mathematical, scientific, and logical tasks. Inspired by human cognitive flexibility - our ability to dynamically switch mental perspec...

Arash Lagzian, Srinivas Anumasa, Dian-Bo Liu · 0 citations
Preprint Aug 2026

Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary

Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to experts and non-experts alike. Large language models could, in principle, bridge this gap, but...

S. A. Hebbar, Peiyao Sheng, Sewoong Oh et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Self-Play Search Distillation for Large Language Model Reasoning

Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a fra...

Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.