Aug 2026· Language Resources and Evaluation· Vol 60· 0 citations· 81 references
TL;DR
This review paper provides a comprehensive overview of hallucinations in GAI and LLMs, and synthesizes a range of correction and mitigation techniques, from proactive measures during training to hybrid approaches that combine detection and intervention.
Abstract
Generative Artificial Intelligence (GAI) and Large Language Models (LLMs) have demonstrated significant capabilities in generating human-like content; however, they exhibit a propensity to fabricate spurious information, a phenomenon often termed hallucination. This review paper provides a comprehensive overview of hallucinations in GAI and LLMs. More specifically, this review encompasses their definitions, underlying mechanisms, taxonomies, commonly used tests, and datasets for evaluating hallucinations. In addition, this review dives into intrinsic and extrinsic factors contributing to these inaccuracies, including limitations in model architectures, training data biases, and inference algorithms, as well as examines various detection strategies [e.g., post-hoc consistency checks, external fact-checking, contrastive learning, uncertainty calibration methods, and Retrieval-Augmented Generation (RAG)]. The review also synthesizes a range of correction and mitigation techniques, from proactive measures during training to hybrid approaches that combine detection and intervention. Finally, this review integrates qualitative assessments and comparative insights to delineate the impact of hallucinations on user trust and acceptability, and to shed light on current challenges and future research trends.
The results suggest that no single architecture guarantees factual reliability, however, contextual grounding and verification mechanisms can significantly improve response quality and highlight the importance of combining language modelling capabilities with grounding strategies to support the development of more reliable AI systems.
This review provides systematic theoretical support for industrial RAG model selection and optimization and summarizes existing research gaps, including lightweight deployment and multimodal expansion, and proposes future research directions for trustworthy RAG systems.
Shujing Liu· Applied and Computational En...· 0 citations
This work proposes AURORA, a novel hallucination detection framework that shifts the focus from static representations to the weight-gradient dynamics of LLMs, and achieves strong hallucination detection performance across four model families and four benchmark datasets.
Z. Zhang, Hainan Zhang, Zhiming Zheng· arXiv.org· 0 citations
This survey provides a comprehensive treatment of the field across five interconnected dimensions, proposing a unified five-class taxonomy that organizes hallucinations by their failure mode: object, attribute, relational, factual, factual, and reasoning.
A. O. Ogar, Joshua Abah, M. Suleiman et al.· 0 citations
This survey addresses hallucination mitigation through the lens of explainability, proposing a taxonomy that distinguishes between internal explainability and post hoc explainability and discusses the constructive role of hallucinations in creative and user‐driven applications.
Wentao Deng, Jiao Li, Hongyu Zhang et al.· WIREs Data Mining and Knowle...· 1 citation
This study proposes a novel conceptual framework and taxonomy for hallucination mitigation in low-code AI environments, integrating retrieval, validation, conflict resolution, and workflow orchestration mechanisms to contribute to the development of more reliable, transparent, and scalable AI systems.
I. K. W. Adnyana, Rosalin Theophilia Tayane, Fahmi Fahmi et al.· EDUKASIA Jurnal Pendidikan d...· 0 citations