Taxonomies of the LLM hallucination phenomena and evaluation benchmarks are presented, existing approaches aiming at mitigating LLM hallucination are analyzed, and potential directions for future research are discussed.
A natural way to decrease hallucinations is every time the authors have new pairs on which to train, they should also again train on the pairs corresponding to well-established facts and rules.
T. Ilina, M. Ceberio, Marcelo F. Frias et al.· 0 citations
Interpretable machine learning for Large Language Models (LLMs) increasingly relies on sparse probing methods that identify small sets of neurons claimed to detect and causally influence behaviors such as factuality recall, safety alignment, and hallucination. These claims have important implications for model auditing...
Huseyin Cavus, Sebin Sabu, J. Spear et al.· 0 citations
Language, truth, and LLMs How can LLMs appear to know facts about the world when they are trained only on patterns in language rather than direct experience of reality? Hallucinations shouldn’t surprise us. The real mystery is how these systems get anything right at all.
This work proposes a method to rank feed-forward neurons at the final prompt token using a custom neuron selection dataset, and transfers the selected neuron identities to train hallucination classifiers on other factual question answering datasets.
Ali Derogar Odolou, Reza Nazari, Mostafa Salehi· 0 citations
Large Language Models (LLMs) frequently exhibit hallucinations, presenting a major barrier to reliability in complex reasoning tasks. While traditional detection methods rely on output-based confidence metrics, these logits are often miscalibrated by modern alignment techniques. In this paper, we investigate the tempor...
AI hallucination is not just about a one-off technical glitch at the model level but is a systemic issue with the interplay between humans, the data, and the model itself. Most existing reviews focus on isolated parts of model development, like model architecture or data governance, and do not consider the whole model...
Jun-Hua Chen· Applied and Computational En...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.