A Draft-Ground-Verify-Revise Framework for Reducing Hallucination in Large Language Models
Abstract
Large language models (LLMs) often generate fluent but factually unsupported or logically invalid text: failure modes broadly referred to as hallucination. A variety of approaches have been proposed to mitigate hallucination, ranging from careful benchmark design, preference modeling, and fine-tuning (e.g., RLHF) to post-hoc grounding via retrieval-augmented generation (RAG) or chain-of-verification. Prior work primarily considers these approaches as competing options, but we posit that they should be viewed as complementary stages in a coherent system for reliable LLM-based reasoning. We introduce the Draft-Ground-Verify-Revise (DGVR) framework that unifies these approaches under a single umbrella. Crucially, we find that the verification step within DGVR needs to distinguish between two types of hallucinations: claims that are unsupported by external facts and invalid inferences that follow logically from true premises. With these distinctions in place, we show precisely how DGVR explains both the limitations of hallucination in RLHF-aligned models and the shortcomings of retrieval-only approaches such as RAG. Lastly, we identify the open problems in developing practical systems based on the DGVR framework. In particular, we highlight the need for decomposition of claims into atomic statements, development of dual-subchecker verifiers, and calibrated abstention from answer generation.