A Survey of Retrieval-Augmented Language Models for Knowledge-Intensive Text Applications
Retrieval-Augmented Language Models (RALMs) have emerged as an effective approach for addressing the limitations of conventional language models in knowledge-intensive text applications. These models combine external knowledge retrieval and language generation, enabling them to deliver more relevant, accurate, and contextually appropriate responses and alleviate the need for internal knowledge. This survey reviews the fundamental concepts, architecture, knowledge sources, retrieval techniques, and major types of Retrieval-Augmented Generation (RAG) systems. It explores how the retrieval-based language generation process is affected by textual documents, scientific literature, databases, knowledge bases, enterprise documents and multimodal sources. The survey also covers the use of RAG for knowledge-intensive tasks, such as question answering, reasoning, document analysis, summarization, and information extraction. Particularly, emerging applications in healthcare and education, where reliable and domain-specific knowledge retrieval is a necessity, are given special attention. Besides, the survey stresses on some challenges related to retrieval quality, knowledge freshness, contextual relevance, hallucination, and system scalability. Last, future research directions on enhancing the robustness, efficiency, reliability and domain adaptability of RALMs are discussed.