Retrieval Augmented Generation (RAG) is a key component for generating accurate and hallucination free answers using Large Language Models (LLMs). LLMs are improving at handling long context, but still suffer from"lost in the middle"problem. Thus, precise and accurate retrieval is important. Current retrievers chunk lo...
Vineet Kumar, Meghanadh Pulivarthi, Vishwajeet Kumar et al.· 1 citation
RAG has become the de facto method for incorporating new, corpus-specific knowledge into an instruction following LLM (Instruct LLM). Although RAG-based prompting improves factual grounding, it fails when retrieval is incorrect or incomplete, leading to hallucinations. Finetuning methods such as RAFT and PA-RAG enhance...
Kushagra Bhushan, Meghanadh Pulivarthi, Sai Krishna Reddy Sathi et al.· 1 citation
This work introduces KItCAT: Knowledge Injection via Corrupted Auto-regressive Training, a lightweight training strategy that reduces the need for paraphrasing in decoder-only LLMs and shows that KItCAT consistently improves over CPT across multiple datasets and model families.
Meghanadh Pulivarthi, Kushagra Bhushan, Vineet Kumar et al.· 0 citations
SearchWiki paired with WikiResearcher-9B demonstrates that learned navigation over structured corpora is a superior alternative to flat retrieval, and optimizing the agent's navigation policy with on-policy reinforcement learning with a multi-component reward function balancing answer correctness, retrieval quality and...
Guransh Singh, Vishwajeet Kumar, Arkadeep Acharya et al.· 0 citations
It is demonstrated that models trained using ColSNAP maintain near full-resolution retrieval performance under substantial compression and that ColSNAP transfers effectively across multiple late-interaction backbones, and achieves most of its improvements via a lightweight adaptation stage applied to a pre-trained retr...