RepoNav is introduced, a lightweight post-retrieval interface that reorganizes retrieved snippets into a file-centered navigation scaffold that improves function-level localization and narrows the file-to-function gap.
Abstract
Solving repository-level code tasks requires LLM-based agents to use code search tools to navigate large codebases and identify a small set of relevant files and functions. However, current retrieval tools typically return flat lists of isolated code snippets: such lists can surface relevant files, but provide insufficient structure for agents to distinguish the target function from semantically similar alternatives in the same file. We introduce RepoNav, a lightweight post-retrieval interface that reorganizes retrieved snippets into a file-centered navigation scaffold. By presenting compact structural cues and candidate targets, this scaffold guides on-demand file-structure browsing, helping agents compare sibling symbols before selecting a target function. Across diverse models on LocBench, RepoNav improves function-level localization and narrows the file-to-function gap. Controlled ablations demonstrate that these gains come from structured evidence organization rather than simply exposing additional file structure, and the approach also improves performance on a repository-level question-answering benchmark.
Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue. LLM agents approach this task iteratively: they identify a set of potentially relevant locations, inspect the corresponding code, and revise their judgments about these candidates as new evidenc...
Recent code generation research has moved from isolated function completion toward repository-level generation in existing codebases. To implement a target function correctly, an LLM must identify reusable repository dependencies such as existing functions, APIs, and cross-file definitions. Existing retrieval methods p...
Xu-Tian Li, Bo Xiong, Yi-Feng Zhu et al.· 0 citations
We introduce C2C (From Codebase to Culprit), a framework for precise bug localization that progressively reduces the debugging search space across multiple levels of granularity: files, functions, and lines of code. To mirror developer's natural top-down debugging workflows, C2C integrates semantic retrieval and Hierar...
Ankur Garg, Corey Yang-Smith, Rishav Rishav et al.· 0 citations
CrossCoder is proposed, a cross-repository code generation framework that explicitly incorporates external libraries into the retrieval context through a unified knowledge graph over repository and library entities that consistently improves both functional correctness and robustness to dependency-version changes.
Minh Le-Anh, Nam Le Hai, Quyen Tran et al.· 0 citations
Large language model (LLM)-powered coding agents have made rapid progress in automating software engineering tasks, yet repository-level issue resolution remains challenging. Beyond generating a plausible patch, an agent must localize relevant code across interdependent files and maintain repository context that is bot...
Yun-Xiang Zhang, Hai-Quan Wang, Jia-Wei Guo et al.· 0 citations
A novel chatbot architecture leveraging OpenAI's GPT-4 model for automated extraction and analysis of repository data is introduced, suggesting that this architecture can make repository data more accessible to technical and non-technical audiences through the production of actionable insights.
Muhammad Jawad Chowdhury, Md. Sakib Khan· 1 citation
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.