Findings indicate that realizing the benefits of agentic RAG depends on selecting models with sufficient tool-use propensity, as tool access alone did not guarantee performance gains in the authors' experiments.
Abstract
In civil law systems, legal professionals navigate sources of law hierarchically, searching for statutes, looking up specific articles, finding relevant cases, and examining full judgment texts in an iterative process. We present an agentic retrieval-augmented generation (RAG) architecture that mirrors this exploration workflow, enabling large language models (LLMs) to autonomously select and iteratively invoke five specialized legal tools implemented via the Model Context Protocol (MCP). On 300 multiple-choice questions from the 2025–2026 Korean Bar Examinations, our system achieved 95.33% accuracy with GPT-5.1 and 92.67% with Claude Sonnet 4.5, outperforming both Closed Book and Naïve RAG baselines, with the gains over Naïve RAG statistically significant (McNemar’s test, $p \lt .001$ ). However, effectiveness proved model-dependent: Gemini 2.5 Pro scored below its Naïve RAG baseline despite identical tool access. To explain this divergence, we analyzed tool-use behaviors and identified three distinct patterns: Intensive Tool Use (GPT), Efficient Utilization (Claude), and Tool Aversion (Gemini). GPT achieved high accuracy with broad, iterative case searches; Claude reached comparable performance with fewer searches, gaining accuracy through extended thinking rather than additional retrieval; and Gemini frequently avoided tool use altogether. An ablation study further revealed that case law tools were associated with the largest accuracy gains, followed by statute tools and full judgment text access. The magnitude of each contribution varied across models. These findings indicate that realizing the benefits of agentic RAG depends on selecting models with sufficient tool-use propensity, as tool access alone did not guarantee performance gains in our experiments.
We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one. Standard legal RAG treats the corpus as static; we argue legal question answering is a temporally-indexed retrieval problem....
This paper introduces SARA, an LLM‐powered platform deployed in a Brazilian court, which demonstrates significant efficiency and quality gains through the integration of LLM agents with knowledge models on jurisprudential and basic legal concepts.
V. Pinheiro, Francisco C. J. Bonfim, S. Silva et al.· The AI Magazine· 0 citations
Recent advances in Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) have significantly democratized access to legal information. Nevertheless, most existing legal assistants remain confined to multi-turn conversational QA, failing to support complex legal tasks that require systematic evidence retr...
Xiaoxia Cheng, Lin-Nan Wang, Jia-Hao Ma et al.· 0 citations
The application of Large Language Models (LLMs) to Thai legal question answering (QA) presents significant opportunities but faces challenges, particularly with complex legal reasoning, accurate citation, and the high cost of advanced alignment techniques. This thesis addresses these issues through a two-fold contribut...
Results show that structure‐aware retrieval combined with targeted information extraction can improve the reliability of AI systems for legal research and analysis.
Ali Hamza, Maaz Bin Rizwan, Maemoona Kayani et al.· Expert systems· 0 citations
This work proposes LexKairos, a comprehensive benchmark for evaluating the temporal capabilities of LLMs in the Chinese legal context across three dimensions: statutory temporal knowledge, case temporal modeling, and statute-case temporal reasoning.
Chenyang Li, Ze-Jia Feng, Yuqi Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.