Skip to content
Open access

Building a Multilingual AI Legal Assistant Using Retrieval-Augmented Generation: A Case Study on the Legal System of Kazakhstan

Aug 2026 · Computers · Vol 15, pp. 556 · 0 citations · 38 references

TL;DR

Based on the proposed architecture and the selected semantic retrieval and language models, an AI legal assistant was developed and integrated into the “Adal Azamat” legal services platform providing users in Kazakhstan with practical access to AI-assisted legal consultation.

Abstract

This study presents the development of an Artificial Intelligence (AI)-based legal assistant using the Retrieval-Augmented Generation (RAG) architecture to provide legal assistance to citizens of the Republic of Kazakhstan. The proposed solution is designed to generate accurate, evidence-based responses to user queries using the regulatory legal acts of the Republic of Kazakhstan as the primary source of information. A legal corpus comprising 101,000 legislative documents and court decisions, with approximately 77 million tokens in Kazakh and Russian, was constructed to support the retrieval component of the system. To identify the most effective semantic retrieval method, three multilingual embedding models—Multilingual-E5-Large, BGE-M3, and KazEmbed-V5—were evaluated for vector search. The experimental results showed retrieval accuracies of 87.6%, 76.8%, and 83.3%, respectively. The GPT-5.4 and Llama-4-Scout-17B-16E-Instruct large language models were used to generate legal reasoning and responses based on documents retrieved through semantic search. The quality of the generated responses was evaluated using two complementary approaches. First, legal experts assessed the factual correctness and legal validity of the answers. Second, automatic evaluation was performed using word-level F1, BLEU, ROUGE, and BERTScore-F1 metrics. Among all evaluated configurations, GPT-5.4 combined with Multilingual-E5-Large achieved the highest overall accuracy (88.5%), whereas Llama-4-Scout-17B-16E-Instruct combined with KazEmbed-V5 achieved an accuracy of 83.6%. Based on the proposed architecture and the selected semantic retrieval and language models, an AI legal assistant was developed and integrated into the “Adal Azamat” legal services platform providing users in Kazakhstan with practical access to AI-assisted legal consultation.

Read PDF

Similar papers

Conference Aug 2026

A Retrieval-Augmented Generation Framework for Legal AI Assistance in Indian Domain

Legal research is a complex and time-consuming process that requires practitioners to analyze numerous statutes, case laws, judicial precedents, and legal interpretations from multiple sources. Although Large Language Models (LLMs) have enabled conversational legal assistants, they often generate hallucinated responses...

S. S, Raja Chellan · 0 citations
Open access Sep 2026

Insaafai: A Retrieval-Assisted Multilingual Legal Information and Document-Analysis Prototype for Indian Criminal Law

To Indian Criminal-Law Information Remains Difficult Due To Complex Legal Terminology, Lengthy Statutory Documents, And The Coexistence Of References To The Indian Penal Code (Ipc) And The Bharatiya Nyaya Sanhita (Bns), 2023. This Paper Presents Insaafai, A Retrieval-Assisted Web-Based System Developed To Make Criminal...

Siddhi Vishwakarma and Riddhi Vishwakarma · 0 citations
Preprint Aug 2026

Rhetorical-Role-Aware Retrieval-Augmented Generation for Legal Question Answering over Indian Supreme Court Judgments

This research paper proposes a Retrieval Augmented Generation framework that is specific to the legal field in order to assist interactive retrieval and reason about judgments from the Supreme Court of India and demonstrates strong performance on metrics including contextual recall and answer relevancy.

Sayed Ayaan Ahmed Sha, Sangeetha Sivanesan, A. Madasamy et al. · 0 citations
Open access Sep 2026

A Retrieval-Augmented Multilingual Chatbot for Intelligent College Information Services

The experimental results obtained were 92% response accuracy, 90% retrieval relevance, 1.8 s response time on average, and 91% user satisfaction, thus proving the efficiency of the suggested architecture in providing reliable and scalable academic information services.

Arun Babu P., P. Sree, Lingala Anusha et al. · 0 citations
Conference Aug 2026

Automated Argument Mining and Entity Recognition in Sri Lankan Legal Judgments Using Large Language Models

Access to legal information remains a significant challenge in low-resource jurisdictions, where judicial documents are often lengthy, unstructured, and difficult to interpret. This paper presents a unified framework for automated argument mining and named entity recognition in Sri Lankan legal judgments using Large La...

Sathmi Sansiluni Jayaratne, Ruvan Weerasinghe · 0 citations
Open access Sep 2026

A Comparative Study of RAG System Performance Using Milvus and Qdrant Vector Databases for Legal Information Retrieval and Question Answering in the Context of Thai Law

This study presents an empirical comparative evaluation of Retrieval-Augmented Generation (RAG) system performance for legal information retrieval and question answering in the context of Thai law. The main objective of this research is to analyze, evaluate, and compare the functional effectiveness of two prominent ope...

Chanon Chamnandechakun, Tha Bounthanh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.