Skip to content
Preprint

PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary

Aug 2026 · 0 citations · 44 references
Computer Science

TL;DR

This work presents PROSLEX (PRediction Of Statutes and LEgal eXplanation), a comprehensive dataset comprising 1,623 expert-annotated legal documents from the Indian context, positioning PROSLEX as a benchmark for developing explainable AI systems that can support legal practitioners while advancing research in interpretable legal NLP.

Abstract

Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research. While recent advances have begun incorporating Large Language Models (LLMs) for statute prediction, current approaches primarily focus on accuracy metrics without addressing the critical need for legal reasoning, a fundamental requirement in judicial contexts where decisions must be explainable and justifiable. To address this research gap, we present PROSLEX (PRediction Of Statutes and LEgal eXplanation), a comprehensive dataset comprising 1,623 expert-annotated legal documents from the Indian context. Each document is paired with statute predictions and detailed explanations, totaling 7,450 explanations, capturing the underlying legal reasoning. Using this dataset, we systematically evaluate various prompting strategies, including zero-shot, few-shot, chain-of-thought, and tree-of-thoughts approaches, to generate both statute predictions and their corresponding legal rationales. Our evaluation framework measures not only predictive performance but also the coherence and legal validity of generated explanations, positioning PROSLEX as a benchmark for developing explainable AI systems that can support legal practitioners while advancing research in interpretable legal NLP. To ensure reproducibility, we have made our PROSLEX dataset and model code available on GitHub: https://github.com/subinay494/Legal_Statute_Prediction_Explanation.

View source

Similar papers

Preprint Aug 2026

ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts

This paper introduces the task of identifying and segmenting legal conditions (Tatbestand) and legal consequences (Rechtsfolge) within German statutory texts and presents ANNOTARES (Annotations of Tatbestand-Rechtsfolge Sequences), a novel dataset comprising German law texts with span-level annotations.

R. Schwarz, Jannik Strötgen · 0 citations
Open access Jul 2026

An Explainable AI-Based Legal Research Assistant for Precedent Analysis and Judicial Outcome Prediction Using Dense Retrieval and Transformer Models

With more than 45 million cases awaiting disposal across Indian courts as of 2024, the judicial system faces an acute need for faster, smarter tools to support legal research. This work introduces an artificial-intelligence-driven legal research assistant tailored to the jurisprudence of the Supreme Court of India. Sta...

Krish P. Gokhale, Tanvi Kshirsagar, Anugraha Kasbe et al. · 0 citations
#natural language process... Preprint Aug 2026

JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction

Juris Policy Optimization (JPO), a post-training framework for structured legal reasoning in Chinese criminal judgment prediction, is proposed and experiments show that JPO consistently improves both judgment prediction and reasoning quality over supervised fine-tuning and reinforcement learning baselines.

Zhao-Lu Kang, Yan-Tao Liu, Tailong Luo et al. · 1 citation
Preprint Aug 2026

LexKairos: Benchmarking Legal Temporal Capabilities in LLMs

This work proposes LexKairos, a comprehensive benchmark for evaluating the temporal capabilities of LLMs in the Chinese legal context across three dimensions: statutory temporal knowledge, case temporal modeling, and statute-case temporal reasoning.

Chenyang Li, Ze-Jia Feng, Yuqi Huang et al. · 0 citations
Review Aug 2026

Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law

We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one. Standard legal RAG treats the corpus as static; we argue legal question answering is a temporally-indexed retrieval problem....

Rose Cymbler, D. Guez, Laurent Fabre · 0 citations
Review Open access Aug 2026

NLP-Driven Extraction of Key Features from Legal Texts: Court Opinions, Briefs, Statutes, and Case Law

This paper outlines a unique method of legal text processing using Natural Language Processing (NLP) technology to extract the information from the legal texts meaningfully and naturally. The proposed system is designed in a data pipeline architecture by integrating the NLP functionalities such as tokenization, part-of...

S. A. Gade, Sivaram Ponnusamy · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.