Skip to content
Open access

Detecting And Classifying Multi-Label Semantic Bias In 3,829 Indonesian Military Criminal Judgments Dataset Using Language Modeling And Ensemble Strategies

Aug 2026 · Jurnal Teknik Informatika (Jutif) · 0 citations

Abstract

Objectivity in military criminal judgments is crucial for judicial legitimacy but is frequently compromised by semantic bias. To the best of our knowledge, this is the first study to specifically address automated bias detection within the Indonesian military legal domain, bridging a significant gap in the literature that has predominantly focused on general civil law. This study aims to develop a multi-label classification model to automatically detect and classify three specific types of bias (emotional, character, and ambiguity) in military legal texts. The methodology involved the acquisition and expert annotation of 3,829 judgment documents (2020–2025). Three feature extraction strategies (TF-IDF, IndoBERT, and Doc2Vec) were comparatively evaluated using KNN, MLP, Random Forest, and Custom Ensemble algorithms. Experimental results demonstrate that the lexical approach using Random Forest with TF-IDF achieved superior performance with a weighted F1-Score of 0.82, outperforming both complex embedding-based models and the ensemble approach (F1-Score 0.77). The findings further reveal that character bias is the most dominant form of distortion in the corpus. This research makes three novel contributions: (1) providing the first annotated legal dataset for the Indonesian military domain; (2) demonstrating the superior efficacy of lexical features (TF-IDF) over complex embeddings in this specific legal domain; and (3) establishing a technical foundation for a decision-support system to enhance judicial objectivity.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.