Skip to content

Cost-Effective Hate Speech Detection in Portuguese Using Lightweight LLMs

· 0 citations · 25 references

TL;DR

This work evaluates the efficiency and competitiveness of seven smaller, more accessible LLMs through zero-shot classification with structured instructions, leveraging the expert-annotated test dataset and annotation scheme of the kNOwHATE project to demonstrate a viable, competitive, and low-cost approach that does not rely on large-scale infrastructure or expensive proprietary models.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu

It is challenging to detect hate speech in Low Resource Languages (LRLs) because of the absence of annotated data, the informality of its language structure, and the lack of standardized grammar. A good example of such a challenge is Roman Urdu which is broadly used by South Asians on social media and has a high variat...

Toneema Zubair, Muhammad Asif, F. Kamiran et al. · 0 citations
Preprint Aug 2026

Hate Speech Classification In Roman Urdu: A Comparative Study On Parameter Efficient Fine-Tuning And Prompt Engineering

Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress and negative societal impacts. Ro-man Urdu, a low-resource language used in Pakistan and among Urdu-speaking communities worldwide, presents additional challenges becaus...

Toneema Zubair · 0 citations
#computer vision Preprint Sep 2026

MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos

Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced si...

Itzel Tlelo-Coyotecatl, Hugo Jair Escalante · 0 citations
Open access Aug 2026

Indonesian Hate Speech Detection Across Diverse Domains Using Parameter-Efficient Fine-Tuning with IndoBERT and LoRA

The findings indicate that IndoBERT+LoRA provides a promising and resource-efficient approach for multi-domain Indonesian hate and abusive speech classification, while stricter leave-one-domain-out evaluation remains an important direction for future work.

Fergie Joanda Kaunang, Bhustomy Hakim, A. P. Thenata · 0 citations
#natural language process... Preprint Sep 2026

Benchmarking Automatic Speech Recognition Tools for Iberian Languages

Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Catalan, Galician, Por...

Fernando López, Pablo Gómez, David Solans et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.