2026· SemEval@ACL· pp. 2729-2743· 0 citations· 16 references
Computer Science
TL;DR
A distribution-aware, LLM-augmented dataset was constructed by selectively paraphrasing minority-class instances to enhance class balance, and its performance was benchmarked against full, rebalanced, and undersampled training configurations.
Abstract
Detecting equivocation is essential, as indirect or evasive responses can shape public perception, influence political narratives, and undermine transparency in democratic discourse. To address the challenge of detecting evasive political responses on digital platforms, participation in the CLARITY SemEval-2026 Task was undertaken, which focuses on (i) clarity-level classification and (ii) fine-grained evasion-type classification in political question-answer contexts. This study introduces a data-centric framework that systematically examines the effects of class distribution and refinement strategies on the performance of Large Language Models (LLMs). A distribution-aware, LLM-augmented dataset was constructed by selectively paraphrasing minority-class instances to enhance class balance, and its performance was benchmarked against full, rebalanced, and undersampled training configurations. To comprehensively assess the proposed method, Qwen3-14B, Phi-4, Gemma-2 9B, and Mistral 7B were evaluated in in-context learning (ICL) settings (zero-shot and few-shot) and with LoRA fine-tuning. Experimental results indicate that fine-tuning Phi-4 with class rebalancing yields strong performance, achieving 74.77% on Subtask-1 and 51.55% on Subtask-2
Political evasion refers to responses that engage with a question while withholding the requested information. Recent NLP work frames political evasion as a classification task using a two-level taxonomy of response clarity and fine-grained evasion strategies. Existing work on response clarity and evasion classificatio...
This work evaluates prompting strategies for subtask 2 of the GermEval 2025 Harmful Content Detection challenge, which involves classifying whether a tweet attacks the free democratic basic order and shows that techniques such as Chain-of-Thought, In-Context Learning or Task Decomposition outperform approaches like Tas...
It is proposed that persona-based evaluation can serve as a scalable diagnostic of what generative systems value and prioritize when depicting humanity, and that persona generations are far from neutral.
N. Corrêa, Rafaela Weber Mallmann, David Kaczér et al.· Artificial Intelligence Revi...· 0 citations
This proposed approach combines cross-validation, structured aggregation and bias-aware evaluation to optimize the robustness–performance trade-off, and achieves 93.19% accuracy with a TCE of 3.13, yielding a strong combined score of 38.56 under the official evaluation metric.
The lack of high-quality labeled datasets remains a major challenge for sentiment analysis in low-resource languages such as Indonesian, particularly in specialized domains like fiscal policy. This study investigates the effectiveness of Large Language Models (LLMs) as automated annotators within a teacher-student know...
Novialdi Ashari, Ulfah Oktarida Sihaloho, Novi Aulia Sari· International Seminar on Int...· 0 citations
Aligned language models fail under two independent pressures: the structural jailbreak class recently formalized as Involuntary In-Context Learning (IICL), which reframes a harmful request as the final missing cell of a data-labeling task completed by pattern rather than judged as content; and the erosion of safety ali...
Tejasvi C. Addagada· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.