This work introduces Pak3H1, the first human-validated, culturally contextualized Urdu benchmark suite for 3H alignment, comprising PakAlpaca (helpfulness), PakBeaverTails (harmlessness), and PakTruthfulQA (honesty), and underscores the necessity of human-guided localization for equitable multilingual evaluation.
Abdullah Hashmat, Usman Naseem, Agha Ali Raza· 0 citations
The Cognitively Bounded User Simulator (CBUS) is introduced, an architectural framework that explicitly models the restricted working memory of young readers through an episodic bottleneck and shows that enforcing architectural constraints is more effective for high-fidelity simulation than simply scaling raw model capabilities.
Results show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation and show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation.
Amit Oren, N. Hertz-Palmor, Dean Ariel et al.· 0 citations
Large language models are increasingly deployed in Arabic-speaking markets, yet standard benchmarks overwhelmingly reward Modern Standard Arabic (MSA) fluency while leaving dialectal and culturally grounded competence unmeasured. This gap is consequential: everyday Arabic is largely dialectal, and dialect encodes social meaning that MSA-centric evaluation cannot capture. We present a rubric-based benchmark for the Saudi dialect, comprising 31 expert-authored prompts spanning idiomatic, pragmatic, lexical, and culturally-embedded phenomena, each paired with an expert-established ground truth. Our methodology separates evaluation into a model-agnostic phase, in which atomic, MECE positive criteria are derived solely from the ground truth, and a model-specific phase, in which four state-of-the-art systems -- Claude Opus 5, Gemini 3.7, GPT-5.6, and Kimi K3 -- are scored against those criteria and penalised for errors they actively introduce. Across 124 model-prompt evaluations we catalogue 466 error instances under a nine-category taxonomy. The four systems cluster within a narrow macro-average band (42.7%-53.1%), with no model exceeding 55% and every model recording at least one negative-scoring prompt, confirming that Saudi dialectal competence remains broadly unsolved. Notably, Ambiguous Framing is the dominant failure mode (37.3% of errors) while outright Hallucination accounts for only 11.2%, indicating that models fail less by stating falsehoods than by distorting register and flattening pragmatic nuance. We further observe a consistency-versus-ceiling trade-off and model-distinctive error signatures. We release the full prompt set, ground truths, and scored rubrics to support reproducible dialectal evaluation.
Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
A simple benchmark is built in which a single word is consistently substituted with another in the generation process, and two distinct axes are measured: metrics that relate to the model's surprise, as well as an evaluation of the textual reaction for 19 different open-weight language models.
IndicDetect provides standard data splits, an evaluation protocol, and baselines to establish a robust, language-aware foundation for AI-generated text detection in Indic scripts, and finds substantial robustness failures.
The rapid advancement of large language models (LLMs) has made AI-generated text detection increasingly critical. Existing zero-shot detectors assume that more token-level evidence leads to more reliable detection. However, our empirical study challenges this consensus: fewer tokens sometimes work better, retaining only 40% can yield optimal performance, yet this benefit is not universal. Using the Entropy Gap Score (EGS), we introduce top-$k$ cumulative probability filtering as a diagnostic probe. Across three representative settings, filtering exhibits strikingly different behaviors. We analyze EGS via typical set theory and quantify its dynamics through entropy calibration and distribution analysis. We find that filtering helps for weak source LMs, where low-entropy tokens are harmful, but fails for strong source LMs, where they are not notably harmful. Our work provides the first systematic analysis showing that some tokens are not merely uninformative but systematically harmful due to entropy miscalibration, revealing a two-sided trade-off in token-level detection.
This work proposes REIGN (Refurbished Embeddings with Integrated Guidance Networks), a contrastively trained bi-encoder that operates on sequences of contextualised chunk embeddings from a frozen Guidance Network rather than on raw tokens for document-to-document retrieval.
Across 12 evaluation settings covering the four principles of the Accountable-Professional Persona studied in this work, R$^2$A overall outperforms both the base model and static Persona elicitation and results show that Persona Representation Learning is critical for preventing Runtime Alignment from producing behaviorally imbalanced policies and for achieving more stable Persona policy learning.
Mo-Han Zhang, Chengsong You, Xiao-Yu Cao et al.· 0 citations
This work proposes MI-Distillation, a framework that constructs a continuous Instruct-Reasoning data spectrum through model interpolation, and introduces SeqLSS, which favors reasoning paths that are both informative and learnable for the student.
Yangsong Lan, Ren-Kai Hu, Hong-Kai Zheng et al.· 0 citations
A hybrid fact-checking framework that integrates Knowledge Graph-based semantic memory with adversarial multi-agent reasoning for explainable misinformation detection and a graph-aware confidence mechanism combines semantic similarity, NLI confidence, and structural graph evidence to determine whether internal knowledge is sufficient, thereby reducing unnecessary web retrieval.
Amelia Petrenciuc, Alexandru Lecu, Adrian Groza· 0 citations
SUP-MIMIC is proposed, a multi-task framework utilizing MIMIC-IV-v3.1 that comprises Basic Assessment, Diagnostic Divergence Task (DDT), and Diagnostic Convergence Task (DCT), designed to evaluate the model's"one-to-many"disambiguation capability among phenotypically similar cases, while DCT assesses the model's ability to identify"many-to-one" diagnostic patterns across different pathophysiological pathways