Skip to content

Category

artificial intelligence

4,653 papers

#artificial intelligence Preprint Aug 2026

Pak3H: Evaluating the Cost of Cultural Mismatch in LLM Alignment with a Human-Contextualized Urdu Benchmark

This work introduces Pak3H1, the first human-validated, culturally contextualized Urdu benchmark suite for 3H alignment, comprising PakAlpaca (helpfulness), PakBeaverTails (harmlessness), and PakTruthfulQA (honesty), and underscores the necessity of human-guided localization for equitable multilingual evaluation.

Abdullah Hashmat, Usman Naseem, Agha Ali Raza · 0 citations
#artificial intelligence Preprint Aug 2026

"Act Like a 5th Grader"is Not Enough: Bounding Knowledge in LLM-Based User Simulators

The Cognitively Bounded User Simulator (CBUS) is introduced, an architectural framework that explicitly models the restricted working memory of young readers through an episodic bottleneck and shows that enforcing architectural constraints is more effective for high-fidelity simulation than simply scaling raw model capabilities.

Krisztian Balog, A. M. Bakken · 0 citations
#artificial intelligence Preprint Aug 2026

Generating Clinical Vignettes that Preserve Cognitive Formulations

Results show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation and show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation.

Amit Oren, N. Hertz-Palmor, Dean Ariel et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

Large language models are increasingly deployed in Arabic-speaking markets, yet standard benchmarks overwhelmingly reward Modern Standard Arabic (MSA) fluency while leaving dialectal and culturally grounded competence unmeasured. This gap is consequential: everyday Arabic is largely dialectal, and dialect encodes social meaning that MSA-centric evaluation cannot capture. We present a rubric-based benchmark for the Saudi dialect, comprising 31 expert-authored prompts spanning idiomatic, pragmatic, lexical, and culturally-embedded phenomena, each paired with an expert-established ground truth. Our methodology separates evaluation into a model-agnostic phase, in which atomic, MECE positive criteria are derived solely from the ground truth, and a model-specific phase, in which four state-of-the-art systems -- Claude Opus 5, Gemini 3.7, GPT-5.6, and Kimi K3 -- are scored against those criteria and penalised for errors they actively introduce. Across 124 model-prompt evaluations we catalogue 466 error instances under a nine-category taxonomy. The four systems cluster within a narrow macro-average band (42.7%-53.1%), with no model exceeding 55% and every model recording at least one negative-scoring prompt, confirming that Saudi dialectal competence remains broadly unsolved. Notably, Ambiguous Framing is the dominant failure mode (37.3% of errors) while outright Hallucination accounts for only 11.2%, indicating that models fail less by stating falsehoods than by distorting register and flattening pragmatic nuance. We further observe a consistency-versus-ceiling trade-off and model-distinctive error signatures. We release the full prompt set, ground truths, and scored rubrics to support reproducible dialectal evaluation.

Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?

A simple benchmark is built in which a single word is consistently substituted with another in the generation process, and two distinct axes are measured: metrics that relate to the model's surprise, as well as an evaluation of the textual reaction for 19 different open-weight language models.

A. Cetoli · 0 citations
#artificial intelligence Preprint Aug 2026

IndicDetect: Evaluating Cross-Lingual LLM-Generated Text Detection for Hindi, Telugu, and Tamil

IndicDetect provides standard data splits, an evaluation protocol, and baselines to establish a robust, language-aware foundation for AI-generated text detection in Indic scripts, and finds substantial robustness failures.

Bhaskar Ganesh Devalla, Junchao Wu, Nilesh Dokuparthi et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

When Less is More: Understanding When Token Filtering Helps and Fails in AI-generated Text Detection

The rapid advancement of large language models (LLMs) has made AI-generated text detection increasingly critical. Existing zero-shot detectors assume that more token-level evidence leads to more reliable detection. However, our empirical study challenges this consensus: fewer tokens sometimes work better, retaining only 40% can yield optimal performance, yet this benefit is not universal. Using the Entropy Gap Score (EGS), we introduce top-$k$ cumulative probability filtering as a diagnostic probe. Across three representative settings, filtering exhibits strikingly different behaviors. We analyze EGS via typical set theory and quantify its dynamics through entropy calibration and distribution analysis. We find that filtering helps for weak source LMs, where low-entropy tokens are harmful, but fails for strong source LMs, where they are not notably harmful. Our work provides the first systematic analysis showing that some tokens are not merely uninformative but systematically harmful due to entropy miscalibration, revealing a two-sided trade-off in token-level detection.

Xiaoyang Han, Lvxiaowei Xu, Ming Cai · 0 citations
#artificial intelligence Preprint Aug 2026

REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling

This work proposes REIGN (Refurbished Embeddings with Integrated Guidance Networks), a contrastively trained bi-encoder that operates on sequences of contextualised chunk embeddings from a frozen Guidance Network rather than on raw tokens for document-to-document retrieval.

Devrim Cavusoglu, Emre Akbas · 0 citations
#artificial intelligence Preprint Aug 2026

R$^2$A: Learning Persona Policies Through Persona Representation Learning and Runtime Alignment

Across 12 evaluation settings covering the four principles of the Accountable-Professional Persona studied in this work, R$^2$A overall outperforms both the base model and static Persona elicitation and results show that Persona Representation Learning is critical for preventing Runtime Alignment from producing behaviorally imbalanced policies and for achieving more stable Persona policy learning.

Mo-Han Zhang, Chengsong You, Xiao-Yu Cao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for Chain-of-Thought Distillation

This work proposes MI-Distillation, a framework that constructs a continuous Instruct-Reasoning data spectrum through model interpolation, and introduces SeqLSS, which favors reasoning paths that are both informative and learnable for the student.

Yangsong Lan, Ren-Kai Hu, Hong-Kai Zheng et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Memory-First Fact-Checking: A Knowledge-Graph-Grounded Multi-Agent System for Misinformation Detection

A hybrid fact-checking framework that integrates Knowledge Graph-based semantic memory with adversarial multi-agent reasoning for explainable misinformation detection and a graph-aware confidence mechanism combines semantic similarity, NLI confidence, and structural graph evidence to determine whether internal knowledge is sufficient, thereby reducing unnecessary web retrieval.

Amelia Petrenciuc, Alexandru Lecu, Adrian Groza · 0 citations
#artificial intelligence Preprint Aug 2026

SUP-MIMIC: A Multi-Task Clinical Diagnosis Benchmark for Evaluating LLMs'Robustness to Contradictory Evidence

SUP-MIMIC is proposed, a multi-task framework utilizing MIMIC-IV-v3.1 that comprises Basic Assessment, Diagnostic Divergence Task (DDT), and Diagnostic Convergence Task (DCT), designed to evaluate the model's"one-to-many"disambiguation capability among phenotypically similar cases, while DCT assesses the model's ability to identify"many-to-one" diagnostic patterns across different pathophysiological pathways

Yi Yu, Bo Wang, Chong Feng et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.