Skip to content

Category

artificial intelligence

4,637 papers

#artificial intelligence Preprint Aug 2026

Stratified Consistency Distillation for Natural Language Formalization

A fine-tuning-based Stratified Consistency Distillation approach that shows significant and consistent improvements in both Pass@K and the novel Equivalent Logical Similarity metrics, demonstrating the potential of advancing logical translation through consistency distillation.

Zhi-Chao Hou, Ferhat Erata, Joseph Lilien et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs

Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be sensitive to small surface-level variations in problem formulation, raising questions about whether models truly follow the underlying logical structure. Studying this behavior is challenging because the symbolic components of logical problems, such as operators and predicates, are difficult to systematically manipulate in natural language. We introduce a tool-driven framework for generating controlled, label-preserving edits to logical reasoning problems. Our method operates on symbolic representations of first-order logic and constraint satisfaction problem tasks, enabling targeted modifications to logical operators and other structural components before translating them back into natural language. Using this framework, we evaluate various LLMs under cumulative and individual operator edits and analyze their behavior in response to these changes. Our quantitative and qualitative analyses show that LLM reasoning behavior under controlled operator edits is inconsistent, regardless of model size or family: models sometimes adapt correctly to structural changes but often fail to track their logical consequences. The results from this automated stress test enable an evaluation of language models across different dimensions and help measure the reliability of their reasoning.

Ramya Keerthy Thatikonda, W. Buntine, Ehsan Shareghi · 0 citations
#artificial intelligence Review Aug 2026

The Differential Reasoning Router: Operationalizing Cost-Aware LLM Annotation in E-commerce

This work introduces the Differential Reasoning Router (DRR), a cost-aware framework for cold-start LLM annotation that jointly optimizes model selection and human escalation, enabling a gradual shift from human-heavy cold-start annotation toward high-confidence automated routing.

Cheng Lyu, Jingyu Zhang, Vinny DeGenova et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Label Semantic Expansion via Label Guided Neural Topic Modeling

A Label-Guided Neural Topic Model (LGNTM) is proposed, which learns dedicated label-aligned topics, grounds them in lexical and document semantic spaces, and preserves consistency between topic structures and label structures.

Hao-Jia Zheng, Yuyin Lu, Jun-Tian Huang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

CPR for LLMs: Critical-Point Routing against Catastrophic Forgetting in Domain Adaptation

This work proposes CPR (Critical-Point Routing), a token-level routing framework between a base model and its expert derivative, based on critical tokens where the base model fails but the expert succeeds, and achieves state-of-the-art across all settings.

Kwangmin Ki, Yunhun Nam, Jongheon Jeong et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators

Over six months with four state-of-the-art LLM agents configured with web search, aggregate nowcast accuracy is broadly comparable to the institutional and professional benchmarks, with performance varying widely across individual indicators.

Xin-Yue Zhao, Ruiyi Zhang, Liqin Ye et al. · 0 citations
#artificial intelligence Preprint Aug 2026

AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP

AtlasNLP is introduced, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced, showing that dataset coverage is highly uneven across countries and tasks and language coverage does not imply geographic representation.

Joan Nwatu, Tsedeniya Solomon Amare, Longju Bai et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer

We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single-file C/CUDA training stack and a Devanagari-aware byte-level BPE tokenizer built for this project. On ARC-Easy and ARC-Challenge, Arkios exceeds three comparably sized open models (Pythia-1.4B, TinyLlama-1.1B, OLMo-1B) despite an order of magnitude fewer training tokens, likely aided by a match between our educational-web-text pretraining data and ARC's grade-school-science format rather than a general capability advantage. We report full evaluation results under standard protocols, including a correction to an earlier partial-sample estimate, and findings specific to evaluating small models in a low-resource language: the standard multiple-choice-letter prompt format used by common evaluation harnesses places this model at chance on Nepali reading comprehension, and simultaneously at chance on English in the same format, which would lead a naive benchmark run to conclude the model has no Nepali ability when in fact it does. Concretely, both languages score at chance in the letter-choice format (0.240 Nepali, 0.236 English, against a chance baseline of 0.250), while scoring the answer text directly reveals genuine, English-favoring comprehension (0.306 Nepali, 0.387 English). We describe a manifest-conditioned tool-use contract introduced during instruction tuning, where tool calls are permitted only when a tool manifest is declared in context and suppressed otherwise, and report where that contract holds and where it does not. We release both the base and instruction-tuned model weights under Apache-2.0. The training code and a small privately-sourced portion of the Nepali pretraining corpus are not released; everything needed to reproduce the reported numbers from the released weights is included here.

S. Regmi, Siddhartha Pudasaini, Chetan Phakami Pun · 0 citations
#artificial intelligence Preprint Aug 2026

Pak3H: Evaluating the Cost of Cultural Mismatch in LLM Alignment with a Human-Contextualized Urdu Benchmark

This work introduces Pak3H1, the first human-validated, culturally contextualized Urdu benchmark suite for 3H alignment, comprising PakAlpaca (helpfulness), PakBeaverTails (harmlessness), and PakTruthfulQA (honesty), and underscores the necessity of human-guided localization for equitable multilingual evaluation.

Abdullah Hashmat, Usman Naseem, Agha Ali Raza · 0 citations
#artificial intelligence Preprint Aug 2026

"Act Like a 5th Grader"is Not Enough: Bounding Knowledge in LLM-Based User Simulators

The Cognitively Bounded User Simulator (CBUS) is introduced, an architectural framework that explicitly models the restricted working memory of young readers through an episodic bottleneck and shows that enforcing architectural constraints is more effective for high-fidelity simulation than simply scaling raw model capabilities.

Krisztian Balog, A. M. Bakken · 0 citations
#artificial intelligence Preprint Aug 2026

Generating Clinical Vignettes that Preserve Cognitive Formulations

Results show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation and show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation.

Amit Oren, N. Hertz-Palmor, Dean Ariel et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

Large language models are increasingly deployed in Arabic-speaking markets, yet standard benchmarks overwhelmingly reward Modern Standard Arabic (MSA) fluency while leaving dialectal and culturally grounded competence unmeasured. This gap is consequential: everyday Arabic is largely dialectal, and dialect encodes social meaning that MSA-centric evaluation cannot capture. We present a rubric-based benchmark for the Saudi dialect, comprising 31 expert-authored prompts spanning idiomatic, pragmatic, lexical, and culturally-embedded phenomena, each paired with an expert-established ground truth. Our methodology separates evaluation into a model-agnostic phase, in which atomic, MECE positive criteria are derived solely from the ground truth, and a model-specific phase, in which four state-of-the-art systems -- Claude Opus 5, Gemini 3.7, GPT-5.6, and Kimi K3 -- are scored against those criteria and penalised for errors they actively introduce. Across 124 model-prompt evaluations we catalogue 466 error instances under a nine-category taxonomy. The four systems cluster within a narrow macro-average band (42.7%-53.1%), with no model exceeding 55% and every model recording at least one negative-scoring prompt, confirming that Saudi dialectal competence remains broadly unsolved. Notably, Ambiguous Framing is the dominant failure mode (37.3% of errors) while outright Hallucination accounts for only 11.2%, indicating that models fail less by stating falsehoods than by distorting register and flattening pragmatic nuance. We further observe a consistency-versus-ceiling trade-off and model-distinctive error signatures. We release the full prompt set, ground truths, and scored rubrics to support reproducible dialectal evaluation.

Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.