Skip to content

Category

artificial intelligence

6,274 papers

#artificial intelligence Preprint Aug 2026

Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text

This work constructs a pipeline in which a misaligned teacher model generates filtered synthetic datasets across domains such as creative writing and code generation, which are then used to fine-tune aligned student models, and shows that benign-looking synthetic data can act as a covert channel for transmitting targeted biases while largely preserving the student model's general task capabilities.

Minkyung Cho, Jihyo Kim, Seungwoo Song et al. · 0 citations
#artificial intelligence Review Aug 2026

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

An overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation is presented, covering spoken visual question answering and image-grounded hallucination detection in English and Modern Standard Arabic, and CRAI-Bench, evaluating the cultural accuracy of text-to-image generation.

Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Enhancing Low-Resource Language Reasoning via High-Resource Language Feature Transfer

A mechanistic intervention framework for identifying and transferring task-relevant sparse latent features across languages and reframes some cross-lingual reasoning gaps as failures of mechanism elicitation rather than capability absence, and offers a causally testable route to feature-mediated transfer without translation, fine-tuning, or changing the user-facing language.

Minju Song, Hyeon Hwang, Junhyun Lee et al. · 0 citations
#artificial intelligence Preprint Aug 2026

SemPOI-RL: Aligning LLM Semantic Reasoning for Interpretable Out-of-Town POI Sequential Generation

SemPOI-RL is proposed, a framework that aligns LLM semantic reasoning with structured sequence generation for interpretable OOT recommendation and consistently outperforms both traditional recommenders and direct LLM baselines, while providing interpretable style attribution across different phases of a trip.

Yunqi Liu, Yang Zhang, Ruixing Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Using Grounded Theory for Agent Behavior Analysis at Scale

This work proposes AutoTraceGT (Automated Trace analysis through Grounded Theory), the first multi-agent pipeline that automates grounded theory on agent trajectories and suggests Grounded Theory offers a scalable analytic tool for ML researchers and agent developers studying what agents actually do.

Zhuoran Lu, Yang-Yang Yu, Zhuoyan Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Stratified Consistency Distillation for Natural Language Formalization

A fine-tuning-based Stratified Consistency Distillation approach that shows significant and consistent improvements in both Pass@K and the novel Equivalent Logical Similarity metrics, demonstrating the potential of advancing logical translation through consistency distillation.

Zhi-Chao Hou, Ferhat Erata, Joseph Lilien et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs

Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be sensitive to small surface-level variations in problem formulation, raising questions about whether models truly follow the underlying logical structure. Studying this behavior is challenging because the symbolic components of logical problems, such as operators and predicates, are difficult to systematically manipulate in natural language. We introduce a tool-driven framework for generating controlled, label-preserving edits to logical reasoning problems. Our method operates on symbolic representations of first-order logic and constraint satisfaction problem tasks, enabling targeted modifications to logical operators and other structural components before translating them back into natural language. Using this framework, we evaluate various LLMs under cumulative and individual operator edits and analyze their behavior in response to these changes. Our quantitative and qualitative analyses show that LLM reasoning behavior under controlled operator edits is inconsistent, regardless of model size or family: models sometimes adapt correctly to structural changes but often fail to track their logical consequences. The results from this automated stress test enable an evaluation of language models across different dimensions and help measure the reliability of their reasoning.

Ramya Keerthy Thatikonda, W. Buntine, Ehsan Shareghi · 0 citations
#artificial intelligence Review Aug 2026

The Differential Reasoning Router: Operationalizing Cost-Aware LLM Annotation in E-commerce

This work introduces the Differential Reasoning Router (DRR), a cost-aware framework for cold-start LLM annotation that jointly optimizes model selection and human escalation, enabling a gradual shift from human-heavy cold-start annotation toward high-confidence automated routing.

Cheng Lyu, Jingyu Zhang, Vinny DeGenova et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Label Semantic Expansion via Label Guided Neural Topic Modeling

A Label-Guided Neural Topic Model (LGNTM) is proposed, which learns dedicated label-aligned topics, grounds them in lexical and document semantic spaces, and preserves consistency between topic structures and label structures.

Hao-Jia Zheng, Yuyin Lu, Jun-Tian Huang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

CPR for LLMs: Critical-Point Routing against Catastrophic Forgetting in Domain Adaptation

This work proposes CPR (Critical-Point Routing), a token-level routing framework between a base model and its expert derivative, based on critical tokens where the base model fails but the expert succeeds, and achieves state-of-the-art across all settings.

Kwangmin Ki, Yunhun Nam, Jongheon Jeong et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators

Over six months with four state-of-the-art LLM agents configured with web search, aggregate nowcast accuracy is broadly comparable to the institutional and professional benchmarks, with performance varying widely across individual indicators.

Xin-Yue Zhao, Ruiyi Zhang, Liqin Ye et al. · 0 citations
#artificial intelligence Preprint Aug 2026

AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP

AtlasNLP is introduced, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced, showing that dataset coverage is highly uneven across countries and tasks and language coverage does not imply geographic representation.

Joan Nwatu, Tsedeniya Solomon Amare, Longju Bai et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.