Skip to content

Category

artificial intelligence

4,637 papers

#artificial intelligence Preprint Aug 2026

Evaluating and Improving LLM Self-Modeling

To improve self-modeling skill, a scalable synthetic-data pipeline is developed that produces self-modeling training data, and reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks.

Si-Qi Zeng, André Assis, Rowan Wang · 0 citations
#artificial intelligence Preprint Aug 2026

CogEvol: Towards Efficient and Reliable Learning Environment Generation

CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass, lowering the unit cost of AI-native education at scale.

Shangqing Tu, Daniel Zhang-Li, Yucheng Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models

A multilingual German-English benchmark dataset that combines community-sourced stereotypes from German-speaking queer individuals with a German translation of WinoQueer is introduced, showing that language models reproduce anti-queer stereotypes, with variation across identities and models.

M. Morch, Daniel Braun · 0 citations
#artificial intelligence Preprint Aug 2026

Personas Differ from Native-Language Generation: Language Pathways Shape LLM Interpersonal Advice

LLMs are increasingly used for interpersonal advice and as tools for studying social behavior across languages and cultures. A common shortcut for eliciting language- or culture-related variation is to ask a model to answer as a native speaker. We test whether this native-speaker persona reproduces the outputs obtained when models instead generate advice in the target language and translate the response back into English. Using 600 interpersonal advice questions across 13 languages and eight LLMs, we compare native-language generation followed by translation (NL) with native-speaker persona prompting (NP), measuring linguistic style, behavioral scaffolding, and forced-choice action recommendations. We find that NP and NL are not interchangeable. Compared to NL, NP often increases lexical social cues, including affiliation and positive tone, while reducing qualities such as concreteness and social attunement; NP also provides less actionable scaffolding in open-ended advice. In forced-choice scenarios, NP changes which action the model selects, favoring confrontation over redirection, with effect sizes varying across languages, topics, and models. Our results show that cross-lingual elicitation strategy is a consequential methodological choice that can change both how advice is framed and which actions models recommend.

J. Won, Xinlan Emily Hu · 0 citations
#artificial intelligence Preprint Aug 2026

Calibrating Small Language Models for Claim Check-Worthiness Detection

N-PPI is proposed, a pointwise extension of Prediction-Powered Inference that calibrates model predictions at inference time as a lightweight post-hoc layer, without re-training the underlying model, to make accurate check-worthiness detection substantially cheaper to operate at scale.

Pratuat Amatya, Venktesh Viswanathan, Vinay Setty · 0 citations
#artificial intelligence Preprint Aug 2026

Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text

This work constructs a pipeline in which a misaligned teacher model generates filtered synthetic datasets across domains such as creative writing and code generation, which are then used to fine-tune aligned student models, and shows that benign-looking synthetic data can act as a covert channel for transmitting targeted biases while largely preserving the student model's general task capabilities.

Minkyung Cho, Jihyo Kim, Seungwoo Song et al. · 0 citations
#artificial intelligence Review Aug 2026

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

An overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation is presented, covering spoken visual question answering and image-grounded hallucination detection in English and Modern Standard Arabic, and CRAI-Bench, evaluating the cultural accuracy of text-to-image generation.

Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Enhancing Low-Resource Language Reasoning via High-Resource Language Feature Transfer

A mechanistic intervention framework for identifying and transferring task-relevant sparse latent features across languages and reframes some cross-lingual reasoning gaps as failures of mechanism elicitation rather than capability absence, and offers a causally testable route to feature-mediated transfer without translation, fine-tuning, or changing the user-facing language.

Minju Song, Hyeon Hwang, Junhyun Lee et al. · 0 citations
#artificial intelligence Preprint Aug 2026

SemPOI-RL: Aligning LLM Semantic Reasoning for Interpretable Out-of-Town POI Sequential Generation

SemPOI-RL is proposed, a framework that aligns LLM semantic reasoning with structured sequence generation for interpretable OOT recommendation and consistently outperforms both traditional recommenders and direct LLM baselines, while providing interpretable style attribution across different phases of a trip.

Yunqi Liu, Yang Zhang, Ruixing Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Using Grounded Theory for Agent Behavior Analysis at Scale

This work proposes AutoTraceGT (Automated Trace analysis through Grounded Theory), the first multi-agent pipeline that automates grounded theory on agent trajectories and suggests Grounded Theory offers a scalable analytic tool for ML researchers and agent developers studying what agents actually do.

Zhuoran Lu, Yang-Yang Yu, Zhuoyan Li et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.