Skip to content

Category

artificial intelligence

6,274 papers

#artificial intelligence Preprint Aug 2026

Automated Researchers Can Mitigate Well-characterized Alignment Failures

Using human ideas as the AARs' initial research direction does not improve performance, suggesting current AARs may not need guidance from experienced researchers, and suggests that automating alignment research on well-characterized failures may be practical in the near term.

Yueh-Han Chen, Jia-Xin Wen, J. Kirchner · 0 citations
#artificial intelligence Review Jul 2026

RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciences

Cross-jurisdiction regulatory divergence detection is introduced: given an FDA requirement and an EMA requirement on the same topic, classify their relationship as AGREE, DIVERGE, or SILENT and three directional observations emerge at pilot scale.

Chu-Chu Wu, Zhi-Ying Zhou, Jing-Zhu Hu et al. · 2 citations
#artificial intelligence Preprint Aug 2026

Evaluating and Improving LLM Self-Modeling

To improve self-modeling skill, a scalable synthetic-data pipeline is developed that produces self-modeling training data, and reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks.

Si-Qi Zeng, André Assis, Rowan Wang · 0 citations
#artificial intelligence Preprint Aug 2026

CogEvol: Towards Efficient and Reliable Learning Environment Generation

CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass, lowering the unit cost of AI-native education at scale.

Shangqing Tu, Daniel Zhang-Li, Yucheng Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models

A multilingual German-English benchmark dataset that combines community-sourced stereotypes from German-speaking queer individuals with a German translation of WinoQueer is introduced, showing that language models reproduce anti-queer stereotypes, with variation across identities and models.

M. Morch, Daniel Braun · 0 citations
#artificial intelligence Preprint Aug 2026

Personas Differ from Native-Language Generation: Language Pathways Shape LLM Interpersonal Advice

LLMs are increasingly used for interpersonal advice and as tools for studying social behavior across languages and cultures. A common shortcut for eliciting language- or culture-related variation is to ask a model to answer as a native speaker. We test whether this native-speaker persona reproduces the outputs obtained when models instead generate advice in the target language and translate the response back into English. Using 600 interpersonal advice questions across 13 languages and eight LLMs, we compare native-language generation followed by translation (NL) with native-speaker persona prompting (NP), measuring linguistic style, behavioral scaffolding, and forced-choice action recommendations. We find that NP and NL are not interchangeable. Compared to NL, NP often increases lexical social cues, including affiliation and positive tone, while reducing qualities such as concreteness and social attunement; NP also provides less actionable scaffolding in open-ended advice. In forced-choice scenarios, NP changes which action the model selects, favoring confrontation over redirection, with effect sizes varying across languages, topics, and models. Our results show that cross-lingual elicitation strategy is a consequential methodological choice that can change both how advice is framed and which actions models recommend.

J. Won, Xinlan Emily Hu · 0 citations
#artificial intelligence Preprint Aug 2026

Calibrating Small Language Models for Claim Check-Worthiness Detection

N-PPI is proposed, a pointwise extension of Prediction-Powered Inference that calibrates model predictions at inference time as a lightweight post-hoc layer, without re-training the underlying model, to make accurate check-worthiness detection substantially cheaper to operate at scale.

Pratuat Amatya, Venktesh Viswanathan, Vinay Setty · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.