It is argued that RCBM provides a promising framework for understanding the nature of cognition, and that it can be used to develop more sophisticated models of cognition in the future.
Teun van Gils, R. Sommers, M. Ostarek et al.· 0 citations
EmoLASP's LLM pipeline demonstrates the potential advantages of using a reasoning approach to ensure emotion prediction consistency and to reduce both the cost of fine-tuning and the cost of prompting with long dialogue histories.
Using human ideas as the AARs' initial research direction does not improve performance, suggesting current AARs may not need guidance from experienced researchers, and suggests that automating alignment research on well-characterized failures may be practical in the near term.
Yueh-Han Chen, Jia-Xin Wen, J. Kirchner· 0 citations
Cross-jurisdiction regulatory divergence detection is introduced: given an FDA requirement and an EMA requirement on the same topic, classify their relationship as AGREE, DIVERGE, or SILENT and three directional observations emerge at pilot scale.
Chu-Chu Wu, Zhi-Ying Zhou, Jing-Zhu Hu et al.· 2 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios, with one failure mode drawn from published scribe-error taxonomies.
Sebastian Fox, L. Markham, Ryan Lail et al.· 0 citations
It is asked whether judges detect omissions in clinical notes, and two methods reach it independently and trade off: a per-fact pipeline, and a GEPA-evolved prompt doing the same in one call.
Sebastian Fox, L. Markham, Ryan Lail et al.· 0 citations
To improve self-modeling skill, a scalable synthetic-data pipeline is developed that produces self-modeling training data, and reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks.
CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass, lowering the unit cost of AI-native education at scale.
Shangqing Tu, Daniel Zhang-Li, Yucheng Wang et al.· 0 citations
A multilingual German-English benchmark dataset that combines community-sourced stereotypes from German-speaking queer individuals with a German translation of WinoQueer is introduced, showing that language models reproduce anti-queer stereotypes, with variation across identities and models.
LLMs are increasingly used for interpersonal advice and as tools for studying social behavior across languages and cultures. A common shortcut for eliciting language- or culture-related variation is to ask a model to answer as a native speaker. We test whether this native-speaker persona reproduces the outputs obtained when models instead generate advice in the target language and translate the response back into English. Using 600 interpersonal advice questions across 13 languages and eight LLMs, we compare native-language generation followed by translation (NL) with native-speaker persona prompting (NP), measuring linguistic style, behavioral scaffolding, and forced-choice action recommendations. We find that NP and NL are not interchangeable. Compared to NL, NP often increases lexical social cues, including affiliation and positive tone, while reducing qualities such as concreteness and social attunement; NP also provides less actionable scaffolding in open-ended advice. In forced-choice scenarios, NP changes which action the model selects, favoring confrontation over redirection, with effect sizes varying across languages, topics, and models. Our results show that cross-lingual elicitation strategy is a consequential methodological choice that can change both how advice is framed and which actions models recommend.
N-PPI is proposed, a pointwise extension of Prediction-Powered Inference that calibrates model predictions at inference time as a lightweight post-hoc layer, without re-training the underlying model, to make accurate check-worthiness detection substantially cheaper to operate at scale.
This work proposes Long Chain-of-Thought Graph Verifier (LCoT-GV), a graph-based framework that represents LCoTs as reasoning graphs, each node in the graph represents a reasoning step and the edges encode semantic and logical relations.