Skip to content

Disentangling Statistical Preemption from Entrenchment in Language Models'Avoidance of Overgeneralization

Sep 2026 · 0 citations · 58 references
Computer Science

Abstract

How do learners avoid overgeneralizations such as Tom laughed me without explicit negative evidence? Constructionists have posited two proposals that describe indirect negative evidence against overgeneralizations: preemption (which privileges exposure to near-synonymous construction---e.g., she made him laugh) vs. entrenchment (all exposures to a verb's grammatical usages, including cases like He laughed). We disentangle these hypotheses by running controlled rearing experiments on LMs trained on child-caregiver conversations, where we systematically remove preemptive vs. non-preemptive evidence. We find that while LMs avoid overgeneralizations, they do not show preemption at a verb-specific level, instead showing weak but non-zero evidence of abstract preemption. Combined with results from analyzing the LMs'training dynamics, we find that LMs treat competing structures as indirect positive---as opposed to negative---evidence in the verb-specific condition. Insofar as preemption is the more plausible route to avoiding overgeneralizations in humans, our results point the need for there to be sensitivities to indirect negative evidence in neural network learners, and suggest new human experiments to test abstract preemption.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis

Large language models are increasingly asked to analyze data and report what the results mean, a task distinct from the belief- or preference-alignment settings studied in most sycophancy research. We test whether editorial framing in the prompt, ranging from a neutral request to an explicit instruction to search exhau...

P. Balani, Subhrakanta Panda · 0 citations
Preprint Aug 2026

Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm

Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional hierarchies, raising a question social psychology answered for humans six decades ago: how far will an agent escalate a harmful action when a legitimate authority insists? We port...

Hidayet Aksu · 0 citations
#machine learning Preprint Sep 2026

Tracing mechanisms of sycophantic agreement in language models

Sycophantic agreement in language models refers to the tendency to overly affirm a user's stated beliefs or preferences, often at the expense of factual accuracy. Although it is widely recognized as an alignment failure, its underlying mechanisms remain poorly understood. In this work, we use causal mediation analysis...

Si-Xing Chen, Zhuo-Fan Ying, L. Smith et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Memory vs. Context? Influential Factors of Factual Recall in Language Models

We reproduce and stress-test the work of Yu et al. (2023), who characterize how language models (LMs) arbitrate between memorized knowledge and contradictory in-context statements. We replicate their world-capitals experiments on 31 models spanning Pythia, GPT-2, Qwen3, and Ministral families, including base and post-t...

Guilhem Fouilhé, Nicholas Asher, Philippe Muller · 0 citations
#artificial intelligence Preprint Sep 2026

Knowing, and Saying It Only When Asked: LLM Endognostics and the Schizognosis of Minerva-7B

Evaluating an aligned language model by reading its answers assumes the answers carry the distinction the evaluator cares about. We introduce LLM endognostics, a white-box internal auditing framework designed to extract and causally manipulate latent knowledge within the residual stream. Applied to Minerva-7B-Instruct-...

Fabrizio Davide, Francesco Collova · 1 citation
2026

There Is No Spoon: Existential Presupposition in Large Language Models

It is found that while all models show sensitivity to existential presupposition across syntactic embeddings, determiner types and contextual cues, their behaviour differs markedly in strength and systematicity, with NLI-fine-tuned autoregressive models exhibiting the most coherent and stable projection patterns.

Marie-Léontine Wörgötter, Shiyang Lai, Sebastian Schuster · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.