It is argued that WG is more plausible as an adversarial threat-requiring careful data engineering-rather than as a significant hazard inherent to routine fine-tuning.
Abstract
Narrow fine-tuning on small, domain-specific datasets can produce broad and surprising changes in model behavior-a phenomenon called weird generalization (WG). Yet, it remains unclear what features of the fine-tuning data are necessary for WG to arise. Here, we address this question by investigating a range of plausibly relevant features, including dataset size, composition, language, presentation style, and novelty relative to a model's parametric knowledge. Further, since WG evaluations rely on small question sets that assess the extent of the generalization, we also analyze how sensitive this measurement is to the set of questions used. Experiments with three open-weight models on four datasets show that the degree of WG (1) depends heavily on dataset composition and language (more than on size); (2) is greater for data familiar from pretraining than for novel data; and (3) is sensitive to the set of evaluation questions used. Collectively, these results indicate that WG is a product of quite fragile properties of both training and evaluation data. As such, we argue that WG is more plausible as an adversarial threat-requiring careful data engineering-rather than as a significant hazard inherent to routine fine-tuning.
Dataset Inference provides a robust framework for auditing data ownership by aggregating statistical signals that traditional Membership Inference Attacks (MIAs) fail to capture. While proven for LLM pre-training, its efficacy during fine-tuning is largely unexplored. We evaluate Dataset Inference on Gemma, Llama, and Qwen models using Full Fine-Tuning (FFT), LoRA, and QLoRA. Our findings reveal a stark architectural divergence: ParameterEfficient Fine-Tuning (PEFT) mitigates data leakage in Llama (AUC ≈ 0.50, p > 0.05), while Gemma and Qwen remain highly vulnerable across all adaptation methods (p < 0.05). Additionally, higher learning rates accelerate data absorption, and QLoRA provides only marginal regularization compared to standard LoRA.
Gabriel Vaz de Oliveira, Arthur Santos Viana de Oliveira, V. A. E. de Farias et al.· Anais do LIII Seminário Inte...· 0 citations
A lot of research attention has been devoted to checking whether large language models (LLMs) are politically biased. This work has largely focused on high-level ideological dimensions, such as left--right or progressive--conservative, and it has been shown that while LLMs are predominantly left and progressive leaning, largely mimicking the biases in the training data, they can be to some extent steered to change their preferences in post-training. In this short note, we check if LLMs have robust stances with regard to major substantive societal issues, on which members of the same ideological camp are often in disagreement, summarised in a novel dataset \textsc{HardChoices}. We show that, faced with this line of questioning, LLMs, both large and small, surprisingly rarely declare neutrality, are often incoherent, and demonstrate a remarkable degree of agreement on issues where they do take stances.
As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term Salience Bias: models become easily hijacked by useless explicit distractors (e.g., numerical values), leading them to ignore the implicit physical or commonsense prerequisites of a task. A critical open question is whether this failure reflects a genuine gap in commonsense knowledge or merely its suppression under misleading task framing. To investigate this, we construct the SaliTrap Benchmark, a high-quality dataset across four trap dimensions. Evaluating 12 state-of-the-art LLMs, we find that all mainstream models suffer significantly from salience bias, with severity scaling with distractor density and detecting the trap often decoupled from actually avoiding it. Crucially, by re-eliciting the same models with the task framing stripped away, we show that this is overwhelmingly a failure of \textbf{knowledge suppression rather than knowledge absence}: a context-free knowledge probe alone recovers over 90\% of sycophantic-compliance failures, revealing that the requisite commonsense is intrinsically present but actively crowded out by salient distractors that lure the model into over-compliant, unnecessary computation. Building on this diagnosis, we further show that lightweight, inference-time prompting alone substantially closes the gap without any retraining. Our findings relocate the bottleneck of commonsense reasoning failures from model competence to elicitation, and we release SaliTrap as a testbed for this blind spot. The codes are available at https://github.com/Wuzheng02/SaliTrap.
Zheng Wu, Chenhao Xue, Shijie Zheng et al.· 0 citations
: As small-scale, open-source Large Language Models (LLMs) proliferate for on-device and privacy-centric applications, understanding the trade-offs between their utility and behavioural reliability becomes critical. This study evaluates a suite of instruction-tuned LLMs, Gemma, Llama, and Qwen ( ≤4 B parameters), treating the model family, and scale as the primary units of analysis. Retrieval-Augmented Generation (RAG) is employed as a controlled experimental condition to assess utility gains on the Natural Questions (NQ) benchmark, while utilizing native, non-augmented configurations to establish a fairness baseline via the Bias Benchmark for QA (BBQ). The findings reveal that while RAG significantly enhances utility, often doubling Exact Match (EM) scores, these gains are non-uniform and architecture-dependent, with certain families exhibiting greater "retrieval-readiness" than others. Paradoxically, the fairness analysis shows that providing explicit context in disambiguated settings can increase stereotype engagement rather than suppressing it. These results suggest a fundamental disconnect between a model's capacity for factual accuracy and its ability to maintain social fairness, highlighting the need for multi-dimensional evaluation frameworks for small-scale systems.
M.J.F. Valdez, Arghir-Nicolae Moldovan· Proceedings of the 15th Inte...· 0 citations
This work induces misalignment by fine-tuning a Qwen2.5-14B-Instruct base model on nar-rowly misaligned data and tests Simple Self-Distillation as a method to recovering alignment in misaligned models.
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.