Skip to content

Category

small language model

383 papers

#small language model Preprint Aug 2026

Most of the LLM Routing Gap Is Task Type

This paper argues that a small win does not show that routing did anything, the authors' or anyone else's, and argues that a small win does not show that routing did anything, theirs or anyone else's.

Janghoon Lee · 0 citations
#small language model Open access Aug 2026

The role of linguistic knowledge in verbal fluency tests: How individual differences in language skills shape the mental lexicon.

"List as many words as you can that start with M." The verbal fluency (VF) task is simple, yet even a typical university student only manages to produce about 15 words within 1 min, and there is substantial variability around this mean. The present study examined how linguistic and domain-general abilities contributed to VF performance in a large sample of healthy adult native speakers of Dutch (N = 571). We assessed the effects of linguistic knowledge, processing speed, short-term/working memory, and fluid intelligence on performance in the VF task. To examine whether linguistic and domain-general abilities contribute differently across VF task types, we included semantic trials (category-based generation: animals, food) and phonemic trials (letter-based generation: words beginning with S or M). We assessed the total number of correct words produced and the time to first response. Mixed-effects modeling showed that linguistic knowledge predicted the total number of correct responses in both semantic and phonemic VF. Short-term/working memory and processing speed were also significant predictors, but with smaller estimated effect sizes. Time to first response showed little effect of linguistic skills. We discuss how linguistic knowledge shapes the structure of the mental lexicon such that it affects both meaning-driven and form-driven access to lexical items. In addition, we provide updated norms for VF performance in Dutch and practical suggestions for using the task. (PsycInfo Database Record (c) 2026 APA, all rights reserved).

Kyla McConnell, Berit Reise, Antje S Meyer · 0 citations
#machine learning Preprint Aug 2026

Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

This work introduces a general steering technique called Semantic Overlays: small learned adapters applied at chosen prefill positions to a frozen model's residual stream that defends against the broad class of prompt injections that add instructions in untrusted context.

Joshua Penman · 0 citations
#small language model Preprint Aug 2026

Mixture of Channel Experts: Static Sparse Supports with Input-Adaptive Mixing for Pointwise Projections

This work introduces Mixture of Channel Experts (MoCE), a structured sparse channel-mixing layer, inspired by MoE, that replaces pointwise (1x1) channel-reduction projections and matches or exceeds dense baselines and prior channel-selection methods while reducing MACs by 16.7% and end-to-end latency.

Elian Iluk, Gil Ben-Artzi · 0 citations
#small language model Preprint Aug 2026

MARS: Multi-Specialist LLM Relay System for Competitive Programming

This work presents MARS (Multi-Agent Relay of Specialized LLMs), a prompt-only framework in which each agent is a topic specialist---dynamic programming, graphs, strings, geometry, and so on---grounded by retrieval-augmented generation over an algorithm-theory corpus.

Andrei Mikhailov, M. Burtsev, Alsu Sagirova · 0 citations
#small language model Preprint Aug 2026

Towards LLM-Enhanced Android Taint Analysis

Whether off-the-shelf Large Language Models (LLMs) can effectively reason about taint flows in Android apps is investigated, and preliminary findings suggest that LLM reasoning may effectively complement traditional static taint analysis.

Nicholas Miazzo, Marco Alecci, Jordan Samhi et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.