Skip to content

Category

small language model

383 papers

#small language model Preprint Aug 2026

Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs

This work presents a production retrieve-then-match cascade that spends computation in proportion to difficulty: retrieval surfaces plausible matches, a lightweight text cross-encoder auto-resolves the high-confidence majority, and an agentic multimodal vision-language model settles the ambiguous remainder.

Jian Wang, Steven Xu, Sanjyot Thete et al. · 0 citations
#small language model Preprint Aug 2026

Belief Cascades Drive Persuasion in LLM Agent Networks

This work introduces a controlled testbed for studying how goal-directed persuaders shift elicited stances in networks of LLM agents grounded in real-world ego-network topologies, and argues for evaluating multi-agent persuasion as a trajectory- and exposure-level process.

Haoyi Qiu, Genglin Liu, P. Venkit et al. · 0 citations
#small language model Preprint Aug 2026

The Von-Neumann State-Space Transformer for neural decoding

A von-Neumann inspired hypothesis of efficient computation as an alternative for neural decoding, a memory-augmented Transformer whose feed-forward block is a low-rank instruction bank: a shared base operator plus a small set of learned low-rank instructions, from which a per-token code synthesizes the weight matrix actually used at that token.

Morteza Sarafyazd · 0 citations
#small language model Preprint Aug 2026

SHIFT-LLM: Distribution Shift Correction in Depth-Pruned LLMs

SHIFT-LLM, a training-free post-pruning correction framework that inserts a Linear Residual Adapter at each pruning site, consistently recovers accuracy lost to depth pruning across most configurations, achieving gains up to +15.7 points on Llama-3.1-8B-Instruct.

Ali Bahri, Hang Li, Hongliang Li et al. · 0 citations
#small language model Open access Aug 2026

Estimated Prevalence of Developmental Language Disorder Risk in West Virginia Schools.

PURPOSE Developmental language disorder (DLD) is linked to long-term academic and social difficulties. Despite these lasting impacts, state-level prevalence estimates are limited. This study estimated prevalence among early-elementary students in West Virginia who may be at risk for DLD and examined variation by grade, locale, and assessment instrument. METHOD A cross-sectional screening battery was administered to 252 students in Grades 1-2 across six Title I schools. Children with Kaufman Brief Intelligence Test-Second Edition Matrices standard scores ≤ 70 were excluded prior to analysis. A dual-criterion definition of DLD risk was used: ≤ 80 on the Clinical Evaluation of Language Fundamentals-Fifth Edition (CELF-5) core composite and ≤ 92 on the Test of Narrative Language-Second Edition (TNL-2). Descriptive statistics and mixed-effects logistic regression (random intercept for school) tested effects of grade, gender, and locale. RESULTS Using the dual-criterion definition, 72 of 252 students (28.6%) met criteria for being at risk for DLD. Adjusted models showed higher odds of DLD risk in town versus urban schools (odds ratio = 3.06, 95% confidence interval [1.31, 7.15], p = .010), whereas grade and gender were not significant predictors after accounting for school-level clustering. Instrument-specific classification rates diverged substantially, with the TNL-2 identifying a higher proportion of students relative to the CELF-5. CONCLUSIONS Although based on a relatively small yet representative sample of West Virginia students, prevalence estimates substantially exceed commonly cited large-scale population estimates. Results underscore the influence of instrument choice on suspected prevalence and the need for locally calibrated screening protocols and targeted service planning.

Megan Israelsen-Augenstein, Michelle Moore, Tracy Toman et al. · 0 citations
#small language model Preprint Aug 2026

Bolt-on, Verifiable Provenance for LLM-Powered Data Processing

BLIP is presented, a bolt-on framework for efficiently inferring a small-sized verifiable provenance for any LLM-powered data processing task, with any LLM, and eight strategies, each guaranteed to find a minimal verifiable provenance are introduced.

Yiming Lin, Sepanta Zeighami, Aditya G. Parameswaran · 0 citations
#small language model Preprint Aug 2026

Query Expansion Is More Than Generation: Improving Dense Retrieval through Better Integration

This work introduces AnchorQE, a training-free method that separately encodes the original query and its expansion before interpolating them, and shows that AnchorQE improves retrieval effectiveness by up to 12.89% when compared to widely-used expansion-only or text-level concatenation baselines across TREC-DL, LoTTE, and BEIR.

Sixia Sun, Mihai Surdeanu · 0 citations
#small language model Preprint Aug 2026

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

This work proposes SMART, an MLLM-guided temporal alignment framework for joint sign recognition and spotting that incorporates CSFormer, a CSLR-guided spotting module that injects recognition-derived gloss evidence into a boundary-aware spotting network.

Eunjee Choi, J. Sung, Seongwhan Cho et al. · 0 citations
#small language model Review Open access Aug 2026

Challenges and opportunities in type III secretion system effector prediction.

The conceptual evolution of T3SE prediction is reviewed, persistent limitations and sources of bias are highlighted, and open questions that must be addressed are outlined to enable robust, interpretable and ecologically inclusive prediction of T3SEs, pointing towards the need for centralized, user-friendly platforms that integrate diverse biological signals into transparent, ranked outputs suitable for experimental validation.

Iva Rosić, Ivan Nikolić · 0 citations
#small language model Preprint Aug 2026

Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More

ProViP is proposed, a training-free progressive visual token pruning framework that removes redundant visual tokens based on the embedding similarity of input tokens before reasoning of the LLM backbone, and then prunes tokens during reasoning via head-aware pruning.

Chaofang Ma, Lin Jiang, Carol Jingyi Li et al. · 0 citations
#small language model Preprint Aug 2026

Short Horizons and Sparse Concepts: a Mathematical View of the Readout in the J-lens

The Jacobian lens (J-lens) has been proposed as a way to read verbalizable representations from language models. However, its principle and meaning lack a detailed and theoretical discussion. We provide a mathematical view of this interpretation and of its assumed causal structure. Besides treating the J-lens as a heuristic probe, we further regard it as a first-order causal transfer operator from intermediate activations to expected future readouts. We study the Jacobian matrix as the optimal local linear approximation of the downstream mapping, analyze its global approximation behavior and bias, and identify its mathematical meaning as an expectation over anticipated future readouts. Further analysis of the Jacobian energy distribution reveals that its causal geometry is highly sparse. The energy decays with depth, concentrates in an extremely small proportion, and decomposes into diagonal pathways and specific critical positions. This decomposition further resolves the expectation of the J-lens over future outputs into short-horizon and sparse concept predictions, providing a more intuitive attribution and explanation for the ability of the J-lens to visualize concepts during the thinking process. Based on the theory, we propose a simple but effective improvement strategy and decoupling method for the J-lens, which significantly enhances the ability of the J-lens to read out correct intermediate concepts.

Shi-Qi Yan, Kai-Xuan Ding, Chao-Hong Tan et al. · 0 citations
#protein folding Review Aug 2026

Emerging experimental and computational methods for studying redox-regulated structural transitions.

Thiol-based redox switches utilize the unique nucleophilicity of cysteine and selenocysteine to dynamically link real-time cellular redox fluctuations to metabolic regulation and signaling pathways. Capturing the precise atomic-level thermodynamic and kinetic mechanisms driving these oxidative modifications has long been limited by the chemical instability of transient intermediates and the immense computational costs of classical molecular dynamics simulations. However, recent advancements in chemoselective small-molecule probes now allow for the high-purity trapping and enrichment of specific sulfenic and sulfinic acid states. In parallel, a paradigm shift toward machine learning, graph neural networks, and protein language models has bypassed traditional computational bottlenecks, enabling high-throughput, proteome-wide predictions of redox switches in seconds. Furthermore, emerging data reveal that these redox modifications do not merely alter well-structured proteins, but actively dictate conditional folding transitions and structural transformations within intrinsically disordered proteins and biomolecular condensates. Here, we provide a comparative overview of these dual experimental and computational advancements and highlight how the integration of generative diffusion models could facilitate the real-time simulation of conditional, multi-state structural ensembles across the redox proteome.

T. Rass, Dana Reichmann, Gábor Erdős · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.