Skip to content

Category

small language model

343 papers

#small language model Open access Aug 2026

Estimated Prevalence of Developmental Language Disorder Risk in West Virginia Schools.

PURPOSE Developmental language disorder (DLD) is linked to long-term academic and social difficulties. Despite these lasting impacts, state-level prevalence estimates are limited. This study estimated prevalence among early-elementary students in West Virginia who may be at risk for DLD and examined variation by grade, locale, and assessment instrument. METHOD A cross-sectional screening battery was administered to 252 students in Grades 1-2 across six Title I schools. Children with Kaufman Brief Intelligence Test-Second Edition Matrices standard scores ≤ 70 were excluded prior to analysis. A dual-criterion definition of DLD risk was used: ≤ 80 on the Clinical Evaluation of Language Fundamentals-Fifth Edition (CELF-5) core composite and ≤ 92 on the Test of Narrative Language-Second Edition (TNL-2). Descriptive statistics and mixed-effects logistic regression (random intercept for school) tested effects of grade, gender, and locale. RESULTS Using the dual-criterion definition, 72 of 252 students (28.6%) met criteria for being at risk for DLD. Adjusted models showed higher odds of DLD risk in town versus urban schools (odds ratio = 3.06, 95% confidence interval [1.31, 7.15], p = .010), whereas grade and gender were not significant predictors after accounting for school-level clustering. Instrument-specific classification rates diverged substantially, with the TNL-2 identifying a higher proportion of students relative to the CELF-5. CONCLUSIONS Although based on a relatively small yet representative sample of West Virginia students, prevalence estimates substantially exceed commonly cited large-scale population estimates. Results underscore the influence of instrument choice on suspected prevalence and the need for locally calibrated screening protocols and targeted service planning.

Megan Israelsen-Augenstein, Michelle Moore, Tracy Toman et al. · 0 citations
#small language model Preprint Aug 2026

Bolt-on, Verifiable Provenance for LLM-Powered Data Processing

BLIP is presented, a bolt-on framework for efficiently inferring a small-sized verifiable provenance for any LLM-powered data processing task, with any LLM, and eight strategies, each guaranteed to find a minimal verifiable provenance are introduced.

Yiming Lin, Sepanta Zeighami, Aditya G. Parameswaran · 0 citations
#small language model Preprint Aug 2026

Query Expansion Is More Than Generation: Improving Dense Retrieval through Better Integration

This work introduces AnchorQE, a training-free method that separately encodes the original query and its expansion before interpolating them, and shows that AnchorQE improves retrieval effectiveness by up to 12.89% when compared to widely-used expansion-only or text-level concatenation baselines across TREC-DL, LoTTE, and BEIR.

Sixia Sun, Mihai Surdeanu · 0 citations
#small language model Preprint Aug 2026

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

This work proposes SMART, an MLLM-guided temporal alignment framework for joint sign recognition and spotting that incorporates CSFormer, a CSLR-guided spotting module that injects recognition-derived gloss evidence into a boundary-aware spotting network.

Eunjee Choi, J. Sung, Seongwhan Cho et al. · 0 citations
#small language model Review Open access Aug 2026

Challenges and opportunities in type III secretion system effector prediction.

The conceptual evolution of T3SE prediction is reviewed, persistent limitations and sources of bias are highlighted, and open questions that must be addressed are outlined to enable robust, interpretable and ecologically inclusive prediction of T3SEs, pointing towards the need for centralized, user-friendly platforms that integrate diverse biological signals into transparent, ranked outputs suitable for experimental validation.

Iva Rosić, Ivan Nikolić · 0 citations
#small language model Preprint Aug 2026

Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More

ProViP is proposed, a training-free progressive visual token pruning framework that removes redundant visual tokens based on the embedding similarity of input tokens before reasoning of the LLM backbone, and then prunes tokens during reasoning via head-aware pruning.

Chaofang Ma, Lin Jiang, Carol Jingyi Li et al. · 0 citations
#small language model Preprint Aug 2026

Short Horizons and Sparse Concepts: a Mathematical View of the Readout in the J-lens

The Jacobian lens (J-lens) has been proposed as a way to read verbalizable representations from language models. However, its principle and meaning lack a detailed and theoretical discussion. We provide a mathematical view of this interpretation and of its assumed causal structure. Besides treating the J-lens as a heuristic probe, we further regard it as a first-order causal transfer operator from intermediate activations to expected future readouts. We study the Jacobian matrix as the optimal local linear approximation of the downstream mapping, analyze its global approximation behavior and bias, and identify its mathematical meaning as an expectation over anticipated future readouts. Further analysis of the Jacobian energy distribution reveals that its causal geometry is highly sparse. The energy decays with depth, concentrates in an extremely small proportion, and decomposes into diagonal pathways and specific critical positions. This decomposition further resolves the expectation of the J-lens over future outputs into short-horizon and sparse concept predictions, providing a more intuitive attribution and explanation for the ability of the J-lens to visualize concepts during the thinking process. Based on the theory, we propose a simple but effective improvement strategy and decoupling method for the J-lens, which significantly enhances the ability of the J-lens to read out correct intermediate concepts.

Shi-Qi Yan, Kai-Xuan Ding, Chao-Hong Tan et al. · 0 citations
#protein folding Review Aug 2026

Emerging experimental and computational methods for studying redox-regulated structural transitions.

Thiol-based redox switches utilize the unique nucleophilicity of cysteine and selenocysteine to dynamically link real-time cellular redox fluctuations to metabolic regulation and signaling pathways. Capturing the precise atomic-level thermodynamic and kinetic mechanisms driving these oxidative modifications has long been limited by the chemical instability of transient intermediates and the immense computational costs of classical molecular dynamics simulations. However, recent advancements in chemoselective small-molecule probes now allow for the high-purity trapping and enrichment of specific sulfenic and sulfinic acid states. In parallel, a paradigm shift toward machine learning, graph neural networks, and protein language models has bypassed traditional computational bottlenecks, enabling high-throughput, proteome-wide predictions of redox switches in seconds. Furthermore, emerging data reveal that these redox modifications do not merely alter well-structured proteins, but actively dictate conditional folding transitions and structural transformations within intrinsically disordered proteins and biomolecular condensates. Here, we provide a comparative overview of these dual experimental and computational advancements and highlight how the integration of generative diffusion models could facilitate the real-time simulation of conditional, multi-state structural ensembles across the redox proteome.

T. Rass, Dana Reichmann, Gábor Erdős · 0 citations
#small language model Open access Aug 2026

CARE: conflict-aware regulation and evolution for continual test-time open-vocabulary semantic segmentation

A continual test-time adaptation framework that updates only a small subset of model parameters that consistently outperforms existing TTA and OVSS adaptation baselines across multiple datasets and corruptions, while maintaining stable performance over extended and cross-corruption continual streams without introducing additional trainable modules.

Fen Luo, Sen Li, Zhanqiang Huo · 0 citations
#small language model Open access Aug 2026

What students ask matters: LLM interaction depth, task quality, and immediate recall in higher education

The findings indicate a dissociation between performance quality and short-term recall in LLM-supported study, which aligns with cognitive-psychology evidence that elaboration improves comprehension while retrieval practice consolidates retention.

V. Tsiligkiris · 0 citations
#small language model Preprint Aug 2026

PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos

PhysMLLMs is a training-stage prior injection architecture that injects physics-inspired spatial continuity priors into Video MLLMs, demonstrating that the injected spatial prior improves video consistency without compromising image-level grounding or general multimodal capability.

Siyao Yan, Bo Han, Jisheng Dang et al. · 0 citations
#small language model Preprint Aug 2026

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems, and the 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form.

Apodex Team B. An, B. Li, B. Wang et al. · 1 citation

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.