Skip to content

Category

small language model

343 papers

#small language model Preprint Aug 2026

A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models

This work proposes a composite metric that combines two orthogonal criteria: information retention and throughput gains and finds that it allocates more resources to the most expressive layers compared to evolutionary search, specialized accelerators, or Shapley-value-based approaches that require expensive approximate inference.

A. Safronov · 1 citation · ⚡1
#small language model Preprint Aug 2026

The Pulse Beneath the Job Title: Monthly Readings of Requirements and Tasks from 750 Million Chinese Job Ads

How do we define an occupation? By its job title? An accountant at a small trading company keeps the books; at a listed firm the same title demands a certified-accountant licence, and the week goes to the reports that regulators and the board read. Same title, different bar, different work. What defines an occupation is who it lets in and what it asks them to do. In a rapidly changing labor market, tracking those requirements and tasks is how to take the market's pulse. Yet no instrument reads both at the speed they change. Official occupational directories like O*NET report one national average per occupation, updated every few years. Job postings are timely but unstructured. Research built on them works from job titles plus proprietary skill keywords, which blur what is asked of a candidate into what a candidate is asked to do. The blur matters, because rising requirements and changing tasks are different events with different causes. We separate them. From 752.6 million job ads posted on China's five leading recruitment platforms between 2022 and 2026, we extract the phrases employers write, unify those that name the same thing, and validate the mapping from text back to entry. By doing so we construct two catalogs, 20,721 requirements a candidate must meet and 44,479 tasks the hire will do. With the entries standardized, we annotate them further. Each task, for example, carries a score for how far a language model could absorb it. Matched back onto every ad, the catalogs read the market month by month. Two examples show what the layer beneath the job title buys. First, the occupational registry records one accountant where the ads record a staircase, the junior certificate at the bottom of the wage range and the intermediate one at the top. Second, counting occupations says the work most exposed to language models is disappearing, and counting tasks says far less of it is.

Qingqing Chen, Ying Fang, Xiang-Yu Wang et al. · 0 citations
#small language model Preprint Aug 2026

Ternary-Valued Finite-Difference Time-Domain Method: Equivalence with the Yee Scheme Through Noise-Shaped Quantisation

It is demonstrated that finite-difference time-domain dynamics can be reproduced with field variables restricted to the ternary alphabet, and its extension to acoustics, Virieux-type elastodynamics and Schrodinger-equation solvers points to a broader class of quantised physics solvers for resource-constrained and specialised hardware.

I. S. Maksymov · 0 citations
#small language model Preprint Aug 2026

Boosting LLM Exploration via Weak-Model Guidance in RLVR

This work empirically study the potential of outer prefixes, revealing the mechanism of the impact of distributional discrepancy to the exploration dynamics in RLVR training and efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting.

Xin Shen, Huishuai Zhang, Peng Li et al. · 0 citations
#small language model Preprint Aug 2026

Disentangling Optimization Scale from Preference Scale in DPO

This work shows that $\beta$ entangles two distinct roles: it governs the effective inverse preference-noise scale and simultaneously rescales the optimization dynamics, coupling this scale with the effective step size, and proposes a centered-softplus reformulation that is argmin-equivalent to DPO for $\beta>0$, while making the inverse preference-noise-scale and learning-rate effects explicit and independently tunable.

Ivan Kruzhilov · 0 citations
#small language model Preprint Aug 2026

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost.

Yu-Fan Wu, Yinghui He, Zhengyi Hu et al. · 1 citation
#small language model Preprint Aug 2026

Reservoir: A Large-Scale Simulated Dataset for Training and Evaluating Epidemiological Models

Large-scale, standardized datasets have driven many advances in AI-based scientific modeling, from protein structure prediction to natural language processing. Infectious disease epidemiology is increasingly adopting AI methods for forecasting, surveillance, and outbreak analytics, but the time-series data available to train them remains orders of magnitude smaller than the corpora behind the advances seen in other fields. Because the scope of real-world epidemiological data cannot practically reach the scale needed to train truly large-scale AI methods, simulated data provides a possible alternative. Here we introduce Reservoir, a large open simulator and dataset of realistic epidemic simulations in which every trajectory carries complete ground-truth labels, including quantities that cannot be measured directly in a real outbreak, such as true infection counts, time-varying reproduction numbers, and counterfactual intervention effects. Reservoir is generated by a stochastic simulator with realistic noise and reporting artifacts, together with interventions with configurable timing, compliance, and age-dependent efficacy. The current release contains 500,000 outbreak trajectories spanning one billion simulated days across diverse pathogen characteristics, population structures, and intervention regimes. Reservoir enables counterfactual experiments, surveillance-design studies, and training of epidemic models at a scale real-world datasets cannot provide.

Carson Dudley, Reiden Magdaleno, Marisa Eisenberg · 0 citations
#small language model Preprint Aug 2026

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

This work introduces Visual Retrieval Heads (VRHs), a small subset of attention heads that are causally responsible for grounding text descriptions to image regions, and shows that scoring attention from output prediction tokens with a sum over the ground-truth referent region most reliably identifies causal heads.

Chanho Park, Daehyeon Choi, Jihyun Lee et al. · 0 citations
#small language model Preprint Aug 2026

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

This report presents an open pretraining recipe that trains a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs, and derives a Puro Cost Scaling Law that relates training cost to average model performance.

Kairong Luo, Jia-Rui Cui, Yao-Rui Yin et al. · 0 citations
#small language model Open access Aug 2026

The Performance of Large Language Models in Extracting Intestinal Symptoms From Electronic Health Records: Retrospective Observational Study

This study provides a systematic comparison of several open-source LLMs on a structured intestinal symptom extraction task and concludes that Qwen3 models offer a favorable balance between accuracy and efficiency, making them suitable for resource-constrained scenarios.

Xinyue Zhang, Quanyu Wang, Beibei Liu et al. · 0 citations
#small language model Preprint Aug 2026

Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning

CMPM, a Chinese Multi-Panel Meme benchmark with 1,214 annotated samples covering five structural types, ordering dependency, panel-order constraints, and optional comment context, is introduced and results indicate that canonical-display accuracy is not by itself evidence of order understanding.

Hai-Han Li, Hai-Hao Li, Zhengjie Xu et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.