Skip to content

Category

small language model

437 papers

#small language model Preprint Aug 2026

Future Querying: Can LLMs Serve as Implicit Medical World Models?

This work introduces future querying, a paradigm that probes whether large language models can function as implicit medical world models by evaluating their ability to answer time-indexed clinical queries about a patient's future, and shows that small, locally fine-tuned open-weight models can match or approach larger proprietary systems, making the framework suitable for privacy-preserving, on-premise deployment.

Siri Willems, James Butterworth, L. Goetschalckx et al. · 0 citations
#small language model Preprint Aug 2026

Cross-Domain, Multi-Task Data-to-Text Generation without In-Domain Training Data

This work compares data-driven knowledge distillation (DDKD) against zero-shot inference and fine-tuning on out-of-domain D2T data, and introduces structure-preserving augmentation via structural subsampling and perturbation in cross-domain D2T generation.

Yifei Song, Kun Efimov-Zhang, Claire Gardent · 0 citations
#small language model Preprint Aug 2026

ProxyFormer: A Dual-Stream Proxy Architecture for Ultra-Long Context and High-Resolution Generation

This work proposes ProxyFormer, a general dual-stream architecture built upon proxy tokens, and introduces factorized multi-level compression/decompression, layer-wise dynamic compression ratios, asymmetric dual embeddings, and a proxy-only KV-cache inference scheme.

Zhongpan Tang · 0 citations
#small language model Preprint Aug 2026

StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models

This work proposes StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility, and experiments show that explicit strategy utility differs substantially across task categories and depends on both strategy generation and execution conditions.

Jinghan Tan, Yuanzheng Wang, Lu Chen et al. · 0 citations
#small language model Preprint Aug 2026

Systematic Bias in Green Patent Classification: Silent Green and False Green

Green-patent indicators based on Cooperative Patent Classification Y02 tags increasingly inform research, industrial policy, and climate-oriented investment, yet their construct validity has not been evaluated at corpus scale. We ask whether Y02 classification errors are random measurement noise or systematic, direction-specific bias. We introduce an Error-as-Signal framework in which disagreement between an administrative label and an independent model is treated as evidence of potential measurement error. Screening 9,075,421 USPTO granted patents from 1962-2024 with a fine-tuned domain model identifies 517,772 disagreements. Two independent open-weight large language models then assess whether each flagged invention has a direct climate-mitigation or adaptation function. Cross-model consensus identifies 180,384 administrative Type I errors (False Green) and 29,465 Type II errors (Silent Green). Correcting these errors reduces the measured green-patent population by 25.5%, from 592,387 to 441,468 patents. Misclassification is systematic rather than random. Atypicality predicts Silent Green in an inverted-U pattern, while reflection complexity independently increases under-recognition: controlling for atypicality and filing year, a one-standard-deviation increase is associated with 1.61 times the odds of Silent Green. Structural complexity has the opposite association. Among consensus-attributed errors, the same increase in reflection complexity is associated with 2.45 times the odds that an error is Silent Green rather than False Green. Event tests show no discrete rise in misclassification when green classification became more salient and only limited evidence of increased explicit green framing after the 2013 CPC launch. The evidence is more consistent with bounded classification capacity than with applicant gaming.

Hamid Bekamiri, Jan Auernhammer, Milad Abbasiharofteh et al. · 0 citations
#small language model Preprint Aug 2026

HelaBERT: Enhancing Sinhala Language Understanding with Dual Pooling Classification Head

We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourced from MADLAD-400, CulturaX, and a custom corpus comprising news articles, Sinhala Wikipedia, and web crawl data. HelaBERT-Small (~23.3M parameters, 6 layers) and HelaBERT-Large (~110M parameters, 12 layers) both use a SentencePiece Unigram tokenizer (vocabulary size 32,000) tailored to Sinhala's agglutinative morphology and complex script. We evaluate both models on four downstream Sinhala text classification tasks: news category classification, news source classification, sentiment analysis, and writing style classification, using 5 independent seed runs with stratified 80/20 train/test splits. We additionally propose a dual pooling classification head and evaluate it systematically across all four tasks, finding consistent improvements on sentiment analysis and a moderate gain on news category classification for HelaBERT-Small, while the standard [CLS]-linear head remains competitive on news source classification, a headline-level task with short average input length. We release both models to support further research in Sinhala NLP.

Thisen Ekanayake, Nisansa de Silva · 0 citations
#small language model Preprint Aug 2026

Predicting the scale limits of social mechanisms in agent societies

An audit is introduced that predicts a mechanism's fate as a population grows, asking how often the mechanism can act, whether agents use the information it supplies, and whether the measurement itself creates apparent scale effects.

Zengqing Wu, Chuan Xiao · 0 citations
#small language model Preprint Aug 2026

FormuEvo: LLM-Guided Evolution for Discovering Solver-Efficient Mixed-Integer Programming Formulations

A solver-informed diagnosis mechanism that exploits fine-grained solver statistics as verbal gradients for targeted refinement and a structured memory abstracts prior experience into reusable modeling strategies, avoiding redundant exploration while enabling zero-shot transfer to unseen problems and bootstrapping smaller LLMs.

Haofeng Yuan, Ji-Ming Peng, Jieyi Bi et al. · 0 citations
#small language model Preprint Aug 2026

A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification

A reproducible, license-aware knowledge-distillation recipe addressing the constraint of deploying a safety layer for large language models on commodity hardware by partitioning the corpus into seven safety categories aligned to a public hazard taxonomy.

Edson Rodrigues da Cruz Filho, Paulo Ricardo Ferreira Neves, P. H. Falsetti et al. · 0 citations
#small language model Preprint Aug 2026

From Subjective Judgments to Auditable Standards:Protocol-Guided AI Auditing of Website Redundancy

Website redundancy does not have a single fixed meaning. The same repeated element may distract during one task and provide backup during another. We introduce CORA (Counterfactual, Observable Redundancy Audit), which measures repetition load, normal-use tax, and failure-domain recovery reserve separately. Each run retains screenshots, stable element identities, and task traces. A versioned vision-language model proposes the annotations. Typed validation and release checks then determine whether a calibrated dimension can be reported; failed or malformed outputs stay in the fixed denominator. On a transparent mechanistic testbed, the factorized CORA representation separated reserve from normal-use tax and predicted perturbed success more accurately than scalar-load baselines. The model studies then showed why repeatability is not enough: two small local vision-language models produced recurring outputs, but neither instrument met all release requirements. CORA therefore withheld automated scores from both instruments while retaining the raw responses and failure records. Separate checker fixtures confirmed that the typed validator and hardened release gates implement their specifications; these tests do not establish semantic grounding or accuracy on production sites. Taken together, the results position CORA as an auditable candidate procedure for the controlled benchmark studied here rather than a general standard. Human agreement, AI-versus-human accuracy, and validation on independent production sites remain open empirical questions.

G. Kong, Yongtong Cao · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.