Skip to content

Category

small language model

437 papers

#small language model Preprint Aug 2026

Decoupled Physical Modeling and Execution for Physics Reasoning

A unified framework is introduced that distills intermediate representations that explicitly encode the physical modeling process and adopt a two-stage post-training strategy, where supervised fine-tuning establishes structured modeling, and reinforcement learning with rubric-based feedback improves the quality of the modeling process.

Ye Zhang, Xuehang Guo, Rui Pan et al. · 0 citations
#small language model Preprint Aug 2026

PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding

PUMA (Polish Unified Multimodal Assessment) is proposed, a novel benchmark of 900 hand-crafted tasks designed to probe the limits of multimodal models in the Polish cultural and linguistic context and open-source the evaluation framework to advance localized multimodal AI research.

Sławomir Dadas, Michał Perełkiewicz, Rafal Poswiata et al. · 0 citations
#small language model Preprint Aug 2026

What actually runs: a measurement study of language model placement and decode speed on the Apple Neural Engine

This work sweeps a 64-shape matrix of LLM primitives that varies how a computation is expressed while holding what it computes fixed, recording per-operation device support and finds that placement is a property of how a computation is expressed, not of what it computes.

A. ShahirM · 0 citations
#small language model Preprint Aug 2026

Monitored free fermions under periodic driving

We investigate analytically and numerically a one-dimensional periodically driven free-fermionic system subjected to monitoring of the local particle density. Based on the analytical approach that describes the long-wavelength physics of the time-dependent Hamiltonian in the field-theoretical language using the nonlinear sigma-model (NLSM), we reveal that driving does not alter the universality class of the problem. As a consequence, the system retains the area-law behavior in the thermodynamic limit, with an intermediate diffusive regime giving rise to logarithmic growth of entanglement entropy for a small monitoring rate. At the same time, driving leads to a renormalization of the bare coupling constant of the NLSM, which controls the space-time ``conductivity''in the diffusive regime. We derive the analytic form of this renormalization, which becomes particularly strong in the case of a ``maximally symmetric''drive and sufficiently short driving period. In addition, we employ the Wiener-Hopf method to investigate the ballistic-diffusive crossover. These analytical predictions are corroborated by numerical simulations of the von-Neumann entanglement entropy and the density correlation function. Our numerical results clearly demonstrate that, with an increase in the system size, there are successive crossovers from ballistic to diffusive behavior and ultimately to localization. Furthermore, in the diffusive regime, we observe weak-localization corrections that are in agreement with the analytical predictions of the NLSM. Overall, our results provide a unified analytical and numerical framework for understanding the effects of monitoring in time-modulated fermionic systems, paving a way for broader investigations of driven quantum matter.

Aditi Chakrabarty, A. Mirlin, I. Poboiko · 0 citations

ProphDR: An Interpretable Deep Learning Model for Predicting Cancer Drug Response via Multi-Omics and Cross-Attention Mechanisms.

ProphDR is an interpretable deep learning framework that integrates multiomics data and drug structural information using a hierarchical attention mechanism, and generates biologically interpretable attention maps that highlight key pharmacophores and resistance-related genes consistent with established mechanisms in NSCLC and BRCA.

Yundian Zeng, Qing Ye, Jike Wang et al. · 0 citations
#small language model Open access Aug 2026

Are reasoning paradigms scale-aware? A cross-paradigm verification of prompting, retrieval, and knowledge-graph scaffolding for small language models

A scale-aware comparative study of reasoning enhancement for SLMs across three major families of methods: prompting-based reasoning, retrieval-based augmentation, and knowledge graph guided scaffolding shows that reasoning-enhancement strategies are not universally transferable across model scales under the evaluated settings.

Zhen-Zhen Gu, Jie Liu, Xian Liu · 0 citations
#small language model Open access Aug 2026

Knowledge Accuracy and Response Characteristics of an On-Device Large Language Model in Building Environmental Engineering: a Preliminary Case Study of EXAONE 4.0 1.2B Using ChatGPT-4o as a Cloud Reference

This preliminary case study examines EXAONE 4.0 1.2B, an on-device large language model (LLM), in building environmental engineering, using ChatGPT-4o as a high-capability cloud reference rather than a size-matched competitor. Ten Korean-language items were evaluated: four multiple-choice, three single-step calculations, and three short-answer questions covering indoor environmental quality, energy, and post-occupancy evaluation. The same simple prompt condition was applied to both models, with no model-specific prompt optimization; therefore, the results represent one standardized condition rather than each model’s maximum performance. Both models answered all seven objective items correctly, although EXAONE showed terminology confusion in its reasoning on a thermal-environment item. Five doctoral-level experts rated the short-answer responses for accuracy, completeness, and logical consistency. Mean scores were 4.18 for ChatGPT-4o and 2.98 for EXAONE. Fleiss’ kappa was 0.020 and 0.044, while free-marginal kappa was 0.333 and 0.222, respectively; the latter values indicate fair, not moderate, agreement, so absolute expert scores require cautious interpretation. Qualitatively, ChatGPT-4o more consistently organized responses around concepts and scope, whereas EXAONE tended to provide specific technical, operational, and Korean-regulatory content but occasionally omitted canonical elements or added unsupported detail. Because the item pool is small and only one on-device model was tested, the findings are exploratory and cannot be generalized to building environmental engineering as a whole or to on-device LLMs as a class. The results support further investigation of on-device LLMs as supervised offline assistants where connectivity or external data transmission is constrained, but not as unsupervised tools for engineering calculation or compliance decisions.

C. Cheong, Jeong-Hoon Lee · 0 citations
#small language model Preprint Aug 2026

Asymmetric Capacity Allocation in Self-Refinement Pipelines

It is concluded that larger generators and refiners generally improve the pipeline, whereas an undersized refiner can even harm performance, and that model capacity should not be allocated uniformly across self-refinement pipelines.

Zhuoyi Yang, Ian G. Harris, Salar Hashemitaheri et al. · 0 citations
#small language model Preprint Aug 2026

Distilling Black-Box Machine Learning into a Small, Self-Explaining Language Model for Learning Analytics

A two-stage fine-tuning pipeline is proposed that distills a fitted black-box estimator and its post hoc interpretation into a small, open-weight large language model (LLM) that returns an individual-level estimate and explains in natural language that predicts and explains offline on a commodity laptop, so student records never leave the machine.

Chenguang Pan, Airui Meng, Youmi Suk · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.