Skip to content

Category

small language model

442 papers

ProphDR: An Interpretable Deep Learning Model for Predicting Cancer Drug Response via Multi-Omics and Cross-Attention Mechanisms.

ProphDR is an interpretable deep learning framework that integrates multiomics data and drug structural information using a hierarchical attention mechanism, and generates biologically interpretable attention maps that highlight key pharmacophores and resistance-related genes consistent with established mechanisms in NSCLC and BRCA.

Yundian Zeng, Qing Ye, Jike Wang et al. · 0 citations
#small language model Open access Aug 2026

Are reasoning paradigms scale-aware? A cross-paradigm verification of prompting, retrieval, and knowledge-graph scaffolding for small language models

A scale-aware comparative study of reasoning enhancement for SLMs across three major families of methods: prompting-based reasoning, retrieval-based augmentation, and knowledge graph guided scaffolding shows that reasoning-enhancement strategies are not universally transferable across model scales under the evaluated settings.

Zhen-Zhen Gu, Jie Liu, Xian Liu · 0 citations
#small language model Open access Aug 2026

Knowledge Accuracy and Response Characteristics of an On-Device Large Language Model in Building Environmental Engineering: a Preliminary Case Study of EXAONE 4.0 1.2B Using ChatGPT-4o as a Cloud Reference

This preliminary case study examines EXAONE 4.0 1.2B, an on-device large language model (LLM), in building environmental engineering, using ChatGPT-4o as a high-capability cloud reference rather than a size-matched competitor. Ten Korean-language items were evaluated: four multiple-choice, three single-step calculations, and three short-answer questions covering indoor environmental quality, energy, and post-occupancy evaluation. The same simple prompt condition was applied to both models, with no model-specific prompt optimization; therefore, the results represent one standardized condition rather than each model’s maximum performance. Both models answered all seven objective items correctly, although EXAONE showed terminology confusion in its reasoning on a thermal-environment item. Five doctoral-level experts rated the short-answer responses for accuracy, completeness, and logical consistency. Mean scores were 4.18 for ChatGPT-4o and 2.98 for EXAONE. Fleiss’ kappa was 0.020 and 0.044, while free-marginal kappa was 0.333 and 0.222, respectively; the latter values indicate fair, not moderate, agreement, so absolute expert scores require cautious interpretation. Qualitatively, ChatGPT-4o more consistently organized responses around concepts and scope, whereas EXAONE tended to provide specific technical, operational, and Korean-regulatory content but occasionally omitted canonical elements or added unsupported detail. Because the item pool is small and only one on-device model was tested, the findings are exploratory and cannot be generalized to building environmental engineering as a whole or to on-device LLMs as a class. The results support further investigation of on-device LLMs as supervised offline assistants where connectivity or external data transmission is constrained, but not as unsupervised tools for engineering calculation or compliance decisions.

C. Cheong, Jeong-Hoon Lee · 0 citations
#small language model Preprint Aug 2026

Asymmetric Capacity Allocation in Self-Refinement Pipelines

It is concluded that larger generators and refiners generally improve the pipeline, whereas an undersized refiner can even harm performance, and that model capacity should not be allocated uniformly across self-refinement pipelines.

Zhuoyi Yang, Ian G. Harris, Salar Hashemitaheri et al. · 0 citations
#small language model Preprint Aug 2026

Distilling Black-Box Machine Learning into a Small, Self-Explaining Language Model for Learning Analytics

A two-stage fine-tuning pipeline is proposed that distills a fitted black-box estimator and its post hoc interpretation into a small, open-weight large language model (LLM) that returns an individual-level estimate and explains in natural language that predicts and explains offline on a commodity laptop, so student records never leave the machine.

Chenguang Pan, Airui Meng, Youmi Suk · 0 citations
#small language model Preprint Aug 2026

COEC: Calibrated Orthogonal-Equivalence Compensation for Structured Pruning of Large Language Models

COEC (Calibrated Orthogonal-Equivalence Compensation), a training-free compensation framework that applies alternating left and right orthogonal rotations to the retained weight, improves perplexity on every model and zero-shot accuracy in most settings over existing compensation methods.

Peiqi Yu, Nam Ling, Wei Wang et al. · 0 citations
#small language model Preprint Aug 2026

BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers

The results establish BF1 as a reproducible sparse operator and selective retrofit primitive with real long-context systems value, and evaluates numerical correctness, selected-interaction scaling, kernel performance, partial-model inference, and matched next-token language modeling.

Hina Dixit · 0 citations
#small language model Preprint Aug 2026

Jokes Aside: Measuring the Semantic Distance of Double Meanings

Large language models have significantly enriched the toolkit for computational humor research, particularly in the automated generation of jokes and puns. A key innovation, contextual embedding vectors, offers new opportunities to revisit and refine earlier hypotheses. Notably, Petrovic and Matthews (2013) proposed a joke generation model based on the scheme"I like my X like I like my Y, Z"(e.g."I like my ice like I like my dreams, crushed"). They suggested that joke hilarity increases with: a) frequent association of Z with X and Y, b) rarity of Z, c) ambiguity of Z, and d) meaning distance between X and Y. Building on this, Winters et al. (2019) proposed a set of metrics, based on Google Ngrams and Word2Vector. In this work, three out of their five metrics are revisited with word embeddings: obviousness, compatibility, and comparison. Another measure, symmetry, defined as closeness of Z to both X and Y, is introduced here for the first time. Two models were used to collect the embedding vectors (OpenAI text-embedding-3-small and MiniLM all-MiniLM-L6-v2) on three datasets: JokeJudger, Expunations, and rJokes. The last two datasets, Expunations, and rJokes, were expanded by adding paired sentences that captured the ambiguous expression at the core of each joke in its two different meanings. Results revealed that models trained on the proposed metrics performed poorly in predicting humor ratings: on JokeJudger, the best model achieved 57.1% accuracy, below the 61.5% baseline, while performance on Expunations and rJokes was even lower. Nevertheless, the symmetry metric seems consistently associated with higher-rated jokes, suggesting it may capture a necessary -though not sufficient- property of humor.

Fabio De Ponte · 0 citations
#small language model Preprint Aug 2026

Vibe Coding and Web Application Security: A Twin-Prompt Study

This work studies six functionally distinct web applications, each generated in two prompt variants that are identical except for an appended security-requirements section: a baseline (A) and a security-aware (B) variant.

Darko Andročec · 0 citations
#small language model Preprint Aug 2026

Interaction Effects Between Learner Characteristics and Dialogue Format in TTS Dialogue-Based Lessons

The results suggest that dialogue format should be selected according to learner characteristics in TTS dialogue-based lessons, with a significant interaction between learner characteristics and dialogue format for ARCS-based motivation.

Fumie Watanabe, Tota Suko, T. Ishida et al. · 0 citations
#small language model Preprint Aug 2026

Minimax Optimality of Score-Entropy Discrete Diffusion

This work establishes a minimax lower bound under the score-entropy loss, and proposes an MLE-based thresholding estimator that matches this lower bound up to constant and polylogarithmic factors that depend on neighboring density ratios.

Chol-Kyoon Cho, Yuchen Wu · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.