Skip to content

Category

small language model

343 papers

#small language model Preprint Aug 2026

Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance

The effect of introducing the manager-worker scaffold over a shared filesystem workspace, with no training and no per-benchmark tuning, measured against the same model answering in a single pass is investigated, finding several mechanisms behind the gains.

Victor Gao, Vida Khosrowshahi, Ali Khosrowshahi et al. · 0 citations
#small language model Preprint Aug 2026

FOCUS&RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models

A token-level analysis of this failure mode is presented by viewing decoding as a dynamical process that enters and persists in a small set of recurrent contexts and shows that persistence is controlled by the escape mass assigned to plausible alternatives within the token sampling set.

Junyoung Lee, Se-Hee Park, Shinhyoung Jang et al. · 0 citations
#small language model Preprint Aug 2026

Planting a Latent Variable in Natural-Looking Text: a More Realistic Test of Belief States in LLMs and Their Link to Concept Geometry

This work plants a controllable latent variable inside natural-looking text and arranges the 8 states themselves on a ring, in the exact order of the Markov chain, which is supporting evidence that a concept's geometry can be formed by the statistical dynamics of the latent variable behind it.

Alexandru-Iulius Jerpelea · 0 citations
#small language model Open access Aug 2026

Intent Drift in LLM-Assisted Brain Computer Interface Communication: An In-Silico Benchmark Under Simulated Decoder Corruption

Background Large language models are increasingly proposed to post-edit decoded text in communication brain-computer interfaces and augmentative communication. A fluent model can substitute a different intent than attempted (intent drift). Whether meaning survives or confidence flags failure is unmeasured. Methods In-silico benchmark of 20 open-weight models post-editing text (4,252,326 labeled generations) corrupted with an empirical P300 confusion matrix at five levels (0-40% character error rate, CER) across the ALS message-banking vocabulary (AUTH), a message-critical probe set, and matched controls. Outputs were scored faithful, degraded, or drift by an ensemble benchmarked against physicians. A substudy re-ran 562 messages under six interface policies (seven-model panel). Findings Detected drift rose steeply with corruption in all three corpora, from 2.2% to 60.3% at 0-40% target CER in AUTH, a stress-test upper bound, not an expected clinical rate (odds ratio 2.30 per 10-percentage-point rise in target CER). Stated confidence discriminated faithful outputs reasonably well (AUROC 0.83, 0.80-0.85) but was poorly calibrated (expected calibration error 0.32, 0.27-0.37): 28.4% of outputs at confidence 90 or higher were not faithful. Message-critical content carried a small excess after matching, surviving detector removal (rule-free OR 1.10). The ratio of faithful rescues to fluent errors exceeded 1 at low corruption but fell below 1 at 20-30% target CER. No interface policy removed drift: conservative editing and abstention lowered it, alternatives and expansion raised it; the best drifted on 18.0 per 100. A 2,281-item panel (16 of 20 models) gave moderate ensemble-versus-consensus agreement (kappa 0.41); correction lowered pooled drift 31.4% to 28.3%, and a CER-stratified physician-corrected re-analysis confirmed the dose-response at each level. Interpretation Language-model post-editing produced fluent semantic substitutions that rose with corruption, confidence did not reliably flag, and no interface policy removed. This does not demonstrate clinical harm; prospective human-in-the-loop evaluation is needed. Funding: A.G. and E.K. were supported in part by the Clinical and Translational Science Awards (CTSA) grant UL1TR002541 from the National Center for Advancing Translational Sciences, through the Harvard Catalyst | The Harvard Clinical and Translational Science Center Pilot Award Program. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. Competing interests: The authors declare that they have no competing interests.

A. Gorenshtein, M. Omar, E. Jia et al. · 0 citations
#small language model Review Open access Aug 2026

Artificial intelligence-enabled sustainability in sme supply chains: a systematic and bibliometric literature review

The growing convergence of sustainability pressures, digital transformation, and supply chain complexity lead to a rising interest in how artificial intelligence (AI) can support sustainable supply chain management, particularly in small and medium-sized enterprises (SMEs). Despite the expanding literature on AI and sustainability, important gaps remain regarding how AI-enabled technologies contribute to economic, environmental, and social performance under the resource constraints typically faced by SMEs. Existing studies frequently address AI adoption through broad discussions of digital transformation, forecasting, or automation, while providing comparatively limited integration of sustainability and SME-specific organizational conditions. This research investigates the application of AI technologies in SME supply chains and their contribution to sustainability performance. The research adopts a systematic and bibliometric literature review approach based on the PRISMA methodology. Peer-reviewed journal articles indexed in Scopus and published between 2022 and 2026 were analyzed. Following a structured screening and eligibility assessment process, 49 articles were included in the final sample. The selected studies were examined through descriptive profiling, thematic coding, and bibliometric analyses covering AI technologies, sustainability dimensions, supply chain applications, implementation barriers, and emerging research trends. The findings indicate that machine learning is the dominant AI technology in SME supply chains, primarily used for forecasting, inventory management, process monitoring, logistics optimization, anomaly detection, and operational decision support. Natural language processing, large language models, and computer vision appear less frequently but are increasingly relevant for communication, information management, and intelligent operational analysis. The review also shows that economic and environmental sustainability dimensions receive substantially greater attention than social sustainability. Recurring barriers to AI adoption include financial constraints, weak digital infrastructures, fragmented data environments, limited analytical capabilities, and organizational readiness challenges. At the same time, managerial commitment, strategic alignment, technological partnerships, and policy support emerge as important enabling factors. This investigation contributes to the literature by integrating research on AI, sustainability, and SME supply chains through a combined systematic and bibliometric perspective. The findings highlight that the sustainability potential of AI depends not only on technological capabilities but also on organizational conditions, data governance, and SMEs’ ability to integrate AI-enabled decision support into operational processes. The study also identifies important theoretical, managerial, and methodological gaps and proposes directions for future research on AI-enabled sustainability in SME supply chains.

L. Fonseca, Luca Esposito, T. Murino et al. · 0 citations

Low Carbon Scheduling of Integrated Energy System Based on Large Language Model-Embedded Multi-Agent Reinforcement Learning

The complex multi-energy coupling characteristics inherent to integrated energy system (IES) present unprecedented challenges for the implementation of low-carbon scheduling. Existing optimization methods often exhibit limitations in system scalability, algorithm adaptivity, and carbon reduction efficacy for complex IES. This paper proposes a Large Language Model (LLM)-Embedded Multi-Agent Reinforcement Learning (LEMARL) to address the aforementioned issues. The proposed method integrates the global perception capability of LLMs with the dynamic optimization capability of MARL. Specifically, the LLM-Embedded module generates high-quality reward functions and policy frameworks from a global perspective, while the MARL module leverages these LLM-generated strategies for distributed interactive iterations—greatly enhancing computation efficiency and scalability. Simulation results demonstrate that LEMARL reduces carbon emissions by 7.76% and simultaneously decreases operating costs by 4.49% in a small-scale IES. Furthermore, LEMARL also exhibits superior applicability and scalability in large-scale IES of the IEEE 141-bus power grid integrated with 51-node thermal system.

Chen Xia, Tong Gou, Yinliang Xu et al. · 1 citation

DuoPIM: RRAM–DRAM Hybrid PIM Acceleration for Flexible-Batch LLM Decoding

Transformer-based large language models (LLMs) primarily consist of weight-intensive fully connected (FC) layers and cache-dependent attention layers. While batching significantly enhances the throughput of FC layers, it paradoxically increases the cache demands of attention layers. This provides no performance benefit and creates substantial memory pressure. Consequently, existing graphics processing unit (GPU)-based LLM acceleration systems face throughput limitations from batch size constraints. Even when DRAM-based processing-in-memory (PIM) is employed to accelerate attention, the utilization remains extremely low under small batch sizes, which is unsuitable for low-batch scenarios. Fortunately, the emerging nonvolatile resistive random access memory (RRAM) technology offers batch size-insensitive acceleration for FC layers through highly parallel in situ computations by eliminating weight loading overhead. This insight leads us to propose a hybrid approach: RRAM for FC layers and DRAM PIM for attention layers to overcome batch size limitations. However, merely scaling existing RRAM architectures misaligned with LLMs’ computation and storage demands will result in prohibitive overheads. Meanwhile, existing DRAM-based PIMs suffer from poor resource utilization due to the computational pattern of attention layers. Implementing an effective scheduling strategy is equally crucial to harness the potential of the hybrid PIM system. To address these challenges, we present DuoPIM, a novel RRAM–DRAM hybrid PIM architecture optimized for LLM decoding. We introduce novel architectural innovations for both the RRAM and DRAM PIM components to address the challenges posed by LLMs. Specifically, we decouple RRAM’s storage and computing capabilities within a hierarchical architecture, implement minimal modifications to DRAM PIM to support online softmax, and devise dedicated strategies across multiple architectural levels to enhance overall resource utilization. Evaluations demonstrate DuoPIM’s ability to fully leverage computing capacity across various batch sizes.

Xiaotian Sun, Xinyu Wang, Wanqian Li et al. · 0 citations
#small language model Preprint Aug 2026

Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study

A preliminary study on the adaptation of Whisper for Automatic Speech Recognition in Baniwa, an indigenous Arawakan language spoken in Brazil, Colombia, and Venezuela, demonstrating that multilingual foundation models can be successfully adapted to extremely low-resource indigenous languages.

Leonardo Duart, T. Fonseca, T. Chacon · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.