Skip to content
#small language model Review Open access

Artificial intelligence-enabled sustainability in sme supply chains: a systematic and bibliometric literature review

Aug 2026 · Management & Marketing · Vol 21 · 0 citations · 52 references

Abstract

The growing convergence of sustainability pressures, digital transformation, and supply chain complexity lead to a rising interest in how artificial intelligence (AI) can support sustainable supply chain management, particularly in small and medium-sized enterprises (SMEs). Despite the expanding literature on AI and sustainability, important gaps remain regarding how AI-enabled technologies contribute to economic, environmental, and social performance under the resource constraints typically faced by SMEs. Existing studies frequently address AI adoption through broad discussions of digital transformation, forecasting, or automation, while providing comparatively limited integration of sustainability and SME-specific organizational conditions. This research investigates the application of AI technologies in SME supply chains and their contribution to sustainability performance. The research adopts a systematic and bibliometric literature review approach based on the PRISMA methodology. Peer-reviewed journal articles indexed in Scopus and published between 2022 and 2026 were analyzed. Following a structured screening and eligibility assessment process, 49 articles were included in the final sample. The selected studies were examined through descriptive profiling, thematic coding, and bibliometric analyses covering AI technologies, sustainability dimensions, supply chain applications, implementation barriers, and emerging research trends. The findings indicate that machine learning is the dominant AI technology in SME supply chains, primarily used for forecasting, inventory management, process monitoring, logistics optimization, anomaly detection, and operational decision support. Natural language processing, large language models, and computer vision appear less frequently but are increasingly relevant for communication, information management, and intelligent operational analysis. The review also shows that economic and environmental sustainability dimensions receive substantially greater attention than social sustainability. Recurring barriers to AI adoption include financial constraints, weak digital infrastructures, fragmented data environments, limited analytical capabilities, and organizational readiness challenges. At the same time, managerial commitment, strategic alignment, technological partnerships, and policy support emerge as important enabling factors. This investigation contributes to the literature by integrating research on AI, sustainability, and SME supply chains through a combined systematic and bibliometric perspective. The findings highlight that the sustainability potential of AI depends not only on technological capabilities but also on organizational conditions, data governance, and SMEs’ ability to integrate AI-enabled decision support into operational processes. The study also identifies important theoretical, managerial, and methodological gaps and proposes directions for future research on AI-enabled sustainability in SME supply chains.

Read PDF

Similar papers

#small language model Open access Aug 2026

PARA: Perception, Action, Reasoning, Adaptation. Four Faculties an Institution Can Revoke

The fourth faculty is Adaptation. Any source rendering it as Reflection is in error, including sources by this author, and the distinction is not cosmetic: reflection is a private act with no external consequence, while adaptation writes to institutional memory, which is why it needs a guardrail and why misnaming it removes the reason for one. No trademark is claimed on PARA or on any of the four faculty names. The construct is offered for use, teaching, assessment, extension and criticism by anyone, with attribution, under CC BY 4.0. An operational agent that watches a system and acts on it is usually described as a perceive-and-act loop, and the description omits the two things an institution needs. It omits the reasoning that justifies an action, which is the only part that can be argued with once the action turns out to have been wrong. And it omits the adaptation that closes the loop, which is where the agent's experience becomes something the institution keeps. PARA names four faculties, each carrying a distinct authority type. Perception has read-only access to system signals and emits structured observations, distinguishing what was measured from what was inferred. Reasoning has read access to observations and runbooks, emits a plan and its justification, and writes nothing at all, which is what makes it safe to give it the widest read access of the four. Action holds the sole authority to change production, through enumerated policy-authorized operations only. Adaptation has write access to institutional knowledge and no write access to production. Two faculties write and two do not, and the two that write are the two that carry guardrails. The substantive requirement is that Adaptation is bounded by the same guardrails as Action, which reads as excessive until the failure it prevents is named. An agent that could both act and rewrite the record of its action could launder its own mistakes into institutional memory, and the institution would then improve its future decisions from a corrected account. Nothing about that is detectable downstream, because the record is the only thing downstream has and there is no second copy to compare against. The failure does not require a deceptive agent: one adapting honestly from a mistaken belief about its own action produces the same result, which makes the guardrail a defence against a normal agent rather than a malicious one. The second requirement is the registry entry that turns a faculty from a description into a contract, carrying the faculty, its allowed actions, its forbidden actions, its governing guardrail and its success metrics. Forbidden actions are named although they are formally the complement of the allowed set, because a reviewer cannot otherwise tell a capability deliberately withheld from one nobody thought of. Success metrics sit in the same entry because the metric is what the agent's optimizer pushes against the guardrail. An agent must not exercise a faculty its entry does not record, and an agent that quietly acquires one usually does so incrementally and with good intent: a reasoning faculty given a small write to make itself useful is an action faculty with no guardrail. The acronym and the loop are in different orders, which the specification states explicitly because the mismatch is a reliable source of confusion. The acronym reads P-A-R-A; the loop runs perception, reasoning, action, adaptation, and reasoning precedes action so that a justification is not constructed afterwards. This is the depth treatment of pattern OP-5 of A Pattern Language for Production LLM Platforms, which is the canonical statement and governs where the two disagree. Documented uses of the full four-part model are emerging rather than established, no implementation unconnected to the author has been evaluated, and the laundering failure is argued rather than observed, which the specification records as a weakness of the argument and not only of the phenomenon. It is a specification, not a certification scheme.

Nabeel A. Khan · 2 citations
#small language model Open access Aug 2026

LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences

LifeSciBench is introduced, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work, with each constituent task paired with a human expert-written rubric.

Amelia Liu, Andrew Ho, Anne Marie Droste et al. · 2 citations
#small language model Preprint Aug 2026

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost.

Yu-Fan Wu, Yinghui He, Zhengyi Hu et al. · 1 citation
#artificial intelligence Preprint Aug 2026

TestifAI: Tomography-Based Testing for Deep Learning Systems

TestifAI, a deep learning testing framework for efficient and accurate estimation of robustness against combinations of perturbations, is proposed and partial model tomography is introduced, a novel approach to reconstructing model behaviour in a multi-perturbation space from tests that apply only a small number of perturbations.

Arooj Arif, T. Hartung, E. Botoeva et al. · 1 citation
#small language model Preprint Aug 2026

HEPToolBench 1.2: Testing How Reliably Language Models Can Drive Particle Physics Software

HEPToolBench is introduced, a benchmark of 28 collider-simulation tasks scored by deterministic, task-specific scorers, plus a three-task structured-debugging extension, and moving syntax generation into deterministic software can substantially improve reliability for both small local and frontier models.

Unknown authors · 1 citation

Related blog posts

Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.