Skip to content

LoRO: Real-Time on-Device Secure Inference for LLMs via TEE-Based Low Rank Obfuscation

2025 · Neural Information Processing Systems · 6 citations · 59 references
Computer Science

TL;DR

LoRO can solve the concerns regarding model thefts on edge devices in an efficient and secure manner, facilitating the wide edge application of LLMs and identifying a statistical vulnerability in existing protection methods.

Abstract

While Large Language Models (LLMs) have gained remarkable success, they are consistently at risk of being stolen when deployed on untrusted edge devices. As a solution, TEE-based secure inference has been proposed to protect valuable model property. However, we identify a statistical vulnerability in existing protection methods, and furtherly compromise their security guarantees by proposed Model Stealing Attack with Prior. To eliminate this vulnerability, LoRO is presented in this paper, which leverages dense mask to completely obfuscate parameters. LoRO includes two innovations: (1) Low Rank Mask, which uses low-rank factors to generate dense masks efficiently. The computing complexity in TEE is hence reduced by an exponential amount to achieve inference speed up, while providing robust model confidentiality. (2) Factors Multiplexing, which reuses several cornerstone factors to generate masks for all layers. Compared to one-mask-per-layer, the secure memory requirement is reduced from GB-level to tens of MB, hence avoiding the hundred-fold latency introduced by secure memory paging. Experimental results indicate that LoRO achieve a 0 . 94 × Model Stealing (MS) accuracy, while SOTA methods presents 3 . 37 × at least. The averaged inference latency of LoRO is only 1 . 49 × , compared to the 112 × of TEE-shielded inference. Moreover, LoRO results no accuracy loss, and requires no re-training and structure modification. LoRO can solve the concerns regarding model thefts on edge devices in an efficient and secure manner, facilitating the wide edge application of LLMs.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing

User prompts provided to large language models (LLMs) may contain sensitive or private information that can be misused by remotely deployed models, such as through inadvertent memorization during retraining. One way to protect user prompts is to execute the LLM inside a trusted execution environment (TEE), with the gua...

Shashie Dilhara Batan Arachchige, Robin Carpentier, H. Asghar et al. · 0 citations
Preprint Sep 2026

Understanding the Security Boundary of Obfuscation-based On-Device LLM Protection

This paper formalizes a set of obfuscation primitives, defined as dual-tuples of linear computations satisfying specific algebraic properties, and demonstrates that the matrix-level weight transformations of several representative efficient TSLP frameworks can be expressed as compositions of these primitives.

Han-Yi Zhou, Chen-Yang Li, Yuan-Zhe Pang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Permutation-Based Stegomalware in Large Language Models: Threats and Countermeasures

This paper demonstrates the full potential of behavior-preserving symmetries as a defense against stegomalware, as well as the risks these symmetries pose when exploited by attackers, and quantifies the loss in model performance associated with applying these methods.

Danny Wood, James Stringer · 0 citations
Preprint Sep 2026

Type-Directed, Secure-by-Construction Enclave Partitioning for LLVM

Trusted Execution Environments (TEEs) provide hardware-supported isolation through enclaves that protect code and data independently of software abstractions. However, TEEs alone cannot enforce information-flow security. This problem is further aggravated in LLVM-like low-level languages that allow unrestricted pointer...

Wesley B. Nuzzo, Samuel Dodson, Benjamin Houle et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.