This paper demonstrates the full potential of behavior-preserving symmetries as a defense against stegomalware, as well as the risks these symmetries pose when exploited by attackers, and quantifies the loss in model performance associated with applying these methods.
Abstract
The difficulty of training large language models (LLMs), together with their ubiquity, raises the threat of stegomalware, where malicious payloads are embedded into model weights. Recent work has demonstrated the use of permutation symmetry in model weights to mitigate these threats, but failed to show neutralization of stegomalware across all weights for LLMs. In this paper, we demonstrate the full potential of behavior-preserving symmetries as a defense against stegomalware, as well as the risks these symmetries pose when exploited by attackers. For stegomalware neutralization, we improve upon previous work, demonstrating that it is possible to select permutations which displace all model parameters. This contrasts with previous methods which left a significant percentage of weights unaltered in LLMs. When used in an attack, we show that permutation symmetries can encode malware into the weights of a model in a way that is theoretically lossless, requires no retraining after encoding, and needs no payload-specific information in the extraction script---a combination of characteristics not previously seen in any single method. While theoretically lossless, permutation can in practice alter model behavior due to the accumulation of numerical error. We therefore quantify the loss in model performance associated with applying these methods, for both attack and defense, showing it to be minimal.
Large language model safety and security research is preoccupied with, among other things, detecting and preventing jailbreak attacks: alignment bypasses that allow an adversarial user to elicit unwanted or harmful outputs from models. Arbitrary cipher, or covert communication, attacks are one such type of jailbreak an...
LoRO can solve the concerns regarding model thefts on edge devices in an efficient and secure manner, facilitating the wide edge application of LLMs and identifying a statistical vulnerability in existing protection methods.
Gao-Jian Xiong, Yu Sun, Jian-Hua Liu et al.· Neural Information Processin...· 6 citations
This paper formalizes a set of obfuscation primitives, defined as dual-tuples of linear computations satisfying specific algebraic properties, and demonstrates that the matrix-level weight transformations of several representative efficient TSLP frameworks can be expressed as compositions of these primitives.
Han-Yi Zhou, Chen-Yang Li, Yuan-Zhe Pang et al.· 0 citations
We introduce Stateless Bernoulli Watermarking (SBW), a new statistical watermark for Large Language Models that determines green list membership through independent per-token Bernoulli trials. Unlike KGW's vocabulary permutation or SynthID's multi-layer tournament, SBW requires only a single comparison per token agains...
This work presents SCRIPTIOC-BENCH, a benchmark for measuring static IOC extraction capability on real-world malicious scripts, and evaluates a broad range of proprietary and open-weight LLMs, showing that IOC recovery without execution remains challenging across model scales.
Hanna Kim, Jian Cui, Minkyoo Song et al.· 0 citations
This work proves that the operational validity of the CLWE backdoor critically hinges on assumptions that are incompatible with the realistic RFF learning deployment and analyzes the adversarial robustness of RFF learning models and provides a concrete certified robustness analysis, enabling a deeper security assessmen...
Tianshuo Cong, Pei Li, Hao-Jie Wu et al.· Proceedings of the 32nd ACM...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 29, 2026
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.