ATLAS is a training-free framework that automates this search by treating each layer's approximation setting as a multi-objective optimization over latency and accuracy, and works across encoder-only, decoder-only, and vision Transformers, complementing parallel work on packing and matrix multiplication.
Abstract
Fully homomorphic encryption (FHE) lets a server run inference on encrypted data with strong privacy guarantees, but running a Transformer under FHE is expensive. Its non-linear operations, such as softmax, normalization, and activation, must be replaced with polynomial approximations that the CKKS scheme supports, and the depth of these approximations dominates inference cost. Existing FHE Transformers use hand-tuned approximation settings, such as iteration count and polynomial degree, applied uniformly across layers, models, and tasks. Hand-tuning is slow and error-prone. Even a single uniform setting has about $10^7$ choices, and manual search cannot exploit layer-wise variation. AutoFHE, the only automated method with multi-objective search, targets ReLU-only CNNs and needs full fine-tuning per candidate, which is too costly for Transformers. Per-layer settings also push the search space to about $10^{85}$ for BERT and ViT and $10^{228}$ for LLaMA3, beyond both manual and fine-tuning-based search. We present ATLAS, a training-free framework that automates this search by treating each layer's approximation setting as a multi-objective optimization over latency and accuracy. The problem is hard: the decision space is large (96 or 256 variables), each configuration takes 70 to 1,000 seconds to evaluate even in cleartext, and 85 to 90 percent of configurations are invalid. ATLAS handles this with a two-stage optimization strategy and a surrogate model, completing the search in about one hour. Compared to an iterative softmax baseline, ATLAS cuts multiplicative depth and end-to-end latency by about 35 percent with little accuracy loss, and works across encoder-only, decoder-only, and vision Transformers, complementing parallel work on packing and matrix multiplication.
Fully homomorphic encryption (FHE) allows a server to run a language model directly on encrypted user prompts, but current approaches remain prohibitively slow. Ciphertexts natively support only addition, multiplication, and rotation, and multiplications may be composed only to a bounded depth before a costly bootstrap...
Alessandro Zirilli, Davide Marincione, Evgenios M. Kornaropoulos et al.· 0 citations
Odin is the first open-source end-to-end GPU CKKS implementation of Llama-3, an FHE inference system that co-designs ciphertext packing and model execution for Llama and uses a feature-major cross-layer layout to unify residual connections and layer interfaces.
Yu-Hang Fan, Yu-Si Chen, Kan-Yu Ye et al.· 0 citations
ROSETTA is proposed, a hybrid CKKS/TFHE framework that overcomes inefficiency in evaluating nonlinear operations, which incur substantial overhead and dominate the decode stage and achieves up to $4.8\times$ Softmax speedup and $1.5$--$2.1\times$ end-to-end speedup over the SOTA framework CacheMir.
Jiang-Rui Yu, Bao-Sheng Zhang, Liang Kong et al.· 0 citations
Fully Homomorphic Encryption (FHE) enables computation on encrypted data, preserving privacy throughout analysis. While its privacy is very strong, FHE is much slower to execute than the original computation. In particular, due to the recent success in accelerating its compute, the performance bottleneck shifts to the...
A. W. B. Yudha, Erwin Eko Wahyudi, R. Rajagede et al.· 0 citations
A training-free looping framework that repeatedly applies selected transformer layers inside each denoising call is introduced, which improves primary and auxiliary quality metrics with competitive quality--efficiency trade-offs across two Scale-RAE model scales.
Yuan-Yi Yan, Xin-Zhe Rao, Can-Yu Shen et al.· 0 citations
ExactMoE, an inference design that applies symmetric group-128 four-bit weight quantization only to routed experts, stores those experts in kernel-native MARLIN form in pinned host memory, and executes all selected experts through a configurable GPU-resident slot cache and fused grouped MoE kernels, identifies a practi...
Amjad Saab· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.