Skip to content

ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour

Jul 2026 · arXiv.org · Vol abs/2607.23478 · 1 citation · 63 references
Computer Science

TL;DR

ATLAS is a training-free framework that automates this search by treating each layer's approximation setting as a multi-objective optimization over latency and accuracy, and works across encoder-only, decoder-only, and vision Transformers, complementing parallel work on packing and matrix multiplication.

Abstract

Fully homomorphic encryption (FHE) lets a server run inference on encrypted data with strong privacy guarantees, but running a Transformer under FHE is expensive. Its non-linear operations, such as softmax, normalization, and activation, must be replaced with polynomial approximations that the CKKS scheme supports, and the depth of these approximations dominates inference cost. Existing FHE Transformers use hand-tuned approximation settings, such as iteration count and polynomial degree, applied uniformly across layers, models, and tasks. Hand-tuning is slow and error-prone. Even a single uniform setting has about $10^7$ choices, and manual search cannot exploit layer-wise variation. AutoFHE, the only automated method with multi-objective search, targets ReLU-only CNNs and needs full fine-tuning per candidate, which is too costly for Transformers. Per-layer settings also push the search space to about $10^{85}$ for BERT and ViT and $10^{228}$ for LLaMA3, beyond both manual and fine-tuning-based search. We present ATLAS, a training-free framework that automates this search by treating each layer's approximation setting as a multi-objective optimization over latency and accuracy. The problem is hard: the decision space is large (96 or 256 variables), each configuration takes 70 to 1,000 seconds to evaluate even in cleartext, and 85 to 90 percent of configurations are invalid. ATLAS handles this with a two-stage optimization strategy and a surrogate model, completing the search in about one hour. Compared to an iterative softmax baseline, ATLAS cuts multiplicative depth and end-to-end latency by about 35 percent with little accuracy loss, and works across encoder-only, decoder-only, and vision Transformers, complementing parallel work on packing and matrix multiplication.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

HEAT: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation

Fully homomorphic encryption (FHE) allows a server to run a language model directly on encrypted user prompts, but current approaches remain prohibitively slow. Ciphertexts natively support only addition, multiplication, and rotation, and multiplications may be composed only to a bounded depth before a costly bootstrap...

Alessandro Zirilli, Davide Marincione, Evgenios M. Kornaropoulos et al. · 0 citations
Preprint Sep 2026

An Open-Source End-to-End FHE Implementation for Privacy-Preserving Llama 3 8B Inference

Odin is the first open-source end-to-end GPU CKKS implementation of Llama-3, an FHE inference system that co-designs ciphertext packing and model execution for Llama and uses a feature-major cross-layer layout to unify residual connections and layer interfaces.

Yu-Hang Fan, Yu-Si Chen, Kan-Yu Ye et al. · 0 citations
Preprint Sep 2026

ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE Evaluation

ROSETTA is proposed, a hybrid CKKS/TFHE framework that overcomes inefficiency in evaluating nonlinear operations, which incur substantial overhead and dominate the decode stage and achieves up to $4.8\times$ Softmax speedup and $1.5$--$2.1\times$ end-to-end speedup over the SOTA framework CacheMir.

Jiang-Rui Yu, Bao-Sheng Zhang, Liang Kong et al. · 0 citations
Preprint Sep 2026

Memory-Efficient Designs for Word-Wise Universal Fully Homomorphic Encryption

Fully Homomorphic Encryption (FHE) enables computation on encrypted data, preserving privacy throughout analysis. While its privacy is very strong, FHE is much slower to execute than the original computation. In particular, due to the recent success in accelerating its compute, the performance bottleneck shifts to the...

A. W. B. Yudha, Erwin Eko Wahyudi, R. Rajagede et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Training-Free Hidden-State Refinement for Flow-Matching Image Generators

A training-free looping framework that repeatedly applies selected transformer layers inside each denoising call is introduced, which improves primary and auxiliary quality metrics with competitive quality--efficiency trade-offs across two Scale-RAE model scales.

Yuan-Yi Yan, Xin-Zhe Rao, Can-Yu Shen et al. · 0 citations
Preprint Aug 2026

Every Expert Counts: ExactMoE for Memory-Efficient W4A16 Inference

ExactMoE, an inference design that applies symmetric group-128 four-bit weight quantization only to routed experts, stores those experts in kernel-native MARLIN form in pinned host memory, and executes all selected experts through a configurable GPU-resident slot cache and fused grouped MoE kernels, identifies a practi...

Amjad Saab · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.