Skip to content
Book Open access

Black-Box Embedding Inversion Attack on Vector Databases

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 4486-4497 · 0 citations · 50 references

TL;DR

A novel black-box image embedding inversion attack that reconstructs high-fidelity images using only query access to the embedding model or API, and introduces an embedding-guided cross-attention mechanism, where image embeddings serve as conditional signals to steer the generation process.

Abstract

Vector databases that index and serve dense embeddings have become central to modern data science applications. Embeddings are often regarded as privacy-preserving surrogates for raw data, motivating practitioners to outsource vector databases to third-party services for scalability. However, recent studies show that even text embeddings alone can leak sensitive information, raising serious privacy concerns. Existing attacks on image embeddings, meanwhile, typically assume access to model architecture or parameters, which does not hold in outsourced settings. In this paper, we propose a novel black-box image embedding inversion attack that reconstructs high-fidelity images using only query access to the embedding model or API. Our approach leverages an in-distribution auxiliary dataset to train a conditional diffusion model, capturing domain-aligned knowledge of the data owner's private images. We introduce an embedding-guided cross-attention mechanism, where image embeddings serve as conditional signals to steer the generation process. To improve efficiency, we perform diffusion in the latent space of a pretrained VQGAN with deterministic decoding, which reduces the computational cost while preserving both structural and perceptual fidelity in reconstructed images. Extensive experiments on three real-world datasets and three popular embedding models demonstrate that our approach significantly outperforms four state-of-the-art baselines. These findings reveal that image embeddings can expose sensitive visual information, highlighting the need for stronger privacy protections in outsourced vector databases.

Read PDF

Similar papers

#artificial intelligence Preprint Aug 2026

Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings

DAEI is proposed, a denoising-aware embedding inversion pipeline that combines a residual denoising autoencoder with generative text inversion where the denoiser is trained in an unsupervised manner using Stein's unbiased risk estimate to enable denoising from noisy observations alone.

Yubo Wang, Shujie Cui, James Bailey et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Shadow Queries for Private Retrieval in Vector Databases

SHAQ (shadow query generation), a semantic-decomposition and embedding-decoupling defense against EIAs, is proposed, which uses a generative language model to create diverse shadow queries that capture different semantic aspects of each document.

Xinguo Feng, Zhongkui Ma, Zi-Han Wang et al. · 0 citations
#natural language process... Preprint Sep 2026

A Novel Semantic Manifold Alignment Attack against Embedding-to-Embedding Obfuscation in Privacy-Preserving LLMs

With the widespread applications of large language models (LLMs), privacy-preserving inference has become increasingly essential for sensitive queries. To balance privacy and utility, a series of lightweight obfuscation approaches has recently been proposed, where users locally transform plaintext embeddings into the f...

Si-Cong Li, Ling-Feng Yao, Xing-Ke Yang et al. · 0 citations
Preprint Aug 2026

DeepInvert: Semi-Supervised Embedding Inversion Against Obfuscated Language Models

This work presents DeepInvert, a semi-supervised embedding inversion attack that recovers original tokens from obfuscated representations with higher accuracy than prior methods, and reveals a task-dependent tension: obfuscation schemes preserving enough signal for utility also retain sufficient structure for inversion...

Zhicong Huang, Cheng Hong, Tao Wei · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.