Skip to content
Preprint

HaloMark: A Spectral Threshold for Embedding-Vector Watermarking under C2PA

Aug 2026 · 0 citations · 38 references
Computer Science

TL;DR

HaloMark, a watermark for embedding vectors cryptographically bound to a C2PA manifest, is presented, a watermark for embedding vectors cryptographically bound to a C2PA manifest bound rigorously for linear and non-adaptive attackers and characterise empirically for the adaptive case.

Abstract

Foundation-model embeddings are now a primary data asset, but the content-provenance machinery built for images and audio does not transfer to them. C2PA binds to an asset with a stable bit-level or perceptual identity; embeddings have neither, since quantisation, projection, fine-tuning, and windowed averaging reshape them in normal use and break any fixed hash. We present HaloMark, a watermark for embedding vectors cryptographically bound to a C2PA manifest. It composes four standard primitives -- a block-diagonal orthogonal rotation, public whitening, an input-dependent LSH commitment, and a per-vector nonce -- around one protocol change: the producer signs the LSH commitment c into the C2PA sidecar, and the verifier reads c from the manifest instead of recomputing it. Recomputing is fragile under whitening, which flips the commitment bucket on 62% of inputs at cos = 0.96; reading the signed c reduces the verifier's score to T = T_null + beta(A)*epsilon, so security turns on a single scalar beta, which we bound rigorously for linear and non-adaptive attackers and characterise empirically for the adaptive case. We evaluate against an adversary holding polynomially many clean/watermarked pairs under one key with full sidecar visibility, across eight baselines and ten adaptive attackers including denoising-autoencoder removal. The eleven encoders separate at an empirical threshold eff_rank(Sigma)/d ~= 0.19: above it, detection AUROC stays at 0.98 or higher across every in-budget attack on the three encoders we sweep in full, and at 0.965 or higher under single-seed DAE removal on the rest; below it every variant we tested fails. Why the threshold is dimension-uniform is left open. Deployed as a Qdrant admission filter, the verifier runs at 284 us and 24 bytes of sidecar per vector, validated end-to-end against three C2PA reference-SDK bindings.

View source

Similar papers

Preprint Aug 2026

KeyBound: Keyed and Host-Bound Learned Audio Watermarking for Speech Provenance

KeyBound is presented, a learned audio watermark that restores the two ingredients classical watermarking supplied and learned schemes set aside, a secret key and a host-aware carrier, so the key governs payload access while the host-conditioned carrier resists direct transplantation.

Bang-Shuo Zhu, Yu-Xin Cao, Wei-Fei Jin et al. · 0 citations
Preprint Aug 2026

SpreadMark: Robust Image Watermarking via Spread-Spectrum Embedding

SpreadMark keeps the embedded watermark imperceptible, maintaining high perceptual quality on both COCO and DIV2K, and is the only evaluated method retaining high detection under both the regeneration and the latent-space sparsification settings the authors test.

Wei Song, Yu-Xin Cao, Zhen-Chang Xing et al. · 0 citations
#machine learning Preprint Sep 2026

Flip, Don't Shuffle: Watermarking LLMs at the Speed of Inference

We introduce Stateless Bernoulli Watermarking (SBW), a new statistical watermark for Large Language Models that determines green list membership through independent per-token Bernoulli trials. Unlike KGW's vocabulary permutation or SynthID's multi-layer tournament, SBW requires only a single comparison per token agains...

Simone Ceppi, Ignacio Sanchez · 0 citations
Preprint Aug 2026

Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs

Experiments across four controlled manipulation types under ideal channel conditions show that the embedded payload, and hence an approximate reconstruction of the authentic content, is always fully recovered without bit errors, and the results indicate that the choice of neural codec is the dominant factor for detecti...

Yigitcan Özer, Xin Wang, Zhe Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

OpenStamp: A Watermark for Open-Source Language Models

This work introduces OpenStamp, a watermarking technique that encodes the watermarking logic directly into the model weights by modifying only the final projection, or unembedding, layer, and shows that OpenStamp achieves superior detection performance, with minimal degradation in model capabilities compared to prior m...

Miroojin Bakshi, Saksham Rastogi, Danish Pruthi · 0 citations
Preprint Aug 2026

Asymmetric Phase Coding Video Watermarking

Existing video watermarking systems are symmetric: the party that can verify a mark holds the extractor weights or generator secret and can therefore also embed one. Benchmarks confirm the consequence, reporting that white-box forgery defeats all evaluated methods. We present a training-free video watermark that remove...

Guang Yang, Feng-Chen Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.