Skip to content
Preprint

Robust Watermarks Meet Backdoored Models: Evading Diffusion Semantic Watermarks via Stealthy Backdoor

Aug 2026 · 0 citations · 52 references
Computer Science

TL;DR

GhostVAE is proposed to plant a stealthy backdoor into the encoder of Variational Autoencoder (VAE), enabling reliable evasion of watermark detection and fundamentally undermines the trustworthiness of semantic watermarking systems.

Abstract

Although semantic watermarking is considered a promising safeguard for images generated by Latent Diffusion Models (LDMs), the reliance of the watermark detection pipeline on neural networks introduces a critical yet underexplored backdoor attack surface. To systematically study this vulnerability, we propose GhostVAE to plant a stealthy backdoor into the encoder of Variational Autoencoder (VAE), enabling reliable evasion of watermark detection. GhostVAE operates in two stages: it first constructs a universal trigger via power spectrum regularization to improve the trigger robustness, and then trains a backdoored VAE encoder with a parameter-aligned objective. Through extensive evaluations across three state-of-the-art semantic watermarking schemes and three widely adopted LDMs, we show that GhostVAE preserves watermark detection performance on benign images (achieving an average true positive rate of 94.4%), while simultaneously enabling highly effective evasion under trigger activation (achieving an average attack success rate of 94.6%). Moreover, we comprehensively analyze seventeen representative defenses and demonstrate that GhostVAE remains stealthy across the input space, parameter space, and latent space. Our work fundamentally undermines the trustworthiness of semantic watermarking systems and highlights that secure deployment of semantic watermarks requires end-to-end security considerations, particularly for neural network components.

View source

Similar papers

Conference Open access Sep 2026

Latents-Inv:Robust Semantic Watermark via Dual-Path Mutual Information Redundancy for Diffusion Models

A dual-path network is proposed to encode watermark information into both the generated image and the owner’s secret key, which achieves superior robustness against various adversarial attacks while maintaining high visual quality across diverse generative models.

Cong-Rong Li, Ling-Yun Yu, Pei-Qi Jiang et al. · 0 citations
Open access Aug 2026

Deep Watermarking-based Proactive Defense for Deepfake Detection, Tracing, and Regulation

Current proactive defense mechanisms, though effective, predominantly concentrate on impeding deepfake models rather than regulating them. In light of the pervasive demand for deepfake creation for legitimate purposes, we introduce a proactive deepfake control framework based on a “whitelist” mechanism. This framework...

Yi-Zhi Guo, Bing-Wen Feng, Xiaotian Wu et al. · 0 citations
Preprint Sep 2026

MarkSec: Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

Results show that apparent attack winners depend on text-quality constraints, attack generality, and capability assumptions, and that apparent attack winners depend on text-quality constraints, attack generality, and capability assumptions.

Kai-Rong Li, Zhi-Kun Zhang, Xiao-Nan Ren et al. · 0 citations
Open access Sep 2026

FakeMark: gradient-guided false watermark claims via robust feature fusion

Model watermarking supports intellectual-property claims by verifying a model’s responses to a secret key set, but this behavior-only interface is vulnerable to fabricated evidence. This work presents FakeMark, a gradient-guided false-claim attack for image classifiers that uses a white-box surrogate but never queries...

Yu-Tong Wu, Wen-Yue Li, He-Wang Nie et al. · 0 citations
Open access 2026

Lattice-Quantization Identity Watermarking for Deep Neural Network Ownership Protection

With the increasing trend of open-sourcing deep neural network (DNN) models, protecting model ownership has become a critical challenge, particularly in high-stakes domains such as medical AI. Existing backdoor-based watermarking methods suffer from two key limitations: visible trigger patterns and the lack of a reliab...

Ying Xu, Zhi-Ying Li, Shanxiang Lyu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.