GhostVAE is proposed to plant a stealthy backdoor into the encoder of Variational Autoencoder (VAE), enabling reliable evasion of watermark detection and fundamentally undermines the trustworthiness of semantic watermarking systems.
Abstract
Although semantic watermarking is considered a promising safeguard for images generated by Latent Diffusion Models (LDMs), the reliance of the watermark detection pipeline on neural networks introduces a critical yet underexplored backdoor attack surface. To systematically study this vulnerability, we propose GhostVAE to plant a stealthy backdoor into the encoder of Variational Autoencoder (VAE), enabling reliable evasion of watermark detection. GhostVAE operates in two stages: it first constructs a universal trigger via power spectrum regularization to improve the trigger robustness, and then trains a backdoored VAE encoder with a parameter-aligned objective. Through extensive evaluations across three state-of-the-art semantic watermarking schemes and three widely adopted LDMs, we show that GhostVAE preserves watermark detection performance on benign images (achieving an average true positive rate of 94.4%), while simultaneously enabling highly effective evasion under trigger activation (achieving an average attack success rate of 94.6%). Moreover, we comprehensively analyze seventeen representative defenses and demonstrate that GhostVAE remains stealthy across the input space, parameter space, and latent space. Our work fundamentally undermines the trustworthiness of semantic watermarking systems and highlights that secure deployment of semantic watermarks requires end-to-end security considerations, particularly for neural network components.
A dual-path network is proposed to encode watermark information into both the generated image and the owner’s secret key, which achieves superior robustness against various adversarial attacks while maintaining high visual quality across diverse generative models.
Cong-Rong Li, Ling-Yun Yu, Pei-Qi Jiang et al.· Proceedings of the Thirty-Fi...· 0 citations
Current proactive defense mechanisms, though effective, predominantly concentrate on impeding deepfake models rather than regulating them. In light of the pervasive demand for deepfake creation for legitimate purposes, we introduce a proactive deepfake control framework based on a “whitelist” mechanism. This framework...
Yi-Zhi Guo, Bing-Wen Feng, Xiaotian Wu et al.· ACM Transactions on Multimed...· 0 citations
Results show that apparent attack winners depend on text-quality constraints, attack generality, and capability assumptions, and that apparent attack winners depend on text-quality constraints, attack generality, and capability assumptions.
Kai-Rong Li, Zhi-Kun Zhang, Xiao-Nan Ren et al.· 0 citations
Model watermarking supports intellectual-property claims by verifying a model’s responses to a secret key set, but this behavior-only interface is vulnerable to fabricated evidence. This work presents FakeMark, a gradient-guided false-claim attack for image classifiers that uses a white-box surrogate but never queries...
Yu-Tong Wu, Wen-Yue Li, He-Wang Nie et al.· Cybersecurity· 0 citations
This work identifies two complementary laundering regimes: OpenAI models produce the strongest payload disruption across the evaluated schemes, whereas Nano Banana 2 shows that DwtDct remains vulnerable under high-fidelity reconstruction.
With the increasing trend of open-sourcing deep neural network (DNN) models, protecting model ownership has become a critical challenge, particularly in high-stakes domains such as medical AI. Existing backdoor-based watermarking methods suffer from two key limitations: visible trigger patterns and the lack of a reliab...