Skip to content
Open access

FakeMark: gradient-guided false watermark claims via robust feature fusion

Sep 2026 · Cybersecurity · Vol 9 · 0 citations · 45 references

Abstract

Model watermarking supports intellectual-property claims by verifying a model’s responses to a secret key set, but this behavior-only interface is vulnerable to fabricated evidence. This work presents FakeMark, a gradient-guided false-claim attack for image classifiers that uses a white-box surrogate but never queries or accesses the victim model during attack construction. Under a simplified linear decision-boundary model, targeted perturbations can acquire a nonzero component along the watermark-trigger direction; experiments on deep networks provide only conditional, setting-dependent support for this intuition. FakeMark caches selected convolutional and fully connected layer outputs from clean surrogate batches and injects them through stochastic multi-layer, channel-wise interpolation to improve transfer. Across 16 distinct architectures and an additional adversarially trained ResNet-50 checkpoint variant, over eight evaluated watermark variants, retrospective best-case behavioral target-label accuracy reaches 1.00 on CIFAR-10 and 0.99 on ImageNet. ImageNet transfer varies substantially across checkpoint and surrogate settings, ranging from near zero to 0.99. Matched baselines and detector analyses motivate provenance-aware, multi-factor ownership protocols.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.