The key idea is to induce the adversary to solve a misspecified inverse problem, in which no plausible label sequence in the sequence space can explain the observed gradients, by inducing inconsistency across three dimensions: objective, direction, and scale.
Shiyu Miao, Yunlong Mao, Zirui Huang et al.· 0 citations
This work introduces a deterministic full-chain memorization mechanism that locks onto token-level secrets in dynamic computation flows via online tensor-rule matching, and leverages value-gradient decoupling to stealthily inject attack gradients, overcoming gradient drowning to force model memorization.
Zi Li, Tianyang Zhou, Wenze Li et al.· arXiv.org· 0 citations
DPA is grounded in a critical insight: regardless of fine-tuning tactics to evade provenance, the practical necessity of maintaining utility constrains the model to preserve the fundamental intersection of semantic substance and lexical form, and captures this persistent lexical-semantic intersection as intrinsic distributional fingerprints.
Zirui Huang, Yunlong Mao, Wei Tong et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.