This work proposes Targeted Anti-obfuscation with Mechanistic Enforcement (TAME), which uses Sparse Autoencoders (SAEs) to combine behavioral feedback with targeted suppression of template-associated activations during RL, providing a path from behavioral monitoring to representation-level oversight for more auditable...
Xu-Tao Mao, Jianing Zhu, Jin-Man Zhao et al.· 0 citations
It is observed that OPD-trained models maintain superior avg@K performance across sampling budgets, while the advantage in pass@K gradually shifts to the pre-OPD base models as K increases, suggesting that OPD primarily improves sampling efficiency rather than consistently expanding the student's reasoning capability b...
This paper proposes a novel adversarial strategy, namely Prompt-Optimized Parameter Shaking (POPS), aiming to recover the supposedly unlearned multi-modality knowledge from the MLLMs, exposing fundamental vulnerabilities that challenge the foundational robustness of representative MMU-based privacy protections.
Zhangheng Li, Jianing Zhu, Junyuan Hong et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.