Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually....
Alexander Nemecek, Osama Zafar, Debargha Ganguly et al.· 1 citation
Generative AI makes social-engineering attacks more fluent, adaptive, and scalable, increasing the need for LLM-based de- fenders that can protect users during ongoing interactions. We ask whether such defenders identify the structural source of risk or merely react to surface cues. We formalize trust-chain localizatio...
Yu-Qiao Xu, Osama Zafar, Alexander Nemecek et al.· 0 citations
Machine learning explanations reveal model behavior beyond predictions, creating attack surfaces for model confidentiality and data privacy. We systematize 25 studies that exploit explanations for model extraction, membership inference, and model inversion, treating attribute inference as partial inversion. Existing wo...
A. Oksuz, Anisa Halimi, Erman Ayday· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.