Skip to content

Author

Robin Staab

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

AutoMark: Enabling Autoresearch to Discover Better LLM Watermarks

With LLM watermarking being deployed commercially and now required by regulations, improving its reliability and effectiveness has become crucial. Yet, recent progress in the field of LLM watermarking has increasingly been driven by improving details of existing methods, an effort fundamentally limited by the pace of h...

Thibaud Gloaguen, Robin Staab, Martin T. Vechev · 0 citations
#artificial intelligence Preprint May 2025

Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning

An attack is proposed, FAB (Finetuning-activated Adversarial Behaviors), which compromises an LLM via meta-learning techniques that simulate downstream finetuning, explicitly optimizing for the emergence of adversarial behaviors in the finetuned models.

Thibaud Gloaguen, Mark Vero, Robin Staab et al. · 4 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.