With LLM watermarking being deployed commercially and now required by regulations, improving its reliability and effectiveness has become crucial. Yet, recent progress in the field of LLM watermarking has increasingly been driven by improving details of existing methods, an effort fundamentally limited by the pace of h...
Thibaud Gloaguen, Robin Staab, Martin T. Vechev· 0 citations
This work evaluates coding agents' completion performance in two complementary settings: established SWE-bench tasks from popular repositories, with LLM-generated context files, and a novel collection of issues from repositories containing developer-committed context files.
Thibaud Gloaguen, Niels Mündler, Mark Niklas Müller et al.· arXiv.org· 28 citations· ⚡1
An attack is proposed, FAB (Finetuning-activated Adversarial Behaviors), which compromises an LLM via meta-learning techniques that simulate downstream finetuning, explicitly optimizing for the emergence of adversarial behaviors in the finetuned models.
Thibaud Gloaguen, Mark Vero, Robin Staab et al.· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.