Preprint
Jul 2026
Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study
This work builds a working instance on a hierarchical multi-agent system, runs it under benign and attacked conditions across five language models and two task domains, and measures how much of that warning rests on removable surface cues of the attack rather than on its distributed structure.
Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang et al.
· 0 citations