Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

GPT-Red: Automated Red Teaming via Self-Play at Scale

GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs, is introduced, a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents.

Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal et al. · 1 citation
Preprint Jul 2026

Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents?

This work instantiates Feedback-Augmented Self-Distillation (FA-SD), a self-distillation algorithm for agentic search that leverages successful demonstrations as privileged information and identifies that models can rely on recurring reasoning-and-search output templates, producing trajectories that appear diverse but are largely agnostic to the input question, making the KL-based self-distillation signal uninformative.

Fan Yang, Rui Meng, Yuxin Wen · 1 citation · ⚡1