Research · curated 15 Jul 2026
GPT-Red: Unlocking Self-Improvement for Robustness
First reported · updated · 2 reports openai.com
Coverage timeline
Why it matters
GPT-Red signals a shift toward scaling automated adversarial training against prompt injection, a defensive technique defenders of LLM and agent deployments should understand as attack surfaces grow through browsers, connected apps, and tool use.
OpenAI describes GPT-Red, an internal-only automated red-teaming model trained via self-play at large compute scale to generate diverse prompt injection attacks against its own models. OpenAI reports using GPT-Red to adversarially train GPT-5.6 Sol, claiming 6x fewer failures on its hardest direct prompt injection benchmark versus a production model from four months earlier.