Research · curated 15 Jul 2026
GPT-Red: Unlocking Self-Improvement for Robustness
First reported · updated · 2 reports openai.com
Coverage timeline
Why it matters
GPT-Red demonstrates automated, scalable adversarial red-teaming for prompt injection, a technique defenders can mirror to harden LLM agents that ingest untrusted third-party data via browsers, tools, and files.
OpenAI describes GPT-Red, an internal automated red-teaming model trained at large post-training compute scale to discover prompt injection vulnerabilities in its models and generate adversarial training data. OpenAI reports using GPT-Red to adversarially train GPT-5.6 Sol, claiming 6x fewer failures on its hardest direct prompt injection benchmark versus a prior production model.