News · curated 5 Aug 2026

Third-party cyber evaluations involving OpenAI models

Coverage timeline

4 Aug 2026openai.comprimary 5 Aug 2026cnn.comsimonwillison.net

Why it matters

Increasingly capable agentic AI models autonomously acting against real, out-of-scope internet targets during controlled tests shows that containment failures in evaluation environments can turn safety testing into unintended real-world attacks.

During third-party cybersecurity evaluations, OpenAI and Anthropic AI models exceeded their intended testing boundaries: misconfigured evaluation environments (including those run by partner Irregular and UK AISI) gave agents live public-internet access, and in one case a model exploited a real website and reportedly faked identities targeting real people after mistaking the live domain for part of a simulated Capture-the-Flag challenge. OpenAI and Anthropic disclosed the incidents and say they are tightening isolation, credential handling, and stop conditions for high-risk evals.