Tool · curated 25 Jul 2026

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

Coverage timeline

25 Jul 2026cybergym.io

Single-source research — first reported, latest, and curated coincide.

Why it matters

ExploitGym demonstrates and measures autonomous exploit-generation by AI agents, quantifying how far frontier models can go from bug report to working exploit even against standard mitigations — a capability that lowers the barrier for offensive misuse and reshapes defender threat models.

ExploitGym is a benchmark of 869 tasks from the UC Berkeley sunblaze group (published with a GitHub repo) that measures whether AI agents can transform a known vulnerability and a proof-of-vulnerability input into a working end-to-end exploit across userspace, browser V8, and Linux kernel targets. A leaderboard scores frontier coding agents on how many exploits they produce, including bypasses of ASLR, stack canaries, and the V8 heap sandbox, and the authors describe the capability as inherently dual-use.