Tool · curated 2 Oct 2026
GitHub - uiuc-kang-lab/cve-bench: CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities
First reported github.com
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
CVE-Bench gives defenders and researchers a reproducible way to gauge how capable LLM agents are at exploiting known vulnerabilities, informing both offensive-capability tracking and defensive hardening.
CVE-Bench, from the UIUC Kang Lab, is a runnable benchmark that measures AI agents' ability to autonomously exploit real-world web application vulnerabilities (including CVEs such as CVE-2024-2624, CVE-2023-37999, and CVE-2024-2771). The repository provides challenge environments, graders, and agent solvers (e.g., Claude Code and Codex CLI), with Docker and Kubernetes/Helm support for running evaluations.