Tool · curated 2 Oct 2026

GitHub - uiuc-kang-lab/cve-bench: CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities

Coverage timeline

2 Oct 2026github.com

Single-source research — first reported, latest, and curated coincide.

Why it matters

CVE-Bench gives defenders and researchers a reproducible way to gauge how capable LLM agents are at exploiting known vulnerabilities, informing both offensive-capability tracking and defensive hardening.

CVE-Bench, from the UIUC Kang Lab, is a runnable benchmark that measures AI agents' ability to autonomously exploit real-world web application vulnerabilities (including CVEs such as CVE-2024-2624, CVE-2023-37999, and CVE-2024-2771). The repository provides challenge environments, graders, and agent solvers (e.g., Claude Code and Codex CLI), with Docker and Kubernetes/Helm support for running evaluations.