Tool · curated 23 Jul 2026

GitHub - GiovanniGatti/cve-bench: A benchmark for evaluating AI agents on fixing real-world security vulnerabilities.

Coverage timeline

23 Jul 2026github.com

Single-source research — first reported, latest, and curated coincide.

Why it matters

cve-bench gives defenders a reproducible way to measure how reliably AI coding agents can remediate real vulnerabilities, informing whether such agents can be trusted in security-critical patching workflows.

cve-bench is a benchmark by GiovanniGatti for evaluating AI agents on their ability to fix real-world security vulnerabilities, shipping a Docker-based harness, results, and a write-up comparing model performance across CVEs such as CVE-2026-33175, CVE-2026-42561, CVE-2026-40864, and CVE-2026-30930.