Research · curated 23 Jul 2026

Snyk VulnBench JS 1.0: LLM Bug Repeatability

Coverage timeline

29 Jun 2026snyk.ioprimary

Single-source research — first reported, latest, and curated coincide.

Why it matters

Snyk VulnBench JS 1.0 quantifies the reliability gaps of using agentic LLMs for security code review, which matters as coding agents are increasingly trusted to vet pull requests before humans see them.

Snyk VulnBench JS 1.0 is a benchmark study that ran 300 repeated vulnerability-finding scans to measure how repeatable an agentic LLM security review is on identical code, prompt, and harness. It found LLM findings unevenly repeatable: reference-matched findings were stable while extra-model reports varied widely, with nearly 50% of LLM-only reports appearing in just one of five identical scans, and the best LLM configuration reaching only 75.4% F1 against deterministic SAST.