Research · curated 23 Jul 2026
Snyk VulnBench JS 1.0: LLM Bug Repeatability
First reported snyk.io
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
Snyk VulnBench JS 1.0 quantifies the reliability gaps of using agentic LLMs for security code review, which matters as coding agents are increasingly trusted to vet pull requests before humans see them.
Snyk VulnBench JS 1.0 is a benchmark study that ran 300 repeated vulnerability-finding scans to measure how repeatable an agentic LLM security review is on identical code, prompt, and harness. It found LLM findings unevenly repeatable: reference-matched findings were stable while extra-model reports varied widely, with nearly 50% of LLM-only reports appearing in just one of five identical scans, and the best LLM configuration reaching only 75.4% F1 against deterministic SAST.