Threat · curated 5 Oct 2026

When the pentester is a fleet of AI agents: inside an autonomous vuln-hunting rig

Dossier

Coverage timeline

5 Oct 2026huntback.io

Single-source incident — first reported, latest, and curated coincide.

Why it matters

The "Brainstorm" rig demonstrates that real-world adversaries are already operationalizing fleets of autonomous LLM agents to find, confirm, and triage vulnerabilities at machine speed, with model-backend swapping (Anthropic/Kimi) that defeats provider-domain detection.

Huntback recovered the working tree of an offensive-security operation run almost entirely by AI agents from an exposed directory on a Contabo server (37.60.248.174). The platform, called "Brainstorm," orchestrates a fleet of named Claude Code workers that scan, read source, and triage vulnerabilities autonomously, backed by a home-built out-of-band collaborator confirming blind XXE/SSRF/SSTI/RCE/deserialization, a nuclei template library, SAST taint configs, and responder for LLMNR/NBT-NS poisoning and relay pointing toward internal intrusion.

campaign

Summary

Researchers at huntback indexed an exposed open directory on 37.60.248.174, a Contabo GmbH host, and recovered 365 files representing the working tree of an offensive-security operation run almost entirely by AI agents. The operator's self-built 'Brainstorm' platform orchestrates a fleet of Claude Code worker agents that scan targets, read source code, fire payloads and triage findings in an automated loop, writing to a transcript the human operator reads later.[0]

The rig pairs a legitimate-looking automated code-review layer with unambiguously offensive capability: a home-built out-of-band collaborator for confirming blind vulnerabilities, a nuclei template library, SAST taint configs, and responder for LLMNR/NBT-NS poisoning and relay. The presence of relay tooling, five C2-style config files and harvested secrets points past external bug-bounty toward internal-network intrusion, though huntback explicitly leaves the bug-bounty-versus-malicious question open.[0]

A key structural finding is that the model backend is interchangeable — the code references Anthropic and Kimi (Moonshot) side by side — so detection keyed to one AI provider's API domains ages quickly. huntback recommends defenders focus on victim-side behavior and opsec leftovers such as autonomy flags set to 'yolo' or bypass permissions.[0]

Attack chain

  1. Deploy platform: The operator stands up the self-built 'Brainstorm' orchestration platform on the VPS.[0]
  2. Spawn agent fleet: The orchestrator spawns named Claude Code worker agents ('germinal','strategist','Ask Code'), each given a role in the vulnerability-research loop, with the Serena LSP wired in via MCP for semantic code navigation across 40+ languages.[0]
  3. Scan and SAST: Agents run active scanning using the nuclei template library and SAST taint 'sink' configs to pull dangerous call-sites out of source at scale.[0]
  4. Confirm via out-of-band: Each emitted payload is registered with a unique token so callbacks to the self-hosted DNS/HTTP/SMTP/LDAP interaction logger map back to the originating request, confirming blind XXE/SSRF/SSTI/RCE/deserialization.[0]
  5. Triage findings: The 'Ask Code' assistant with read access to the VPS and authenticated access to the Brainstorm HTTP API mass-triages the Code findings; results are written to a transcript the operator reads later.[0]

Disclosure timeline

DateEvent
2026-09-30huntback sensors place the host 37.60.248.174 as active; observed activity window begins and ends same day.[0]
2026-10-05huntback publishes the teardown of the exposed offensive-automation rig.[0]

Indicators of Compromise

TypeIndicatorContext
ip37.60.248.174Contabo GmbH host whose exposed HTTP root contained the offensive-automation rig; observed 2026-09-30.[0]
otherUnexpected DNS/LDAP/SMTP egress to a single external token domainThe tell of blind-vuln confirmation via the self-hosted out-of-band collaborator; defenders should watch their own egress for such callbacks.[0]

Key takeaways

  • AI-assisted attacks in practice look like a single human pointing a fleet of capable agents at a target to scan, read code, fire payloads and triage callbacks in a loop, compressing recon and triage from days to minutes.[0]
  • Because the AI backend is interchangeable by config, durable detection should key on victim-side behavior and exposed opsec leftovers rather than any one model provider's infrastructure.[0]

Defensive actions

  • Disable LLMNR/NBT-NS and enforce SMB signing.: Removes the internal-relay path that the rig's responder tooling (T1557.001) depends on, so poisoning/relay has nothing to catch.[0]
  • Monitor outbound egress for out-of-band callbacks to a single external token domain across DNS, LDAP and SMTP.: Such callbacks are the signature of the self-hosted interaction logger confirming blind vulnerabilities.[0]
  • Deploy deception/decoys.: Any touch of a decoy during machine-speed automated scanning is high-confidence hostile signal with no false positives.[0]
  • Detect provider-agnostically on behavior and opsec leftovers rather than a single AI provider's API domains.: The model backend is swappable by config (Anthropic, Kimi/Moonshot), so provider-domain detection rots; autonomy flags such as approval mode 'yolo' or bypass permissions survive a backend swap.[0]