Proof

Findings are cheap. Proof is the product.

Anyone can hand you a list. The value is in the finding that has already been reproduced, with the exact steps that make it true, and a person willing to sign it. That is the standard the agent is built to.

The Standard

Reproduce it, or it does not exist

Before a finding reaches you, the agent proves it against the live target: it re-runs the exploit, captures the request and response, and records the steps to repeat it. Anything it cannot reproduce is never shown to you.

Every result carries a CVSS v3.1 vector you can verify yourself, the components it affects, and remediation aimed at the layer that must change. When you want a person accountable for it, a senior practitioner validates it before it lands. If you want to see the format before you commit to anything, see a sample report.

CRITICAL Broker token in JS bundle grants admin API
  • discovered a public JavaScript bundle served by the API
  • extracted a broker token committed into the bundle
  • replayed the token against the admin API and authenticated
  • reached 14,213 order records across tenant boundaries
SeverityCritical · CVSS 9.1
EvidenceRequests, responses, repro steps
VerifiedReproduced 2 of 2
StandardOWASP API Security Top 10

Example finding. Target details redacted. Illustrative example, not a specific customer result. It reads like a mix of BOLA and BFLA findings, the authorization flaws scanners miss most often.

The Difference

What runnable actually means

Most pentest reports offer narrated proof: screenshots, a captured session, a paragraph asserting "confirmed exploitability," and reproduction steps written for a human to follow by hand. That is proof of exploit as a story, and a story has to be trusted and then re-typed by whoever needs to check it. A runnable proof-of-concept is different in kind, not degree. It is a script or request chain the agent already executed against the live target, packaged so the customer can paste it into their own terminal and watch the same exploit fire again, with no interpretation step in between.

That distinction is also the honest answer to the question of whether this is just glorified vulnerability scanning: a scanner outputs a probable weakness for a human to chase down and confirm, while exploit validation here means the agent has already chained the request sequence, authenticated as the right role, and captured the exact traffic that broke the boundary, so the artifact you receive is the exploit itself rather than a claim about one. A finding without a runnable artifact behind it is treated as unproven and never shown to you.

This matters most when a finding is disputed. A narrated writeup invites a debate about whether the tester's environment, session, or assumptions matched production. A runnable proof-of-concept sidesteps the debate: your engineers run it against the same target and see the same result, on their own machine, on their own terms. That is a materially shorter path from finding to fix.

Every Finding Carries

Evidence a reviewer can act on in minutes

  • Reproduction steps: the exact sequence to trigger the issue again, by hand
  • Request and response data: the traffic that proves the behavior, not a description of it
  • CVSS v3.1 vector: a severity you can recompute, not a number we assert
  • Affected components: the assets, endpoints, and parameters in scope of the finding
  • Attack chain: how individually low findings combine into the path that matters
  • Remediation: guidance aimed at the layer that must change, written for engineers
  • Standard mapping: the recognized attack class the finding traces back to
  • Retest: confirmation that a fix closed the class, not just the one sample
How We Keep It Honest

Two checks before a finding is yours

Autonomy is only useful if you can trust the output. The agent verifies its own work, and you can add a human signature on top.

Machine

Independent verification

A separate step re-exploits each candidate finding against the target and discards anything it cannot reproduce, the way a peer reviewer would before a report goes out.

Method

Tested against known-vulnerable targets

Our test library is exercised against public, intentionally vulnerable applications and standard benchmarks, so behavior is measured against ground truth rather than our own marketing.

Human

Validation on demand

Route any finding, or an entire run, through a senior practitioner before it reaches your tracker, when a framework or a stakeholder needs a person to stand behind it.

Standards

Traceable to recognized method

A finding maps to a published attack class, and a severity you can verify, rather than a tester's improvisation.

OWASP WSTG OWASP API SECURITY TOP 10 OWASP ASVS OWASP LLM TOP 10 PTES NIST SP 800-115 MITRE ATT&CK CVSS V3.1
Get Started

See a real run against your surface

Give us a domain and the rules of engagement. We will return a scoped run and a sample of exactly what a proven finding looks like.