Penetrify, a Czech company building autonomous AI penetration testing, solved all 104 challenges in the XBOW benchmark (XBEN) with a black-box agent, achieving 100% across 26 vulnerability classes. The company emphasizes transparency by publishing all per-challenge transcripts and also discloses its weaker performance on CVE-Bench.
-- Brno, Czech Republic.
Penetrify, a company building autonomous AI penetration testing, ran its engine against the full XBOW benchmark (XBEN) and solved all 104 web-security challenges. The run was black-box: the agent got a running application and nothing else, with no source code, no hints, and no human at the keyboard.

Grading is deliberately hard to fake. Each target hides a unique, unguessable flag, and a challenge counts as solved only when that exact flag appears in the agent's own log. Every difficulty level came out at 100 percent, across 26 vulnerability classes including IDOR, SSTI, command injection, XXE, SSRF and deserialization.
Penetrify is deliberate about not overselling the number. XBOW, which publishes the suite, has said it is now largely solved across the industry, and the company states the same on its own results page.
"A perfect score here is a sanity check on our engine, not a claim that we solved web security," said Penetrify founder Viktor Bulanek. "What matters is what it costs, how it runs, and whether you can verify it."
Each of the 104 pentests ran in the company's fastest and cheapest tier, the same one most customer scans use. The engine averaged 9.7 minutes per challenge and about 29 US dollars per scan at list price, and the full suite took roughly 16.8 hours of compute.
For comparison, a single traditional human web-application pentest usually costs into the five figures and takes one to three weeks. At this cost and speed, a full pentest can run on every release inside a CI/CD pipeline (https://www.penetrify.cloud/en/ci-cd/) instead of once a quarter.
The transparency is the part Penetrify stresses most. The company released the test harness and all 104 per-challenge transcripts, sanitized only for internal identifiers, so anyone can rebuild the targets and run the suite again.
The company is also open about the caveats. Grading is lenient by design, the way XBOW intends, and Penetrify published a second, far less flattering result next to this one. On CVE-Bench, a benchmark of real-world CVEs with much stricter grading, the same engine performs markedly worse, and the company posted those numbers rather than leaving them out.
"Any vendor can show you a green benchmark," Bulanek said. "Fewer will show you the run that went badly, the exact grading, and the raw logs."
Full results and the downloadable harness and logs are published at https://www.penetrify.cloud/en/benchmark, with a longer write-up on the company blog.
About Penetrify
Penetrify is a Czech company building autonomous AI penetration testing. Its engine attacks a running web application the way a human tester would, black-box and unattended, and reports validated findings with reproduction steps. Tests run in minutes and integrate into CI/CD pipelines, at a fraction of the cost of a traditional manual engagement.
Contact
Penetrify Viktor Bulanek, Founder Brno, Czech Republic [email protected] https://www.penetrify.cloud
Keywords: AI penetration testing, XBOW benchmark, web application security, autonomous penetration testing, cybersecurity, pentest automation

Contact Info:
Name: Viktor Bulanek
Email: Send Email
Organization: Penetrify.cloud
Address: Nove sady 988/2, 602 00 Brno
Website: https://www.penetrify.cloud
Release ID: 89202689
Should there be any problems, inaccuracies, or doubts arising from the content provided in this press release that require attention or if a press release needs to be taken down, we urge you to notify us immediately by contacting [email protected] (it is important to note that this email is the authorized channel for such matters, sending multiple emails to multiple addresses does not necessarily help expedite your request). Our efficient team will promptly address your concerns within 8 hours, taking necessary steps to rectify identified issues or assist with the removal process. Providing accurate and dependable information is central to our commitment.






