XRanges for AI: A Realistic Benchmark for Autonomous Security Agents

Learn how XRanges provides a realistic benchmark for AI security agent evaluation, ensuring accurate and reliable autonomous pentesting results.

A security researcher manually verifying an autonomous agent's findings on a laptop.

Intro

I have spent too many nights reading reports written by autonomous pentesting agents that sound confident while hiding the fact that half their findings were wrong. The problem is not that the agents are bad. It is that we measure them with the same tool they use to report their own work. An agent writes a PDF. A human with a security background opens it and checks every claim against a live target. Which findings are real. Which are duplicates. Which the agent invented out of thin air. And what it never even tried to touch. That is a day of expert work for one run. Multiply it by three models, four prompt variants and ten repetitions each, and you are looking at a review queue longer than your experiment timeline. This is what makes AI security agent evaluation such an awkward problem for anyone running these tools in production.

Most teams I talk to are actually doing bug bounty style automated scanning now, not just CTF puzzles. They point an agent at a staging app, let it run overnight, and then sit with the output in the morning. The output is almost always a narrative. It says the agent found an access control flaw, maybe shows a curl command or two, and moves on. What it does not show is the forty legitimate features the agent never opened, the API endpoint where it found the first bug but ignored the second bug sitting right next to it, or the fact that it deleted the user table on its way out. No client would accept that result, and nothing in a findings list records it.

Manual expert review works fine when you have one run to check. It breaks down under the experiment matrix an AI engineering team actually needs. XRanges for AI exists inside that gap. It is built by CTF.ae and hosted at ai.xranges.com. Instead of letting the agent write its own report, XRanges instruments every service in a benchmark target with hand-written OpenTelemetry telemetry and scores four independent signals live while the agent is still working. Coverage, boundaries, exploitation, and integrity. You get numbers, not prose, and they update in real time. I have seen this flip the workflow from reading paragraphs to looking at a dashboard that tells you which features your agent ignored and which rules it broke before it even finished. That shift matters when you are comparing models across multiple targets and trying to figure out which configuration actually improves your agent instead of just making it sound more confident.

Background

Autonomous pentesting agents have gotten genuinely better over the last couple of years. They can chain vulnerabilities across service boundaries, discover misconfigurations that would take a human twenty minutes to map out, and produce findings that are often accurate enough to route straight into a ticketing system. The technology has moved fast. The evaluation methods have not.

The workflow most teams land on is straightforward in theory. You spin up a target application that looks like something a real company would run. Maybe it is an internal tool with authentication, job posting management, and a shared workspace feature. You point your agent at it and give it the rules of engagement. It runs. Then you read whatever it outputs.

The output is where this workflow falls apart. An agent will return a report that sounds authoritative. It lists vulnerabilities with timestamps, curl commands, and confidence scores. The prose reads like a consultant wrote it. What the report does not contain is nearly as important as what it does contain. It never mentions the fifteen user-facing features it skipped entirely. It never flags the API endpoint adjacent to the one it exploited where a second vulnerability sits waiting. And if the agent trashed data while hunting for a finding, the report usually frames that as a successful discovery rather than a rule violation.

A human reviewer can catch most of this. Someone with security experience sits down with the target and walks through the claims. Which findings are real, which are partial exploits, which are flatly invented. This works for one run. I have done it myself and it takes about two hours for a single target with a moderately complex agent.

It does not work when you multiply by three models, four prompt variants, and ten repetitions. That is forty runs requiring roughly eighty hours of expert time before you have enough data to say anything useful about which configuration actually performs better. The review queue becomes the bottleneck, not the agent development. Any team doing serious AI security agent evaluation hits this wall within a few weeks.

What’s happening now

Q: What exactly is XRanges for AI and how does it work?

XRanges for AI lives at ai.xranges.com (opens in new tab), built by the CTF.ae team. It solves the evaluation bottleneck by combining two things into a single workspace: realistic benchmark targets with instrumentation baked in, and a live scoring layer that tracks what your agent actually does rather than what it claims to have done.

Each benchmark target is a full multi-service application. Think of it like a small company website with a job board, an assessment platform, shared conversations, and background jobs. It runs across several languages and frameworks because that mirrors how real software gets built. Every target carries twenty or more injected vulnerabilities, ranging from single-step flaws to chains that cross service boundaries. Some of these are zero-days discovered by CTF.ae’s own researchers. None of it appears in public training corpora, which matters more with every passing month as agents scrape the web more aggressively.

The second piece is the telemetry layer. Every service in every target emits structured telemetry through OpenTelemetry. This is hand-written instrumentation, not generic HTTP logs. A standard access log will tell you a request hit an endpoint. It will not tell you whether the agent traversed the right authorization chain to reach it, or whether it stopped halfway through a multi-step exploit. The instrumentation answers those questions. The platform ingests the telemetry per deployment and turns it into four scores that update in real time while the agent is still running.

Q: How does scoring actually work inside XRanges?

The four signals were chosen specifically so they do not overlap. An agent cannot improve one score by gaming another, and that independence is the entire point of the design.

Coverage measures how thoroughly the agent explored the target. Every user-facing feature becomes a coverage point phrased as a business action rather than a URL. Did the agent register an account? Browse job postings? Open a shared conversation? Run code in an assessment? These points are only reachable through normal use, never through an exploit, so the score reflects genuine exploration of the legitimate surface. The unhit points are often more revealing than the score itself. Most teams find that list tells them exactly where their agent has blind spots.

Boundaries tracks rule violations. Each target ships with guard rules such as “must not delete hiring content” or “must not revoke API keys.” A violation is recorded the moment it happens, with the container and timestamp attached. Zero violations is the expectation. Anything else is a finding about the agent, not the target, and usually the more urgent one.

Exploited records which vulnerabilities the agent actually exploited. Every vulnerability is defined as a kill chain of ordered phases. Because each phase is detected from inside the target, the platform knows exactly which step the agent completed and where it stalled, regardless of what the agent wrote in its report. A three-step access control chain that stopped at step two shows up as two of three, with timestamps.

Integrity checks run every minute and confirm the application is still functionally correct. Seed data present. Services answering with the right content. Cross-service trust intact. A failed check is a penalty regardless of cause. It catches the agent that found a bug by breaking the environment around it.

What It Means in Practice

Let me walk through what a live XRanges evaluation actually looks like from start to finish, because the difference between reading about it and watching it run is significant.

You deploy a target from the library. It spins up as a multi-service application, each component emitting telemetry through OpenTelemetry. You point your agent at the endpoint. It starts working. And while it works, four scorecards update in real time. Not after the fact. Not in a report you review tomorrow. Now.

The coverage score tells you what the agent touched and what it missed. Here is where the design choice matters. Coverage points are phrased as business actions: registered an account, browsed job postings, opened a shared conversation, ran code in an assessment. They are only reachable through normal use, never through an exploit. That is intentional. It keeps the metric honest about exploration versus exploitation.

Most teams I talk to end up spending more time on the unhit list than on the score itself. The score might read 72 percent and feel acceptable. But the unhit points are where the agent quietly failed. Maybe it never touched the admin panel. Maybe it never exercised the file upload path. That second one is where an attacker would go. Your agent did not even look there.

The boundaries score catches the agents that find a vulnerability by breaking everything around it. I have seen this in my own work. An agent hunting for SQL injection in a booking system will sometimes drop tables along the way, or revoke API keys belonging to other services, or delete seed data that simulates real user content. In a traditional workflow, nobody notices until the client complains that the demo environment is broken. With XRanges, the violation is logged the moment it happens, container and timestamp included. Zero violations is the baseline. Anything else is a finding about the agent.

A full run takes roughly twenty to forty minutes depending on the target complexity and agent behavior. You get back four live scores, a complete unhit feature list, any boundary violations with exact timestamps, and a phase-by-phase breakdown of each exploited vulnerability. The comparison between three different models on the same target becomes a single screen instead of three separate reports you try to reconcile manually.

There is a limitation I want to flag. This has been stress-tested against CTF.ae’s own targets. Those are carefully instrumented, intentionally seeded applications. I do not know how this holds up against custom internal applications where the attack surface is messier, the telemetry gaps are larger, and the rules of engagement were written by humans who missed edge cases. If you are running this against your own production-like environments, expect to spend time mapping your guard rules before the first real evaluation.

That said, the end-to-end experience is faster than anything I have used for comparing autonomous pentesting agents. Even accounting for the custom setup time, a single XRanges run produces more actionable data than three days of manual review.

What to expect next

The obvious risk with a benchmark platform like this is that the targets leak into training data. As autonomous security agents get more aggressive about web scraping and content harvesting, any publicly deployed target application becomes a liability for the people who built it. CTF.ae has been working with pentesters for years, and they have likely already thought about this, but it is worth watching. Once a target gets thoroughly crawled, the whole coverage metric loses value. The platform will need to rotate or version its benchmarks faster than most teams can consume them.

I expect the four-signal model to either become a quiet standard for AI security agent evaluation or to stay niche behind a wall of proprietary targets. The independence of the signals matters because it prevents gaming. An agent could theoretically optimize for coverage by browsing every page without finding anything. It could optimize for exploitation by brute-forcing endpoints and missing legitimate features entirely. Those single-dimensional approaches no longer work when all four scores are tracked independently and live. That structural honesty is what separates this from the usual benchmarking playgrounds where teams mine a leaderboard by tuning prompts to a fixed test set.

Another thing to watch is whether the telemetry layer becomes portable outside of CTF.ae’s stack. Right now the platform ingests OpenTelemetry per deployment, but the guard rules and coverage points are tied to each specific target. If a team wants to run this against their own internal applications, they have to build or buy that instrumentation work themselves. I would not be surprised if the market splits between one-click CTF.ae targets and custom deployments where the evaluation pipeline looks a lot like the messiness teams already deal with in production monitoring.

Before adopting this into your own AI security agent evaluation workflow, I would start by running a single model against one of the published targets and comparing the output side-by-side with what your current manual process produces. Not everything needs to be automated for you to see where the gaps are. The unhit feature list alone often reveals blind spots that a findings report never surfaces.

Take the next step

If you run pentesting agents or bug bounty automation yourself, the fastest way to understand whether XRanges for AI would actually save you time is to stop reading and launch a real run. The platform lives at ai.xranges.com and does not require a paid account just to deploy one of the benchmark targets and watch an agent work against it cite (opens in new tab).

Start with a single model and a single target. Deploy the target through the dashboard. Configure your agent to run against the live application using the connection details provided. Let it finish. Then open the live scoring view while the agent is still working, not after it has already stopped. Watching the four signals move in real time is where the whole exercise becomes useful. The coverage counter will tick up as the agent touches legitimate features. The boundaries panel will stay blank unless something breaks. The exploited score will climb in discrete steps as the agent completes actual kill chains. The integrity metric will hold steady unless the environment itself starts to degrade.

Then do it again with a second model against the same target. Run both at the same time if the platform allows it, or run them sequentially and compare the final screens side by side. Most of what matters lives in the differences between the two runs, not in either score in isolation.

After that comparison, take your current manual review process and lay it out next to what the dashboard produced in under twenty minutes. I would expect you to notice three gaps immediately. First, the unhit coverage points. Those are the features your agent never touched, and they reveal blind spots a findings report never shows. Second, the per-phase exploit progress. Your agent might have claimed it found a vulnerability. The telemetry shows it reached phase two of a three-step chain and stalled, which changes how you interpret the result entirely. Third, the violation log, even when it stays empty. Knowing your agent respected the rules of engagement is as important as knowing what it found, and nothing in a PDF report captures that.

If the whole exercise takes less than an hour and reveals more than your last round of manual reviews, you have your answer. If it does not, you now know exactly what part of the pipeline needs fixing before you invest further.