Intro
$255.5 million. That is the Series B Armadin just closed, at a valuation above $2.5 billion, for a company that is roughly a year old. Kevin Mandia founded it. He also founded Mandiant, the incident response firm Google bought for $5.4 billion in 2022, and that history is why institutional money shows up this fast and this large.
The plain definition of agent swarm security is this: instead of paying a human crew to break into your systems once a year, you run teams of AI agents that attack your environment continuously, share what they learn with each other, and chain small weaknesses into a working attack path.
This is written for the people who carry that work. If you own a production environment, if the patching queue has your name on it, if you have signed off on a penetration test invoice and then watched the findings sit untouched for eight months, you are the audience. Sales decks about autonomous security are built for your VP. The operational questions land on you.
Two things worth pulling apart before going further. The funding news tells you where enterprise security budgets are heading. Whether the technology does what the pitch claims is a separate question, and not one anybody outside the early customer list can answer with confidence today.
I have run external scans across my own VPS fleet for years, and the pattern is boring. An outdated PHP build on a subdomain someone spun up for a demo. A forgotten staging host still accepting default credentials. A certificate that expired four months ago because the renewal reminder went to an address nobody reads. None of that is clever. All of it is exploitable. The appeal of continuous testing is that it catches the dull stuff on the day it appears rather than the week an auditor asks about it.
Whether swarms of agents do that job better than a cron job and a scanner subscription is the part I keep turning over. The dollar figure says investors believe it does.
Background
Kevin Mandia built Mandiant into the name enterprises called when a breach was already underway, and Google bought the company for $5.4 billion in 2022. TechCrunch leads its report on the new round (opens in new tab) with that history, which is the right place to start. It explains the velocity better than any product walkthrough could.
Institutional money in security does not really buy a feature list. It buys the person who has answered the phone at 3 a.m. for a bank that just lost its domain controller, and who can tell a convincing story about why the next five years look structurally different from the last ten. Mandia has that story. Most founders walking into enterprise security do not, and the spread between a $190M Series A and a $255.5M Series B six months later is partly the price of that credibility.
Against that, look at what a normal penetration test actually leaves you with. It is a scoped engagement with a start date, an end date, and a rules-of-engagement document saying what the testers are allowed to touch. Two or three weeks of human creativity, then a PDF. The findings are usually real and often uncomfortable in a useful way. They are also a photograph of a moving thing.
On my own infrastructure, a scan run in January and the same scan run in June are different documents. Not because the scanner improved. Because the boxes changed underneath it: a package bumped, a subdomain repointed, a service that someone turned on for one client and never turned off. The report was accurate on the day it was signed and drifted from there.
The other half of Armadin’s pitch is the part I find more interesting than the tooling. The company argues that bad actors, and eventually AI labs with agents nobody is watching closely, will use this same agentic capability offensively. I cannot verify how far along that is, and plenty of vendor decks use hypothetical future attackers to justify present spending. But the underlying shift is real enough. When attack execution scales on compute rather than on the number of skilled humans available, annual testing stops being a defensible cadence.
What’s happening now
The numbers themselves are worth laying out plainly, because the structure of the round says as much as the total. Armadin announced a $255.5 million Series B at a valuation above $2.5 billion. Andreessen Horowitz and Accel co-led, with Bain Capital Ventures, Redpoint, 8VC, Ballistic Ventures, Google Ventures, In-Q-Tel, Kleiner Perkins, and Menlo Ventures all participating. In-Q-Tel showing up in that list is not decoration. It is the strategic investment arm tied to US intelligence community interests, and it tends to appear in security companies it expects to matter to government buyers.
The timing is the part that gets attention. Armadin raised a $190 million Series A in March. This round closed six months later, in October, bringing total funding past $445 million. A company that has been operating for roughly a year has raised more than most security vendors raise across a decade, and it did so without the usual multi-year gap between rounds. Either the pipeline is unusually strong for that stage, or the market for agentic security testing is being priced on expectation rather than revenue. I cannot tell which from the outside, and the TechCrunch report on the raise (opens in new tab) does not break out customer counts or ARR, which is the number I would want before judging any of this.
What the product claims to do is specific enough to be checkable later. Instead of a scheduled engagement, Armadin runs always-on agentic swarms that attack an environment continuously. The agents chain vulnerabilities together, meaning they take findings that a scanner would report as low or informational in isolation and combine them into a working path. Armadin frames the purpose as finding holes before a human attacker, or another lab’s unsupervised agents, reach them first.
Chaining is the interesting claim. On a shared hosting setup, that could look like a stale PHP script on one subdomain leaking a database credential that also works for a staging panel, which then exposes an API key for something else entirely. None of those three items is severe alone. Together they are a breach. Whether a swarm finds that reliably across a messy production environment is an open question, and it is the one I would want answered with a trial on real infrastructure rather than a demo tenant.
Agent Swarm Security, Explained
What it actually is
Agent swarm security is a method of testing where multiple autonomous agents run in parallel against the same environment, share what each one learns, and combine low-severity findings into a single working attack path. The second half of that sentence is the part that matters. A single agent on its own is a smarter scanner. A swarm that pools intermediate results is doing something closer to what a red team does, and it does it on a schedule that never stops.
Continuous is not a marketing word here. It changes the artifact you get. A pentest produces a PDF with a date on the cover. A swarm produces a stream. Findings show up when they become true, which means the useful life of a report is measured in hours rather than in quarters.
How is it different from a vulnerability scanner or a normal pentest?
A scanner matches known signatures against known versions and stops there. It will tell you that a PHP release is old. It will not tell you whether someone can reach the file that uses it, whether that file leaks a credential, or whether the credential is reused somewhere that matters. Most of the noise in a typical scanner output is exactly this gap.
A pentest covers the gap but only for the days you booked. Human testers bring creativity that no agent has matched yet, and then they leave, and your environment keeps changing.
Agents sit between the two ends. They attempt exploitation rather than matching strings, they adjust when a technique fails, and they keep running when you push a new deploy on Tuesday afternoon. That last property is the one I would pay for, because deployments are when my own hosts tend to break in ways nobody intended.
Does this replace human pentesters?
For repetitive enumeration and known exploit paths, largely yes, and the economics point that way whether testers like it or not. For business logic abuse, anything requiring context about how your company actually operates, I would not bet on it yet. An agent can learn that an endpoint accepts a negative quantity. It cannot easily learn that the negative quantity is a refund fraud path specific to your billing rules. That judgement still comes from a person who has read your code.
The likelier outcome over the next two years is that human testers spend more of their hours validating what the swarm reports and less of them running nmap by hand. Fewer testers, more reviewers.
What Agent Swarm Security Means in Practice
A continuous testing platform produces continuous findings. That is the whole selling point and also the part nobody puts in the pricing conversation. Your queue gets a new entry every morning instead of a PDF once a year, and if no one owns triage on a daily basis, you have not bought security. You have bought a faster way to build a backlog. On my own boxes I already carry a mental list of low-severity items I keep meaning to fix, and adding an agent that surfaces ten more per week does not help me unless I also add the process to close them.
Here is the kind of chain a swarm is actually good at, on a setup a lot of my readers run. A VPS customer has WordPress on a cPanel account. Their site runs an abandoned backup plugin from 2019. The agent finds a path traversal in it, reads wp-config.php, pulls the database credentials, notices the same password works for the cPanel login because the customer reused it, and now has the whole hosting account. Every individual step is boring. The vulnerability scanner flags the old plugin version and stops. The annual pentest would have caught it in the three days it was booked, and missed it in the eleven months after you installed the plugin. That chain is the pitch, and it is a legitimate one.
Then there is the part the vendor cannot do for you. Anything you let probe production needs scoped credentials, a defined blast radius, and a kill switch that actually works when someone hits it at 2 a.m. Deciding which internal hostnames the agent may touch is an internal conversation about your own architecture. I have watched people hand a monitoring tool a root-equivalent token because it was faster than mapping the permission model, and that decision is how a testing tool becomes an incident.
For single-operator shops and small hosts, my honest read is that a full agentic platform is out of reach on price right now. Continuous external scanning plus a manual test once or twice a year covers most of the realistic exposure for a few hundred dollars a year. Where I genuinely do not know how this holds up is messy environments: old cPanel stacks, custom billing panels, the half-documented internal app someone built in 2016. Automated testing tends to disappoint exactly there, and that is where a large chunk of small business infrastructure lives.
What to expect next
Within a year I expect every established scanner vendor to ship something with “agent” on the label, and most of those products will mean something narrower than what Armadin is describing. A scanner that runs a language model over its existing findings and writes a friendlier remediation paragraph will get called an agent in the launch post. In practice it is a report generator with a better vocabulary. The distinction worth testing in a demo is whether the thing actually attempts exploitation, adapts when an attempt fails, and passes what it learned to the agents running beside it. Ask that directly and watch how quickly the conversation moves to roadmap slides.
Pricing and packaging will be the battleground. Per asset, per agent, per finding, or one flat platform fee, and the gap between those models gets ugly once you pass a few dozen hosts. The number I would watch is where continuous testing lands in the budget, because if enterprises start paying for it out of the annual pentest line item instead of the tooling line item, the manual testing shops have a real problem on their hands. My expectation is that the first contracts get sold as a supplement to the pentest and quietly become the thing auditors ask about two renewal cycles later.
Liability is the piece nobody is pricing yet. When an autonomous agent knocks a service over during a test, the clause about who owns that outage matters more than the detection rate on the slide. A swarm holding production credentials with no rate limit is a self-inflicted denial of service waiting for a Tuesday deploy window. I would want the vendor’s incident response and notification timeline written into the contract, not described verbally on a call, and I would test my own kill switch before the first run rather than during it.
I also expect consolidation. Armadin has the funding and Mandia’s reputation to set the reference point for what agentic pentesting is supposed to look like, and the larger platform vendors will either buy their way in or rebuild something similar badly. Whether any of them publish the false positive rate, which is the number that decides whether a security team keeps the tool switched on, is a different question.
Closing: What I Would Do First
Before any of this becomes a purchase decision, do the part that costs nothing. List every asset you own that answers from the public internet, and next to each one write the date it was last tested by anyone, human or machine. If you have never done that exercise, you will find things. A staging subdomain still resolving to a box you rebuilt in 2023. Port 2087 open on a WHM server because one client needed it from a hotel wifi once. An old API token still baked into a mobile app your contractor shipped two years ago. I have found all three of those on infrastructure I was responsible for, and none of them appeared in a pentest report, because none of them were in scope of the engagement.
Then pick one system that will not page you at 2 a.m. if it falls over, and give a testing tool read-only or tightly scoped access to it for a week. The findings are not the point of the pilot. The noise ratio is. If you get forty alerts and thirty-eight of them are a stale version banner on a package you patched in March, you have your answer about whether your team keeps the tool switched on once the trial ends. A swarm that nobody reads is worse than no swarm, because it puts a subscription line on the invoice and a false sense of coverage in someone’s status report.
Write your permission rules before a vendor asks for credentials. Three lines in a shared doc is enough to start. What the agent may touch. What it may never touch, and I would put the billing database, the mail queue, and anything holding customer data on that list by name rather than relying on a scope document nobody reads. Who can stop it, with the actual name of a person and a second name for when the first is on leave. Give every credential an expiry so a forgotten agent token dies on its own instead of sitting in a vault for two years.
Start with the never-touch list. Write it this week, even if you have no intention of buying anything, because the vendor conversation goes a lot better when you already know the answer.