Intro: A Number Worth Looking At Twice
GitHub published a claim that an AI security agent surfaced 24 vulnerabilities in the Android codebase. The number itself is not what I find interesting. What catches my attention is that a single tool produced that many candidate findings, because in my experience the hard part is never finding more bugs. The hard part is deciding which ones matter.
I run a small VPS hosting business on the side, and I probably receive more vulnerability scanner output in a month than I could read in a week. Discovery has never been my bottleneck. Triage has always been. Every tool I have ever used, from the free ones that ship with the OS package manager to the commercial static analysis suites, generates more findings than any engineer can work through in a reasonable timeframe. The difference between a useful tool and a noisy one is not how many things it finds. It is how many of those findings turn out to be real problems that someone actually needs to fix.
I have not read the underlying reports for this claim. I am treating the number 24 as something to examine, not as a benchmark to quote or compete with. A count of findings means very little without knowing what percentage were false positives, how long the verification took, and whether the confirmed ones were already known. GitHub’s own post on this, How we found 24 Android vulnerabilities using our open source AI security agent (opens in new tab), is worth reading for the methodology, though I would still want to see the raw data before drawing any firm conclusions.
What this post will do is look at the shape of the problem, not the headline number. I will compare the traditional static analysis approach with AI vulnerability scanning, consider where an AI security agent adds value and where it does not, and talk about what this means for the kind of work I do on my own servers. I will also be honest about the places where I am guessing, because the public information does not go as deep as I would like.
Background: Static Analysis vs AI Scanning
Traditional static analysis works the way a rules engine always has. You feed it a rule set, it walks the parse tree, it tracks tainted data from a source to a sink, and it returns a verdict. The same input produces the same output every time, and that reproducibility is the property I value most, because it lets me rerun a scan after a refactor and diff the results instead of wondering whether the tool drifted. Semgrep rules, CodeQL queries, the older lint-style checkers, all of them fit this model.
The cost is precision. A rule that says “this string came from an intent extra and reached a file path” will fire on a great deal of code that is perfectly safe. It is roughly the same failure mode as a web application firewall blocking every request that contains ../ in a query string. You catch the bad ones, you also catch the honest ones, and somebody has to open each hit and decide. On a repository I maintain, one Semgrep run returns a few hundred findings. Most are duplicates or unreachable paths.
An AI security agent changes the shape of the problem rather than shrinking it. Instead of matching a pattern, a model reads a file and asks whether the code does what it claims to do. It can follow a value across function boundaries, notice that a permission check happens in one caller but not its sibling, and flag code that violates an assumption the author clearly held. That is a class of bug a rules engine has no clean way to express, because the rule would have to encode your application’s specific invariants rather than a generic pattern.
The downside is symmetric. A model can describe a vulnerability that does not exist, in fluent and confident prose, and you will not know it is wrong until you have read the code yourself. Detection gets cheaper. The cost of confirming that a finding is real, reachable, and worth fixing does not move at all.
Android makes both approaches harder. Binder IPC crosses process boundaries that a single-file analysis cannot follow. Permissions get checked at several layers, sometimes by the framework and sometimes by the app itself. Intent handling moves data across component boundaries with implicit contracts. And a large share of the interesting logic sits in C++ behind JNI, called from Kotlin, so any tool that parses only Java or Kotlin is looking at part of the picture.
That last point matters for how I read the 24-finding claim, because the mix of languages in a target determines how much of the codebase a scanner can even see.
What’s Happening Now: How the Agent Was Pointed at Android
As GitHub’s blog post (opens in new tab) describes it, the agent was run across the Android codebase and produced a set of candidate findings that engineers then reviewed by hand. That word “candidate” is carrying most of the weight, and it is the part I keep returning to.
“Found” can mean several different things. It can mean the model emitted a file path and a line number. It can mean a human read that output, opened the file, and agreed there was a real problem. It can mean the bug was reachable, reproduces in a build, and got a patch merged upstream. Those are three separate claims with three different confidence levels, and a headline figure flattens all of them into one number that reads the same regardless of which one happened. I have not read the underlying analysis, so I am treating 24 as a claim worth examining rather than a figure worth repeating.
Where this sits in a normal pipeline is the practical question. A scheduled sweep over an entire repository produces one kind of output: a long list with no urgency attached to any line of it, reviewed whenever someone has the spare afternoon. A pull request comment produces something different, because it arrives while a developer is already reading that exact code and deciding whether to merge. The same finding is far more likely to get fixed in the second case than the first, and I would expect the tolerance for false positives to be lower there too, since a noisy comment on a PR is genuinely annoying and a noisy weekly email is easy to archive.
On my own VPS hosts, I live with the scheduled-sweep version constantly, mostly malware scanning on customer WordPress installs. The output arrives, it is longer than I want to read, and it is packed with hits on minified JavaScript that turn out to be nothing. The tool did not remove work. It converted work from searching for problems into reading a list, and that is only a gain if the list stays short.
So the test I would apply to any AI pass over a target as large as Android is not the count. It is whether an engineer can clear the output in an afternoon. If clearing it takes a week, the number is a cost with good publicity.
Do AI Security Agents Replace Static Analysis?
Not for the work I do. Semgrep, CodeQL, and the linters bundled into Android’s own build tooling answer a narrow question well. Given this rule, does this code path match? They run at the same speed every time, the output is byte-identical across runs, and their silence means something you can hand to an auditor. When a CodeQL query returns nothing, I can state with confidence that the specific pattern it looks for is absent from the scope I scanned. That property, the ability to prove a negative inside a defined boundary, is the part an AI security agent cannot give me, because a model reporting nothing may just have had an off day. GitHub’s own writeup (opens in new tab) describes an agent run over the Android codebase that produced candidate findings for engineers to review, which is a different claim from a deterministic scanner going quiet.
The AI pass covers a failure mode that is genuinely real. Picture a WordPress plugin with a custom AJAX handler. WPScan will tell a site owner the plugin version carries a published CVE. It will not tell them that the handler calls current_user_can() four lines after it has already written the request payload into wp_options. Every individual line is well formed. The ordering is wrong, and a model reading the file top to bottom can catch that the same way a human reviewer would. A pattern matcher has no opinion about ordering unless someone wrote a rule specifically for it.
So I run both rather than picking a side. The catch is that running both is not free. If the agent’s findings land on top of the scanner’s findings and the review queue only grows, I have bought detection and paid for it in engineering hours, which is the worst trade in security tooling. The arrangement worth paying for is the one where the AI pass retires some deterministic output, say by confirming that a flagged string is a false positive and attaching the surrounding context so nobody has to open the file.
I do not have data on whether that netting actually happens at Android’s scale. On my own repositories the honest answer right now is that it does not, and I am clearing both lists.
What It Means in Practice
The number that matters is not 24. It is how many of those 24 a security engineer can confirm in a single sitting. If the answer is most of them, that is a good afternoon and the agent paid for itself. If the answer is that two are real and the other twenty-two each need an hour of reading Binder call sites to rule out, then the tool has moved work from detection to triage and called it progress. I have run jobs like that on my own boxes. A scanner dumps 300 lines into a log at 3am, and by the time I have dismissed the noise I have not hardened anything. Detection was never my bottleneck.
The cost accounting is where a lot of teams lose the thread. Model calls are cheap and CI minutes are cheap, and those two numbers are visible on a bill, so they get tracked. The engineering hours spent validating candidates do not appear anywhere, because those hours come out of the same pool as feature work and nobody files an invoice for them. Multiply a few hours per finding across a month of scheduled sweeps and the real line item is a part of an engineer’s week. If a vendor pitches finding counts and never mentions validation throughput, that silence is the thing I would ask about.
There are judgments the agent simply cannot make from reading a repository. Whether a bug is reachable behind a reverse proxy that rejects the request before it hits the handler. Whether a defect matters in your deployment because the affected module is not compiled into the build you ship. Whether two medium findings chain into something critical when a user can trigger them in sequence. Those answers live in your infrastructure, your build flags, and your auth layer. No static view of the source has them.
Where I would put it, based on what I actually run: a scheduled sweep over repositories where I control the whole stack, so I can make the reachability call myself, and not as a blocking check on every pull request. A gate that can be wrong in a confident tone teaches developers to click past it.
A concrete example. On a small WordPress host I run, a plugin flagged for an unescaped $_GET parameter looked alarming until I checked that the parameter feeds a string comparison and never touches the database or the page output. The report was technically accurate and operationally useless, and confirming that took longer than the scan did. Multiply that by a queue and you see the shape of the problem.
What to Expect Next
Finding counts are cheap to publish and slow to verify, so my expectation is that we get more of them, from more vendors, before we get better ways to check them. GitHub at least named the tool and published the agent itself (the writeup is here (opens in new tab)), which is more than most announcement posts bother to do.
The pressure I expect to build is on methodology. A number like 24 means very little without the shape of the pipeline behind it. What task template did the agent run with? How many candidates came back before anyone filtered them? How many engineers validated the shortlist, over how many days, and what share got patched versus closed as not worth fixing? If vendors want these counts read as results rather than marketing, somebody eventually has to publish that detail. I would not hold my breath for the first wave of announcements.
Verification is where I expect the visible bottleneck to sit for a while yet. Detection keeps getting cheaper, since a model pass over a repository costs less than the coffee the reviewing engineer is drinking while reading the output. Confirming that a finding is real, reachable from an untrusted input, and worth the patch still costs a human an afternoon. Tooling that helps me triage candidates, dedupe them against findings I already have open, and automatically close the ones that provably cannot be reached would be worth more to me than another detector with a better recall figure.
Adoption across open source will stay uneven, and I do not read that as laziness on anyone’s part. A project with a security team and a line in the budget will wire an agent into a scheduled job and actually read what comes back. A project with one maintainer and four hundred open issues will look at twenty-four fresh candidates stacked on top of that pile and close the tab. Same tool, same output, and in one case a backlog and in the other a notification nobody opens.
Where I could be wrong: if triage tooling improves faster than I expect, the validation tax drops and my preference for scheduled sweeps over blocking gates gets harder to defend. I would be happy to be wrong about that one.
What I Would Do This Week
Pick one repository you own end to end, small enough that you already know what most of the files do. On my side that is a Go service handling provisioning jobs, roughly four thousand lines, one contributor, and a CI config I can edit without asking anyone’s permission. A first pass against something that size takes an afternoon to read properly, which is the whole point.
Run the agent once over the entire repo and write the raw output to a file before you look at any single finding. The order matters more than it sounds like it should, because the moment you start reading, you start rationalising about which findings are worth keeping and which ones you would rather not deal with.
Then, one candidate at a time, add a line to a plain CSV with the file path, the line number, what the tool claims the problem is, and one of three verdicts: real, false positive, or unclear. Log the minutes you spent on each. That last column will be embarrassing on the first pass and genuinely useful on the second, because it turns “this tool is noisy” into a number you can defend in a budget conversation.
Confirmed findings go into the tracker your team already has open in a tab. On my repos that is a single gh issue create --label security call against the same repository where everything else lives, not a fresh board with its own columns and its own login. Creating a separate security queue feels tidy for about a week, and then it becomes the spreadsheet nobody has bookmarked. I have watched one of those sit untouched for eleven months while three of its entries shipped to production unchanged and two of them were reachable from an unauthenticated endpoint.
Set a reminder for thirty days out and rerun the identical scan against the same commit range. Compare the two CSVs on the verdict breakdown rather than the total finding count. If the false positive share has not moved, the tool has not earned a place in your pipeline yet, and now you have the receipts to say that out loud instead of arguing about it.
Type the header row before you run anything. It takes ten seconds and it is the part people skip.