Agent-to-Agent Handoff Security: The Audit Gap

Agent-to-agent handoff security is the AI security gap most audits miss: authority passed between agents is invisible to single-agent checks.

A woman and a man in business casual clothes stand near a reception desk in a sunlit office lobby, one handing a key to the other.

The Handoff Is the Gap

The log for a payment that should never have gone out looks fine. One agent reviewed the account and wrote status: cleared. A second agent read that status and released funds. There is no forged message, no stolen credential, and the control that should have caught it never had anything to catch. Both agents used approved tools and wrote complete audit records.

Agent-to-agent handoff security is the part of that picture most audits never reach. Identity checks, tool allowlists, scoped credentials and periodic access reviews all answer questions about one agent at a time. None of them inspect what authority travels between two agents, and that is precisely where the failure sat.

If you are running agents that plan, call tools and take action with thinner human review than you had a year ago, this is your problem now rather than next quarter. The workflows shipped faster than the review process around them.

I have watched this play out in hosting, which is where I spend most of my working hours. A support agent closes a ticket with the note “verified, billing ok.” That ticket was verified for the purposes of a refund lookup and nothing else. A provisioning agent downstream reads “verified” as confirmation that the customer’s identity has been established, and hands over control of an account. Different decision, same word, and the second agent granted itself authority the first one never issued. Every step logged, and every step green.

The same shape shows up in the cases that prompted this post. I am leaning on three anchors: that “cleared” pattern, the individual cases OpenAI has disclosed where agents found coordination paths their developers did not intend (including public file-hosting services when local sharing failed), and the AI security guidance South Korea’s KISA is developing for autonomous agents. Reporting on that KISA work, As South Korea Tightens AI Security, Agent-to-Agent Handoffs Deserve More Attention (opens in new tab), is careful to note the checklist is still in draft, so treat the regulatory side as direction rather than settled requirement.

The OpenAI disclosures deserve a careful read too. They are individual observations, not evidence of an attack, and they say nothing about frequency. What they do show is that when the expected route is blocked, an agent may find another. That is capability rather than malice, and it is exactly the behaviour that breaks a handoff assumption built on the idea that an agent stops when it cannot proceed.

Background

KISA has said its work on autonomous systems may produce a checklist for agentic AI services, sitting alongside common controls for physical AI. Both are in development, and the reporting around them is explicit that no final text exists yet. What matters at this stage is the direction: the agency is treating agents as a distinct class of system rather than as a feature inside an application, which sounds like a small distinction until you try to map an agent onto an existing control catalog and find there is no row that fits. Access controls describe what an identity may reach. Tool inventories describe what it may call. Neither describes what a conclusion is allowed to mean once it leaves the agent that formed it.

The reason conventional models do not transfer cleanly comes down to the shape of the interface. A function call returns data under a fixed contract. The caller knows the type, the fields, the range, and what an error code signifies, because someone wrote that down and the runtime enforces it. An agent returns a conclusion. It arrives as language, and the receiving agent resolves its meaning from policy, evidence, and whatever context it happens to hold. There is no schema that says “cleared means only that my own review found no exception.” The contract is implicit, and it gets rewritten on the receiving side without anyone deciding to rewrite it. A warehouse worker marking a parcel “checked” for weight is not clearing it for export, but a courier reading the same word can reasonably conclude otherwise.

The OpenAI disclosures the source cites fit that gap. These are individual cases, not evidence of an attack, and they carry no information about how often the behaviour occurs. What they show is that an agent will locate a coordination path its developers did not intend, including public file-hosting services when local sharing failed. A blocked route did not stop the workflow. It rerouted it. Any handoff design that assumes an agent halts when it cannot proceed along the expected path is resting on something the evidence does not support.

Doyle-Spare calls the working meaning an agent resolves from policy, evidence, and context its Operational Interpretation, and the useful part of that framing is that the interpretation outlives the agent that produced it. It gets copied into the next agent’s context as established fact, shapes the next decision, and then travels again. By the third hop, nobody in the chain is examining the assumption, only the word.

What’s happening now

KISA has said the work may include a checklist for agentic AI services and common controls for physical AI, as reported in this analysis of South Korea’s direction (opens in new tab). The guidance is still in development, so nobody should be building a compliance plan around a document that does not exist yet. The direction is what matters. Korean regulators are moving toward naming the multi-agent workflow as the thing that needs controlling. Most enterprise audit programs are still pointed at individual agents. The asset inventory lists the agent, its identity, its tool scope. The handoff between two agents usually appears nowhere in the control set, which means an auditor examining the deployment gets a complete picture of the parts and no picture of the seams.

Field behaviour makes that omission expensive. Agents route around blocked paths. When an expected tool call fails, or a local sharing mechanism is unavailable, the agent looks for another way to finish the task instead of stopping and asking. I do not read that as malice and neither does the source. It is what these systems are built to do, and it is exactly what breaks an assumption baked into a lot of handoff designs: that a downstream agent only receives a conclusion if the upstream path completed normally. It will not hold. The downstream agent receives conclusions produced by routes nobody modelled, and it has no way to tell the difference.

Audit coverage fails in a specific and quiet way here. The logs around a “cleared” handoff are clean. Authentic sender, approved tool, valid message format, a complete record of what was sent. Nothing in that set of facts is false, so no rule fires and no alert opens. What actually changed, the expansion of what “cleared” was taken to permit, is not a field in the log. It never gets written down, because the systems involved were never built to write it down.

The timing is awkward. Agentic guidance is being drafted at roughly the pace teams are shipping multi-agent workflows into production. I expect the first published checklist to trail what is already running in production by a wide margin, and I would rather be wrong about that than plan around it. Waiting for the framework to arrive before looking at your own handoffs means waiting a long time.

What agent-to-agent handoff security looks like in practice

The rule that makes any of this tractable is short: a conclusion arriving from another agent is input to a new decision, never permission that travels along with the message. Practice pulls the other way. A field that reads status: cleared looks like a green light to a downstream model, and a lot of system prompts say something close to “proceed if the upstream check passed.” You will find that sentence in production prompts today.

Work the “cleared” case through once and the problem shows its shape. Agent one reviews a record set against a policy and finds no exception. What it is actually saying is “my review turned up nothing.” Agent two needs a different fact, that every condition required to release the payment, open the account, or move the customer’s DNS zone has been satisfied. The string is byte-identical in both hops. The authority is not.

Take a domain transfer, which any site owner has run into. Agent one checks the ownership record against the contact email on file and reports “verified.” Agent two reads “verified” and releases the auth code to whoever asked for it. The first agent confirmed an email address matched a record. The second treated that as proof the requester was entitled to move the domain away. Nothing was forged, no credential was stolen, and a registrar’s audit trail would show two correct steps followed by an outcome nobody intended.

Before the receiving workflow acts, the source says three questions need answers. What did the upstream conclusion actually mean? Was that meaning authorized for this new decision? Which downstream actions will end up depending on it? That third question is what separates one bad handoff from a week of unwinding payments, provisioning, and access grants that all traced back to it.

This is semantic failure propagation, and it is quiet by construction. Agent one passes its review because its review was correct. Agent two passes because it acted on an input that was well-formed and came from an authorized sender using an approved tool. The workflow carries the earlier error forward at machine speed, and no single-agent test will surface it because no single agent is wrong.

Does authenticating an agent message make the handoff safe? (Q&A)

No, and the reason fits in one sentence: authentication tells you who sent the message, while authorization for the receiving decision is a question the message cannot answer on its own.

If the sender is authorized, the format is valid, and the logs are complete, what is actually broken?

The receiving agent is relying on a meaning that was never authorized for its specific decision. Sign the payload, pin the certificate, validate the schema, and you still have not established that “cleared” was cleared for releasing funds. Those controls verify that a message arrived intact from a trusted party, and they say nothing about whether the conclusion inside it was permitted to carry that much weight at this hop. A signed message can still be an unauthorized instruction.

Is this the same as prompt injection or a compromised agent?

No, and the source is careful about that distinction. Prompt injection assumes someone is manipulating the agent’s input. A compromised agent assumes the agent itself has been taken over. Semantic failure propagation needs neither. Two well-behaved agents, each following its own instructions correctly, can produce the failure on their own. That is what makes it awkward to write a detection rule for. There is no adversary to find, only a meaning that traveled further than anyone authorized.

What has to be preserved for an audit to catch it?

The origin of the upstream conclusion, and whether it was authorized at that origin. Provenance of meaning, not provenance of the message. Most audit trails I have looked at can tell you which service signed a payload and when it went out. Very few record which meanings that service was permitted to emit. So when the payment releases, the log shows a valid signature on a valid message from a valid sender, and every field an auditor is trained to check passes. The one question that would have caught it, whether this agent was allowed to treat “cleared” as permission to pay, is not a field anyone is writing down yet.

What agent-to-agent handoff security guidance is likely to change next

The first drafts of any agentic checklist tend to look like an inventory. Count the agents. Name the model behind each one. Record the tools each is allowed to call and the data each can touch. That is useful work, and it maps neatly onto how asset registers already function, which is probably why it comes first. What it does not describe is the moment between two agents, because a handoff is not an asset you can list. It is a relationship, and control catalogs are generally bad at relationships.

I expect the next round of guidance, KISA’s included when it lands, to name that relationship explicitly. The shape will probably be a control that asks a team to state, for each edge in a workflow, what authority the upstream conclusion carries and what the downstream agent is entitled to do with it. That is a narrower question than “is this agent secure,” and it is far harder to answer with a checkbox.

The phrase I would watch for is authorization at origin. Today, almost every audit trail answers who sent a conclusion and when. The requirement that follows from this problem is different and harder to build: recording what a conclusion was authorized to mean at the point it was formed, and by whom that meaning was granted. If a review agent emits “cleared,” someone in governance has to have decided in advance whether that token is allowed to authorize payment, or only to authorize the next review step. Right now, that decision is usually made by whichever engineer wrote the second agent’s prompt.

The friction here is real and worth naming before anyone promises you a quick control. Authorization-at-origin asks governance teams to reason about meaning, intent, and context. Control catalogs are built around access and identity, which are enumerable. Meaning is not. I have watched access reviews turn into a spreadsheet exercise, and there is no obvious equivalent spreadsheet for “what this word was permitted to imply.”

How well any of this holds up once a workflow chains six or seven agents deep, I genuinely do not know. The source does not claim to either. At some depth the question of origin stops being tractable, and someone will need to decide whether to cut the chain into separately authorized segments instead. I suspect the first real guidance will dodge that depth problem, and the second version will have to face it.

Close the gap on one workflow this week

Pick one multi-agent workflow that is already in production and touches money, access, or customer records. Not the one on next quarter’s roadmap, the one running right now. Then trace a single conclusion end to end and write down, in plain sentences, what each agent understood that conclusion to permit.

The mechanics of this are dull, which is why it works. Find the token itself. If your review agent emits a status like cleared, verified, or approved, grep for that string in the downstream service and open the code path that consumes it. You will usually find a match statement with no comment explaining why that value is enough to proceed. Ask the engineer who wrote it what the upstream agent was promising. The reply is often something like “it means the case is fine,” and that is the exact collapse: a narrow review outcome being read as blanket authorization.

For that one conclusion, record three things:

  • Which agent formed it, and what its own policy actually allowed it to conclude
  • What the receiving agent treated it as permission to do
  • Whether anyone in governance ever approved that second reading

Most teams can answer the first two in ten minutes. The third is usually blank. That blank is your finding, and it belongs in a ticket with an owner and a date, not in a wiki page that nobody opens again.

The other change worth making this week is smaller and cheaper. Look at your structured logs for the handoff and see what fields you capture. Sender identity, message ID, timestamp, payload, maybe a trace ID. What is almost never there is the scope of the conclusion, meaning what the origin agent was authorized to assert. Add two fields where you can: where the conclusion originated, and what meaning it was authorized to carry. Even a free-text field populated by the emitting agent is better than nothing, because it forces whoever writes the prompt to state the scope out loud. I have done this on internal tooling and the first pass usually reveals that two agents disagree about what one word means.

None of this requires buying a platform or waiting for KISA to publish. Open the ticket for the one workflow where a misread status would release money or grant access, and finish the trace before the next sprint adds a third agent to the chain.