Altar-1 Security Model: Aikido’s Open-Weight Solution

Discover how Aikido Security built the Altar-1 security model for autonomous pentesting in air-gapped environments without compromising data residency.

A security engineer in a dark office reviewing code on a laptop with a server rack visible in the background

Why Aikido built a security model instead of renting one

Altar-1 is an open-weight security model that Aikido Security built by compressing Z.AI’s GLM-5.3, and the reason it exists is narrower than the phrase “open-weight” usually suggests. Aikido sells an autonomous pentesting appliance for on-prem and air-gapped networks, and it kept hitting the same wall: a frontier model runs on someone else’s hardware, so using one means sending source code, architecture docs, and unremediated findings off the network. Banks working under data-residency mandates cannot do that. Neither can OT operators whose plant floor has no internet route at all. If you have ever had to tell a client that their config file is not allowed to leave the building, you already know the shape of the problem.

That is the residency half. The harder half is that open weights on their own do not make a model deployable, and two mechanical problems sit in the way.

The first is that mixture-of-experts checkpoints have to store every expert, even for a workload that touches only a handful of them. GLM-5.3 routes each token to 8 of 256 experts per layer, but all 256 have to exist on disk and in memory. The second problem is specific to security agents: they build long-running context, and that KV cache competes with the model weights for the same GPU memory. On a 4x H200 node you get roughly 564 GB of HBM. Altar-1’s 328 GB of weights leaves real room for a 128k-context cache at production batch sizes. Put the same model on a 4x H100 80 GB node and you have 320 GB total, which is less than the weights alone. The deployment either fits or it does not, and there is no clever kernel that rescues you.

One thing worth reading before anyone builds a product on this. Altar-1 inherits the GLM-5.3 License. Commercial use, modification, and redistribution are permitted, which is more permissive than a lot of people assume when they hear “open.” The clause to notice is that model-as-a-service operators booking more than $10B in revenue over twelve months must first pass a Z.AI security review. That is not a problem for the vast majority of teams, including mine. But Altar-1 is open-weight, not OSI-approved open source, and those two labels get used interchangeably in release posts when they should not be.

I have watched hosting customers assume “open” means “no strings” and then discover a licensing clause eighteen months into a product. Read it now, while it costs you ten minutes instead of a rewrite.

Background: how GLM-5.3 becomes a 328 GB checkpoint

GLM-5.3 is a 753 billion parameter mixture-of-experts model. Each token routes to 8 of 256 experts per layer, which works out to roughly 40B active parameters per forward pass. Aikido did not retrain anything. It took Z.AI’s model through two compression steps and published the result, and both steps are described in the release write-up (opens in new tab).

Step one is quantization. The starting point is the cyankiwi GLM-5.3-AWQ-INT4 checkpoint, an existing community quantization rather than something Aikido produced. AWQ stores the routed expert weights in 4 bits with 16-bit activations, the W4A16 layout most inference stacks already know how to run. Attention layers, the shared expert, the dense layers, and the output head stay in BF16. That split matters. The parts of the network that touch every token keep their precision, and only the parts that get selected sparsely are squeezed.

Then step two, pruning. Aikido used Cerebras REAP, short for Router-weighted Expert Activation Pruning, which scores each expert on router weight and output magnitude rather than on how often the router picks it. Selection frequency is the obvious metric and also the wrong one, because a specialist expert for a rare language or a structured output format may fire rarely and still matter a great deal. Altar-1 keeps 168 of 256 routed experts per layer and drops 88, or 34.4%. Routing itself is unchanged. The router still picks 8 experts per token, now from a smaller pool, with the same approximate 40B active parameters.

Calibration traces came from Aikido’s own pentesting use, plus coding, tool calling, reasoning, and multilingual Wikipedia text. The company states no customer data was used. Each expert is scored by its largest share of any single domain’s routed work, and that is the mechanism protecting the specialists rather than raw traffic counts.

The storage progression is where the compression becomes concrete:

Checkpoint Stored weights
GLM-5.3, BF16 1,506.7 GB
GLM-5.3, AWQ INT4 488.2 GB
Altar-1, pruned W4A16 328.0 GB

328 GB is 78.2% below the BF16 parent and 32.8% below the AWQ checkpoint it was pruned from. If you have ever deleted preinstalled themes from a WordPress install to reclaim disk space, the shape of the idea is familiar, except nothing here reinstalls itself and no amount of retraining fixes a prune that cut the wrong expert.

What’s happening now: the numbers Aikido published

Aikido ran the pruned checkpoint against its internal CVE benchmark: 32 known vulnerabilities spread across 30 repositories, three runs per case. The pipeline uses other models for the surrounding stages, so this measures whether Altar-1 rediscovers a specific bug it is pointed at, not whether it finds anything interesting on its own.

Model Avg recall per run Found at least once
GLM-5.3, BF16 65.6% 25 of 32
GLM-5.3, AWQ INT4 61.5% 23 of 32
Altar-1 60.4% 23 of 32

Read the middle row before you read the bottom one. Quantization to INT4 already cost 4.1 points of recall and two vulnerabilities of coverage. Pruning added another 1.1 points of recall loss and nothing else. So the second compression step is nearly free next to the first, which is not the result I expected going in. Aikido’s own framing is that pruning kept 23 of the parent’s 25 covered vulnerabilities at 5.2 points lower recall, or 92%.

Fidelity is reported separately as KL divergence of 0.506 nats against full BF16 on a sealed 25-prompt panel. An EXL3 build of the same cut scores 0.511, which is close enough that the two quantization paths look interchangeable on this measure. I will say plainly that KL divergence is not a metric most operators can act on. It tells you the output distribution drifted; it does not tell you whether the drift lands on the code paths that matter to you.

What the benchmark does not measure is the part I would keep in mind before quoting any of these numbers in a procurement document. Blind discovery is untested. Exploit validation is untested. Fix proposals are untested. Those are three different jobs that a pentesting agent does, and only the narrowest of them has numbers behind it. The full write-up is in the release coverage at Marktechpost (opens in new tab).

Aikido also reports that Altar-1 surfaced a valid critical-severity vulnerability during a client’s production pentest. That is a single result, self-reported by the vendor, with no way for an outside party to check the finding or the conditions around it. It is worth noting and not worth weighting. I have run enough scans to know that one good catch in production tells you almost nothing about the rate at which the next ten will land, which is exactly the number a buyer needs and the one nobody has published.