Reflection Beam open-weight MoE model: 501B guide

Reflection Beam open-weight MoE model: 501B total, 23B active params, Apache 2.0 weights, reasoning effort, and self-hosting notes.

Two engineers in a rented data center aisle crouch beside an open server rack, one holding a closed laptop and the other pointing at a fiber cable, with warm overhead lights reflecting on the tiled floor.

What Reflection AI Actually Shipped

Reflection AI has published the architecture and training numbers for Beam, its first model. The announcement from Marktechpost (opens in new tab) describes it as a sparse Mixture-of-Experts model with 501B total parameters and 23B active per token, trained from scratch for coding, reasoning and agentic workloads. So the Reflection Beam open-weight MoE model is not a fine-tune of somebody else’s checkpoint, and it is not a dense model with a big parameter count bolted on for the press release. It is a router that picks a small slice of experts per token and leaves the rest of the weights cold.

The active number is the one that decides your cloud bill. A dense 501B model would touch all 501B parameters on every token it generates. Beam touches 23B. That is a rough twenty-fold difference in the memory bandwidth and compute you pay for per token, and it is why the claim of 3 to 4x less inference compute than larger open models on reasoning benchmarks is plausible rather than marketing arithmetic. I have run enough vLLM and TGI deployments on rented GPUs to know that throughput scales closer to active parameters than to total ones, even though all the weights still have to fit in VRAM. You need the whole 501B resident. You only calculate on 23B of it.

There is one caveat that shapes everything below. You cannot download Beam today. It is in final red-teaming, early access runs through a waitlist on the Reflection platform, and the Apache 2.0 license plus the weight files are promised for later in October 2026. I have not benchmarked this model on my own boxes, so every figure in this post comes from Reflection’s own table or from the third-party sources they credit. Treat the efficiency claim as a claim until someone runs the use independently.

What I find interesting about the release is where Reflection chose to compete. They are not arguing that Beam is the most capable open model, because they will lose that argument to Kimi K3 and say so themselves. They are arguing that you can get most of the capability at a fraction of the serving cost per task, which is the argument that actually matters once you are paying for an agent that loops for twenty minutes on a refactor. A site owner comparing hosted coding assistants sees the difference as either a monthly invoice they can absorb or one they cannot.

Background: how the Reflection Beam open-weight MoE model was built

Reflection trained Beam from scratch rather than fine-tuning an existing base, which is unusual for a first open-weight release and explains most of the numbers in the announcement. The pretraining corpus ran to 23.8 trillion tokens drawn from the web, public datasets and licensed proprietary sources. The curation funnel is the part worth a closer look. As reported in the Marktechpost write-up of the release (opens in new tab), Reflection removed roughly 95% of raw internet tokens while keeping about 1.8 trillion tokens that conventional filters would have discarded. That second figure is the interesting one. Anyone can strip spam and boilerplate. Holding onto the technically dense pages that a naive filter mistakes for garbage is where a coding model picks up its edge, and Reflection is claiming they did that on purpose rather than by accident.

The architecture interleaves local and global attention and uses fine-grained routed experts. Expert routing works a bit like a support inbox that sends each ticket to a specialist instead of making every agent read every ticket. You keep a very large staff on the payroll, but only two or three people touch any given ticket. To keep that routing balanced, the team built on DeepSeek-V3’s auxiliary-loss-free balancing and added cosine decay to the expert-bias updates. The busiest expert finished pretraining at 1.04x average load, which is a tidy way of saying no single specialist was buried while the rest sat idle.

Holding a 52-layer model steady at that scale takes more than routing policy. Reflection used depth-based scaling, SandwichNorm, attention gating and FP32 residual accumulation to keep residual norms bounded across every layer. The run itself finished in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs, reached 92.3% goodput near the end, and needed nine semi-automatic rewinds where training had to back up and recover. Midtraining then stretched the effective context window to 1M tokens, which is what makes the agentic pitch realistic for a large repository rather than a single file.

What’s happening now with the Reflection Beam open-weight MoE model

Can you self-host Beam today? No, and it is worth being blunt about that before anyone starts sizing GPUs. The weights are in final red-teaming. Early access runs through a waitlist on the Reflection platform, and the Apache 2.0 license plus the actual download are promised for later in October 2026. So the “open-weight” label is currently a commitment rather than a link. I have watched enough of these announcements slip by a few weeks to treat the date as a plan, not a schedule, though Reflection has been specific about the license, which is the part that matters for anyone planning to ship a commercial product on top of it.

How does it score against the models it sits next to? Reflection reports 80.9 on SWE-bench Verified, against 70.7 for Nemotron 3 Ultra. On Terminal Bench v2.1 the picture is tighter and less flattering: Beam at 80.1, GLM-5.2 at 81.0, DeepSeek V4.1 Flash at 90.6 and Kimi K3 at 88.3. Beam does not win that column. It lands just behind the model it is explicitly competing with, and well behind two others.

What is the reasoning effort parameter? It is a user-facing dial that trades token spend for deliberation. At low settings Beam answers quickly and cheaply. At high settings it is allowed to think longer before committing to an answer. The practical effect is that you decide per request how much you are willing to pay for accuracy, which is closer to how you already handle support tickets internally: a password reset gets a canned reply, a broken checkout flow gets a real investigation. Teams can map that dial to task difficulty instead of paying peak reasoning cost on every call.

Two things I would hold loosely. The rival scores in that table were sourced from Artificial Analysis and DataCurve rather than run in Reflection’s own use, so the comparison is not apples to apples in the strict sense. And Reflection is upfront that Kimi K3 still leads on raw capability, which is a more honest framing than most launch posts manage. The full breakdown is in the Marktechpost coverage of the Beam announcement (opens in new tab).

What it means in practice

The number that decides whether Beam is useful to you is not 501B. It is 23B. On every token the model generates, roughly 23 billion parameters do work, which puts Beam below GLM-5.2 at about 40B active, Nemotron 3 Ultra at 55B and Kimi K3 at 104B. Picture a support desk with 501 agents on the roster and 23 of them actually on calls at any given moment. You pay for the 23, not the roster. That is the selling point, and it is a quieter one than topping a benchmark table.

The reinforcement learning side is where the real engineering sits. The run used 10.5K GB300 GPUs for four weeks and produced over 100 million rollouts at up to 256K tokens of context. Training and grading burned through roughly 1.3 billion sandboxes across nearly a million coding, agentic and STEM environments. Reflection trained with fully asynchronous policy gradients, tagging every token with the policy version that produced it, and reported stable learning even at one-day staleness, 107 weight versions behind the current policy. No plateau as RL compute scaled.

The infrastructure numbers are the ones I would want to reproduce before trusting. The system sustained 110K concurrent rollouts on average, new weights reached the inference fleet in a median of about 12 seconds, and 71 inference incidents were absorbed without pausing training. I have never run anything at that scale. On my own boxes, pushing fresh weights to a live serving fleet is a deploy script and a certain amount of hope, and a 12 second median sounds plausible for a controlled fleet but I do not know how it holds up against real customer traffic.

Licensing is where this gets practical. Beam is slated for Apache 2.0. GLM-5.2 and DeepSeek V4.1 Flash ship under MIT, NVIDIA’s Nemotron 3 Ultra under OpenMDW-1.1, and Kimi K3 under a custom license that adds attribution requirements once your product crosses certain size thresholds. For most teams building an internal coding agent, MIT and Apache 2.0 are interchangeable. The moment you embed a model in something you distribute, that distinction stops being academic and becomes a line item your lawyer reads.

My expectation is that the reasoning effort dial ends up mattering more to operating cost than the active parameter count does. A cheaper token is only cheaper if you stop asking for maximum deliberation on every request.

What to expect next

The weights and the technical report are promised for later in October 2026, and the safety evaluations are still missing entirely. Reflection trained a separate safety and alignment teacher from the pretrained checkpoint, merged it with the RL teacher through multi-teacher on-policy distillation, and applied deliberative alignment on top of that. It is a real pipeline rather than a bullet point, but a pipeline is not a result. Until the evals land in the report, the safety story describes a method and not an outcome.

What I actually care about is what happens when people run Beam outside Reflection’s use. Their benchmark table is self-reported, and they are straightforward that rival numbers came from Artificial Analysis and DataCurve rather than from in-house runs. That is a reasonable way to build a comparison, and also the exact setup where a 3 to 4x inference compute advantage can quietly shrink. My expectation, and it is only an expectation, is that the efficiency claim softens under independent measurement. The 23B active parameter count is a hard architectural fact, so the floor is real. Whether you land anywhere near that floor depends on how many tokens the model burns to get there, which is where the reasoning effort dial and the length-penalty training start to matter.

Two things I will be watching. The first is whether Apache 2.0 ships as stated. “Planned” in a spec table and a LICENSE file in a repo are different states of the world, and I have watched an open-weight promise slip a quarter before. The second is whether the length penalty shows up as real token savings on a normal workload. Teaching a model to solve tasks in fewer tokens is a clean claim inside a controlled RL loop. On your repo, with your half-finished TypeScript and your folder of files named utils2_final.js, the token count will depend more on the task than on the penalty term.

The browsing result is the one I find genuinely curious. Beam improved at browsing without a single browsing task in the RL mix. If that transfers as reliably as reported, it says something about how much agentic skill is shared structure rather than separate training. I cannot verify it from out here, and neither can anyone who has not had the weights in hand.

Next step: get in line, then test it on your own repo

Get in line now, before the weights drop. Early access is the only window where you can measure latency and token spend on your own prompts without competing for bandwidth against everyone who just pulled a 501B checkpoint. When a few thousand people hit the same download at once, you learn nothing useful about throughput, you just learn that Hugging Face is slow that afternoon. On the Reflection platform the queue is per-account and the request path belongs to the model, which is closer to the shape of what you would actually serve to customers.

Join the waitlist. It costs an email address and a form, and the early access program is the only place to see real numbers before release day.

The part worth planning now is the testing you will run the moment you have access. Pick ten to twenty real issues from your own repository, ones where you already know what the correct fix looks like. Then run the SWE-bench Verified use against Beam at each reasoning effort setting rather than leaning on the 80.9 headline:

python -m swebench.use.run_evaluation \
  --dataset_name princeton-nlp/SWE-bench_Verified \
  --predictions_path ./beam_predictions.jsonl

That gives you a solve rate. It does not give you the bill. Log prompt and completion tokens per instance, multiply by your provider rate, and divide by solved tasks. A model that solves two extra issues out of twenty while spending four times the tokens at high effort has cost you money, not saved it. This is the same error people make picking a VPS plan: they compare core counts and never compare the invoice after a traffic spike.

I would track three numbers at each effort setting. Tokens per solved task, wall clock per solved task, and how often the model loops or gives up entirely. That third one matters most for agentic work, because a run that thrashes through twenty tool calls and then fails costs you the same as one that succeeded, minus the fix.

While you wait, read the Apache 2.0 text next to the Kimi K3 license and the GLM-5.2 MIT terms. Attribution requirements and permissive terms look identical in a summary table and very different in front of a lawyer. Do that reading during the waitlist period, not the week you plan to ship.