Reflection Beam: A 501B/23B Open-Weight MoE Aims Squarely at GLM-5.2 With 3-4× Lower Token Cost
Reflection AI officially launched Beam on Oct 5 — a text-only 501B total / 23B active sparse MoE targeting GLM-5.2 with 3-4× lower token cost and inference compute. Apache 2.0 weights drop later in October. Trained on 10.5K NVIDIA GB300 GPUs.
Reflection AI shipped Beam on October 5, 2026 — its first frontier open-weight system and the most direct shot the open-weight camp has taken at China’s GLM-5.2 to date. The headline numbers: 501B total parameters, 23B active per token, sparse Mixture-of-Experts, trained on 10.5K NVIDIA GB300 GPUs over four weeks, with weights dropping later this month under Apache 2.0. The company is claiming 3-4× lower token cost and inference-time compute versus rival Western open-weight models, and Reflection is putting its benchmark scorecard behind the claim.
The numbers and the claims:
- 501B total / 23B active per token — sparse MoE; only ~5% of parameters fire on any given token. Source: Reflection’s own announcement.
- 10.5K NVIDIA GB300 GPUs over four weeks of training, with 100M+ rollouts in the high-compute RL run. NVIDIA-backed, run by former Google DeepMind researchers.
- 80.9 SWE-bench Verified, 80.1 Terminal-Bench v2.1, 78.0 SWE-bench Multilingual, 77.2 SWE-bench Pro v2-Hard — the four-benchmark coding-agent lineup.
- 97.8 AIME 2026, 90.5 GPQA Diamond — reasoning + general-knowledge scores that put Beam in the same conversation as the top closed-weight frontier models.
- 3-4× lower cost vs Western open-weight rivals — the explicit target is GLM-5.2.
- Apache 2.0 weights later this month, plus an OpenAI-compatible beta API live the same day as the announcement.
The shape of the model
501B total / 23B active per token is the same pattern DeepSeek, Mixtral, and the recent Llama-derived open-weight efforts converged on — keep the total parameter count frontier-class, keep the per-token compute modest, route via sparse expert selection. The architecture choice says “we want a model that scores well at the top of the benchmark table while staying deployable” — the 23B active footprint is roughly in the same ballpark as Llama 3 70B dense at inference, even though the total capacity is several multiples beyond it.
The 10.5K GB300 GPU training run is the other headline. GB300 is NVIDIA’s Blackwell-generation accelerator and the current ceiling on per-GPU training compute. A run on 10.5K of them for four weeks is the kind of compute bill that explains why the open-weight frontier has consolidated around players with serious funding — Reflection is NVIDIA-backed, which is the relevant context here.
The 100M+ rollouts figure is the RL signal. Most frontier training runs report single-digit-million rollouts; Beam’s RL is reporting two orders of magnitude beyond that. Whether that translates to a real reasoning gap or is just a training-budget flex is the empirical question the weights release will answer.
The benchmark scorecard
The published numbers, taken together, are the strongest part of the announcement:
- 80.9 SWE-bench Verified — this is the same benchmark the Anthropic enterprise moat post (Oct 5) anchored Sonnet 5.5’s SWE-bench Pro 81.3% claim against. Note Beam’s number is SWE-bench Verified, not Pro; SWE-bench Pro is the harder subset. Don’t conflate them.
- 80.1 Terminal-Bench v2.1 — agentic-coding in real shells. This is one of the more honest agent benchmarks; partial-mode inflated scores have been a category-wide problem, and Terminal-Bench’s strict mode is the more conservative read.
- 97.8 AIME 2026 — math. This is the AIME 2026 set, not 2025. The 2026 set is harder than the 2025 set was at this point last year, so the same numeric score today is a stronger claim than the same score would have been 12 months ago.
- 90.5 GPQA Diamond — graduate-level reasoning. Solid. Not category-leading but in the conversation.
The honest read: these are self-reported, on a harness Reflection controls, on benchmarks where the variance between vendor-published numbers and independent replications has historically been several points. The independent arena runs (lmarena, the-chatbot-arena) are what will settle whether 80.9 SWE-bench Verified is “real 80.9” or “harness-tuned 80.9.” Don’t migrate production workload on the announcement number.
The 3-4× cost claim — what’s actually being compared
The “3-4× lower token cost and inference-time compute” claim is the load-bearing one. Reflection is comparing Beam against Western open-weight rivals — the implicit target is the Llama-derived and DeepSeek-derived systems that have dominated the open-weight frontier for the last 18 months.
The comparison framing matters. Token cost in 2026 has compressed to the $2/$10 input/output reference tier for frontier closed-weight (OpenAI’s GPT-6.1 Sol, Anthropic’s Sonnet 5.5, Google’s Gemini 4 Argon in beta). Open-weight systems have been priced for self-hosting economics, where the relevant comparison is GPU-hours per million tokens rather than API pricing. If Beam’s 3-4× claim is against self-hosted Llama 3 405B or DeepSeek V4-Pro at production token throughput, that’s a real efficiency story. If it’s against Western API pricing, the math doesn’t quite work the same way — the closed-weight vendors have a pricing lever open-weight can’t match.
The China angle is also load-bearing. The fact that the announcement explicitly targets GLM-5.2 (Z.ai’s frontier model) says Reflection is positioning Beam as the Western open-weight answer to the Chinese open-weight frontier. The “open-weight race” framing is now East-vs-West inside the open-weight category, not just “open vs closed.” That’s a meaningful shift from 12 months ago.
What it means for builders
Three takeaways:
-
The open-weight frontier in 2026 is consolidating around deep-pocketed-execution players. Reflection is NVIDIA-backed. DeepSeek is state-backed. Mistral is VC-backed. The Llama “everyone can train a frontier” era is over; the 501B / 10.5K GPU / 100M-rollout-RL regime requires compute budgets that put training behind serious funding. The weights still drop under Apache 2.0 — that’s the open-weight promise preserved — but the training is increasingly concentrated.
-
Apache 2.0 + OpenAI-compatible API is the new open-weight release shape. Reflection is shipping weights + a hosted API endpoint on the same launch day. That matches the Mistral / DeepSeek / Llama 4 pattern from the last 12 months: open weights for self-hosting plus a managed endpoint for those who don’t want to operate it. If you’re an enterprise buyer evaluating “open-weight first,” this is now the baseline expectation, not a bonus.
-
Wait for the independent-arena replication before benchmarking seriously. Reflection’s announcement is a vendor scorecard on vendor benchmarks. The independent verifications matter: lmarena and the-chatbot-arena will surface the real numbers in the next 1-2 weeks, and history says the delta between vendor-published and independent-replicated is several SWE-bench points in either direction. The 80.9 SWE-bench Verified number is the one to watch most closely — it’s the benchmark with the most agentic-coding economic weight.
What we still don’t know: how the 501B/23B active ratio behaves on real production traffic at sustained throughput, what the per-GPU-hour cost looks like when self-hosting, and whether the RL-trained coding-agent capabilities hold up outside the SWE-bench distribution. The weights drop later this month answers the first two; the third needs production workload.
Sources
- Reflection AI — Introducing Beam (canonical spec sheet) — primary; 501B/23B MoE, Apache 2.0, GLM-5.2 targeting, NVIDIA-backed
- TechCrunch — Reflection debuts Beam, an open-weight AI model to rival Chinese models at lower compute cost — primary; 3-4× efficiency vs Western open-weight, GLM-5.2 framing
- Bloomberg — Nvidia-backed Reflection unveils open AI model taking on China — NVIDIA backing, China competition framing
- Unite.ai — Reflection AI unveils Beam, a 501B-parameter open-weight model — 501B/23B MoE architecture details
- AI TLDR — Reflection Beam release summary — benchmark scorecard: SWE-bench 80.9, Terminal-Bench 80.1, AIME 97.8, GPQA 90.5
- The Next Web — Reflection AI Beam open-weight model — GLM-5.2 target, Apache 2.0 release timeline
- Implicator.ai — Reflection AI Beam open-weight model GLM-5.2 — text-only MoE confirmation, 501B/23B
- Runtime Wire — Reflection AI Beam open-weight model — independent scorecard analysis
- Orca Router — Reflection Beam 501B explained — 501B/23B active details + beta endpoint
- Startup Fortune — Reflection AI launches Beam, an open model it says beats the West on efficiency — 3-4× efficiency claim source