Sparse MoE neural network with glowing nodes representing Reflection Beam 501B architecture

Reflection Beam 501B: Open-Weight MoE Announced with Apache 2.0 Release Planned

Reflection AI unveiled Beam on October 5, 2026 — a 501 billion parameter sparse Mixture-of-Experts model with 23 billion active parameters per token, aimed at coding, reasoning and agentic workloads. The announcement says Beam was pretrained on 23.8 trillion curated tokens and refined with a four-week high-compute reinforcement learning run on 10.5K NVIDIA GB300 GPUs, generating over 100 million rollouts. Weights are promised later in October under Apache 2.0, with a technical report, model card and serving stack.

What Beam Is and Why It Matters

Beam is Reflection’s first open-weight model. It is a sparse MoE: 501B total parameters, 23B active per token. The architecture combines interleaved local and global attention with fine-grained routed experts and controlled residual streams. The goal is frontier-level coding and agentic performance with significantly lower inference cost than dense models of similar capability.

Reflection positions Beam as a Western open frontier response to Z.ai GLM-5.2 and Qwen 3.8-Max. Early benchmark tables in the announcement show competitive numbers on coding and terminal tasks: SWE-Bench Verified 80.9, Terminal Bench v2.1 80.1, SWE-Bench Multilingual 78.0, AIME 2026 97.8, GPQA Diamond 90.5. The company claims scores comparable to GLM-5.2 on advanced reasoning while using 3–4× less inference compute.

Training and Scale

The pretraining corpus covers 23.8 trillion diverse, curated tokens from web and licensed datasets. Mid-training extended effective context to 1 million tokens. Reinforcement learning was the scaling axis: 10.5K GB300 GPUs for four weeks, ~100 million rollouts, up to 256K context per rollout, with ~1.3 billion sandboxes and about one million high-quality environments for coding, agentic and STEM tasks.

Reflection emphasizes stable asynchronous RL: policy gradients with up to 107 weight versions staleness, and new algorithms to reduce training-inference mismatch. A controllable length penalty rewarded correct solutions while discouraging unnecessary tokens; early RL reduced token use, later RL increased reasoning length for higher performance on demanding tasks.

Capability Highlights

The announcement focuses on coding and agentic work. Beam is competitive with larger open models on coding and agentic tasks, approaching Qwen 3.8-Max. Where Kimi K3 remains ahead on raw capability, Beam’s advantage is efficiency.

  • Agentic coding: SWE-Bench Verified 80.9, Terminal Bench v2.1 80.1, MCP Atlas 78.7
  • Reasoning: AIME 2026 97.8, GPQA Diamond 90.5, HLE no tools 36.2
  • Tool calling/search: DeepSearchQA with context management 80.1, BrowseComp 77.4

The model is text-only in the documented API; no native image/audio/file inputs. Beta context is 262,144 tokens combined input/output, with generation capped at 131,072 tokens.

Open Release Status

As of October 5, Beam is in final red-teaming. Early access is via a waitlist at platform.reflection.ai. No weights are available yet. Reflection says the weights, technical report, model card and developer artifacts will ship later this month under Apache 2.0, with quantized FP8 and NVFP4 numerics for efficient deployment and distribution partners.

Apache 2.0 allows commercial use, redistribution and modification. If the release lands as promised, developers could self-host, quantize and fine-tune Beam on their own hardware.

What This Means

Beam signals a push toward efficient, open MoE workhorses for enterprises and developers who want to run models locally. If the Apache 2.0 weights arrive in October, 501B total / 23B active is a compelling inference efficiency story for coding and agentic pipelines. Verification will require independent reproduction of the benchmark claims and real-world token-per-second per GPU measurements.

Hardware That Runs Models Like This

Running a 501B parameter MoE locally demands high-end GPUs and lots of VRAM. The following Amazon listings are currently circulating in deal round-ups and provide the kind of GPU horsepower needed for large MoE inference.

The ASUS ROG Astral RTX 5090 32GB GDDR7 White OC Edition Graphics Card offers 32GB GDDR7 on a 512-bit bus and is repeatedly cited in PC Guide deals as ideal for local AI and ML workloads.

The Skytech Gaming Legacy RTX 5090 PC with Intel Ultra 9 285K, 64GB DDR5 and 4TB SSD bundles the RTX 5090 with 64GB DDR5 and a 4TB Gen4 NVMe SSD, a configuration frequently recommended for running large local models.

The HP OMEN MAX 45L RTX 5090 Ryzen 9 9900X3D 32GB/2TB Gaming Desktop pairs the RTX 5090 with a Ryzen 9 9900X3D and 32GB DDR5, a prebuilt option for 4K gaming and workstation-class AI tasks.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *