On October 1, 2026, the Allen Institute for AI (Ai2) released Olmo-core 3, a redesigned open training framework for mixture-of-experts (MoE) language models. The headline claim is hard to overstate: Ai2 says it has benchmarked the stack at 1.2 trillion total parameters across 512 NVIDIA B300 GPUs, and that a 47-billion-parameter MoE now trains at roughly 2.7x the throughput of Olmo-core 2 on identical hardware. More important than any number is what got released — not weights, but the training infrastructure underneath them, under Apache 2.0. For the first time, a fully open lab has published the systems design that makes trillion-parameter sparse models trainable at reasonable cost.
The Core Change: From Fully Sharded to Resident Experts
Every serious MoE training stack has the same structural problem. A mixture-of-experts model routes each token to a small subset of a much larger pool of specialized feed-forward blocks. Inference gets cheaper — only a few experts fire per token — but training still has to keep every expert’s weights, gradients, and optimizer state somewhere in memory, and it has to move each token’s activations to whichever GPU happens to hold the selected expert.
Olmo-core 2 handled the sharding side with fully sharded data parallelism (FSDP): expert weights were split across GPUs and re-gathered for every small training batch. That gather-and-reshare cycle is where MoE efficiency quietly died as models grew — the communication overhead grew with total parameter count even though compute per token stayed flat.
Olmo-core 3 flips the strategy. It switches to distributed data parallelism (DDP), keeping experts resident on their assigned GPUs and routing the relevant token data to them instead of repeatedly gathering the weights for each batch. The expensive operation stops being “move the model to the data” and becomes “move a small batch of data to the model.” That single inversion is where most of the 2.7x comes from.
The Memory Ladder: How a Trillion-Parameter MoE Fits at All
Resident experts are only half the story — no single GPU holds a trillion-parameter model even in pieces. Ai2 composes three partitioning schemes, and their interactive walkthrough shows how they stack from one GPU to 512:
- Expert parallelism assigns different experts to different GPUs, so each device stores only part of the expert pool rather than a full replica.
- Pipeline parallelism splits successive model layers across GPU groups, cutting how much of the dense backbone any single card must hold.
- A distributed optimizer shards optimizer state across GPUs instead of keeping a full copy of it on every device — usually the largest single memory consumer in mixed-precision training.
On top of that, precision does the rest. Olmo-core 3 supports MXFP8, an 8-bit floating-point format with separate scaling applied to small blocks of values. Ai2’s numbers put MXFP8 at roughly a 21% throughput gain over BF16 while cutting peak active memory from about 103 GiB to 95 GiB. Block-level scaling is what keeps an 8-bit format accurate enough for gradients that a plain per-tensor scale would destroy.
Routing Without the CPU Round-Trip
The second engineering frontier is routing itself. When every token needs a decision about which expert handles it, naive implementations turn the CPU into a bottleneck: routing metadata gets copied back from GPU to host memory, the host inspects it, then dispatches all-to-all communication. Olmo-core 3 attacks each leg of that path:
- GPU-resident routing keeps routing metadata on the GPUs, so the CPU can queue subsequent work without waiting for routing information to be copied back across the device boundary.
- Rowwise expert parallelism places routed token data directly into expert input buffers, skipping intermediate staging copies.
- Grouped GEMM fuses the many small per-expert matrix computations into batched kernels, because dozens of tiny matrix multiplies run far worse on a GPU than one grouped call.
The framework was explicitly built to bridge dense and sparse regimes: Ai2 describes growing the expert pool from eight experts to 128 while still selecting four per token. At roughly 3.2 billion active parameters, that takes total model capacity from about 4.6 billion to 47 billion parameters with under 5% throughput loss — capacity growth that costs almost nothing at inference time.
The Benchmark Numbers, Read Correctly
Two sets of figures ship with the release, and both need calibration:
- The 2.7x claim: on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed about 52,000 tokens per second per GPU under Olmo-core 3 versus about 19,400 under the prior FSDP-based implementation.
- The scaling point: a configuration with roughly 58.36 billion parameters active per token out of 1.2 trillion total, run across 512 B300 GPUs, reached a peak observed throughput of 858 TFLOP/s per GPU.
The caveat Ai2 states plainly: both benchmarks used random routing — an artificial stand-in for real token distribution that measures system throughput without training an actual model. The trillion-parameter headline is a systems test, not evidence that a useful trillion-parameter model has been trained. Coverage from SiliconAngle and Unite.AI flags the same point: these are Ai2’s own numbers with no independent validation yet, and the next-generation Olmo — which Ai2 says this stack will underpin — is still being built. Anyone comparing these figures against Megatron or other industrial stacks should treat 2.7x as a self-reported, methodology-adjacent result until third parties reproduce it.
Why Open Infrastructure Changes the Economics of Open Models
The Olmo-core repository is Apache-2.0, installable from PyPI as ai2-olmo-core, tagged v3.0.0 on GitHub, and ships with more than a reference implementation: official pretraining scripts for Olmo 3 7B and 32B (including midtraining and long-context extension stages), training logs from the actual runs behind those models, checkpoint conversion tools between OlmoCore and Hugging Face formats, and a native chat interface that avoids format round-trips. The accompanying technical report, “Supercharging Olmo-core for Efficient and Scalable MoE Training” (October 2026, Ai2 and University of Washington authors), lists measured operating points from a 12.9-billion-parameter configuration on 16 GPUs up through the 512-GPU trillion-parameter layout.
The phrase Ai2 uses in its own announcement thread captures the philosophy: “weights free, GPUs not.” Model weights being open was already half the argument for open development; publishing the training stack extends that transparency one layer down. Researchers get the implementation choices, benchmarks, and failed experiments — including which routing and parallelism decisions were tried and rejected — rather than a weight dump nobody can reproduce affordably. For anyone with modest hardware, this is where practical MoE fine-tuning research becomes viable: the framework itself is not exotic tooling locked to one lab’s cluster.
Gear for Local MoE and Sparse-Model Work
The compute side of these tradeoffs shows up at consumer scale too. A resident-experts design is memory-bound, and the local analog of “keep experts on the GPU” is keeping the active expert set inside VRAM. That is why high-VRAM cards dominate sparse-model inference:
- The ASUS ROG Astral RTX 5090 OC carries the same 32 GB of GDDR7 as every other 5090 variant and is the card most local-LLM builders standardize on when a mid-size MoE’s total footprint exceeds what smaller VRAM budgets allow — list prices currently run from roughly $3,800 to over $7,500 depending on stock.
- If you would rather buy the whole rig than upgrade piecemeal, the Corsair Vengeance i7500 prebuilt pairs that same RTX 5090 with an i9-14900KF, 64 GB of DDR5-6000 and 2 TB of NVMe storage — the system-memory headroom matters when tokenizers, data loaders and optimizer state sit beside the model.
- Sparse-model experiments are CPU-cheap and memory-hungry, so a processor like the AMD Ryzen 9 9900X (recently discounted from $499 to about $337) covers the data-pipeline side without dominating the build budget.
- Checkpoint sprawl is the practical tax of MoE work — expert pools multiply checkpoint size — and a Samsung 9100 PRO 1 TB NVMe drive (about $249 in current deals, down from $339) keeps conversion and storage stages off the network filesystem.
What This Means
Olmo-core 3 is not a model release; it is the missing rung of the open-source ladder. The 2.7x throughput figure is Ai2’s own, measured under random routing, and deserves independent reproduction before anyone treats it as settled. What is already solid: an Apache-2.0 stack that keeps experts GPU-resident, supports MXFP8 end-to-end, ships with the actual training scripts and logs behind Olmo 3, and demonstrably scales capacity (4.6B to 47B total parameters at four active experts per token) without collapsing throughput. If Ai2’s next flagship Olmo is trained on this infrastructure — which they say it is — then trillion-parameter-class sparse models move from “closed labs only” to a documented, reproducible, open recipe.

