Strata is not a model. It is an inference engine, and the distinction matters: the 125-billion-parameter Qwen3.8-Flash-Next it runs would not fit on a gaming GPU by a wide margin, yet Strata runs it on a 12 GB card. It does this by refusing to treat the graphics card as the whole system. The work is spread across the GPU, the CPU, system RAM, and the SSD at the same time, so no part of the machine sits idle while another part waits.
Most people read MoE throughput as a GPU number. Strata changes what that number means, and the effect is visible on hardware that is two generations old. This article covers how the engine works, what its published benchmarks claim, and what it actually measured on a Ryzen 9 5900X with an RTX 3090 and 64 GB of DDR4 — the machine this site’s articles are written on.
What Strata is
Strata is a C++ inference engine released under the MIT license by a developer working as Niko1221, hosted at github.com/Niko1221/Strata. It targets one model family rather than many: Qwen3.8-Flash-Next, the Qwen team’s 125-billion-parameter mixture-of-experts model. The repository has roughly 17,000 stars, 1,500 forks, and 1,424 commits, with releases arriving every few days — v0.1.38 on October 3, v0.1.39 on October 4, v0.1.40 two days later, and v0.1.40.3 on October 7.
It installs from a cloned repository with one script per platform, downloads about 70 GB of model weights from Hugging Face, and starts a local server. It speaks an OpenAI-compatible API and an Anthropic Messages API on localhost, and since PR #451 it also serves OpenAI’s Responses API, which is what makes coding agents like Codex CLI able to point at it. Nothing leaves the machine.
The architecture that makes it fit
Qwen3.8-Flash-Next is built as a team of 24,576 small experts. Each token activates about ten of them, so roughly six billion parameters participate in any given computation even though 125 billion exist. The catch is that all 125 billion must remain reachable — a conventional engine would need them resident on the GPU, which is where a 12 GB card stops.
Strata splits the model across the whole computer instead:
- The GPU keeps the few thousand experts that get routed most often.
- System RAM holds all 24,576 experts at once.
- The CPU computes the experts that did not fit in VRAM, at the same time the GPU works on the ones that did.
- The SSD holds a large n-gram lookup table, read a few rows at a time per token.
That n-gram layer is the unusual part. It is roughly 51 billion parameters, and it lets the model pay for its scale in memory rather than in compute — a lookup instead of a matrix multiply. Memory is cheap and plentiful on a desktop; GPU memory is not. Strata’s design is essentially an argument that a desktop has more usable memory than a single GPU does, and that a model designed around lookups can exploit that.
Quantization comes from ISTA-DASLab’s GSQ-RCO technique, with presets ranging from ultra-fast Q2_0 and IQ2_XS up to IQ3_S and coder-optimized variants. The engine adjusts the tier based on available resources rather than requiring you to pick one.
Speculative decoding, measured
Strata adds a small draft layer of about four billion parameters that guesses the next few tokens, then the large model checks them in a single pass. Guess, then check. The published claim is a 1.6–1.8x speedup over a straight pass.
Here is what that looks like on the machine this site runs on. A single request through the local API, 2,212 prompt tokens and 200 generated tokens, temperature 0:
| Measurement | Result |
|---|---|
| Prompt processing (end-to-end, wall clock) | 1,453 tokens/s |
| Decode (end-to-end, wall clock) | 79.6 tokens/s |
| Decode (engine-reported) | 101.8 tokens/s |
| Draft tokens generated | 138 |
| Draft tokens accepted | 103 — 74.6% acceptance |
| Hardware | Ryzen 9 5900X, RTX 3090 24 GB, 64 GB DDR4 |
The gap between the two decode figures is worth understanding. The engine reports 101.8 tokens/s for its own generation loop; the wall-clock number through the API was 79.6. The difference is request overhead, tokenization, and network — which is the part an agent loop actually experiences. If you are benchmarking for a chat UI, the engine number is fine. If you are benchmarking for an agent that makes hundreds of calls, the wall-clock number is the honest one.
The 74.6% draft acceptance is the number that explains the speedup. Speculative decoding only helps when the guesses are right. At roughly three in four, the guess-then-check loop is doing what the project claims. On a longer prompt with a warm cache, acceptance typically rises, because the draft layer sees more context.
Why a 3090 keeps up with a 5070
Strata’s own published figures put an RTX 5070 at 94.6 tokens/s on a 4K context and 65.1 tokens/s at 128K with the smallest quantization. The measured RTX 3090 result above sits in the same band, on a card two generations older and on DDR4 rather than DDR5. That is not a coincidence, and it is the most useful thing this engine tells you about hardware choices.
Two effects cancel out:
- More VRAM means fewer PCIe trips. A 24 GB card holds more hot experts resident than a 12 GB card, so a larger share of tokens never need to leave the GPU. That is a genuine advantage for the 3090.
- DDR4 is slower than DDR5. The RAM side carries all 24,576 experts, so memory bandwidth is a real constraint. An AM5 platform with DDR5 has more bandwidth available to the offload path.
Net result: comparable throughput. The bottleneck in an engine like this is not raw FLOPs — it is routing, cache residency, and how much traffic crosses PCIe. A card with more memory and a slightly slower platform can beat a newer card with less memory on the same workload. That is a different buying decision than the one most upgrade guides assume, and it is worth reading alongside the hardware guide before spending anything.
Prefill is where the split design shows most clearly. Strata reads long inputs in chunks up to 8,192 tokens at over 1,000 tokens per second, and as of October 7 an opt-in flag (STRATA_PREFILL_CPU_SHARE=auto) hands a small chunk’s least-routed experts to the otherwise idle CPU pool during prompt processing. For agent workloads — where a tool result of a thousand tokens arrives and the CPU pool is sitting idle while the GPU streams — that is the difference between a fast model and a fast system.
What to hold against it
- Every published benchmark is the project’s own. There is no independent reproduction posted yet. Treat 94.6 tokens/s as a claim until someone else measures it.
- The model is an explicit preview. Qwen describes Flash-Next as an under-trained preview, released on August 26 so the community could test a new architecture ahead of the full Qwen 4 family. Quality is not the point of this model; the architecture is.
- Output is not always byte-identical. v0.1.39 advertised identical output to v0.1.38, but the project itself notes that bits genuinely change on long prompts because experts pass through a different cached mix. Reproducibility across versions is not guaranteed.
- Hardware support is compiled and unit-tested, not owned. The project states plainly that it compiles and tests paths for hardware it does not have. Intel Arc and Strix Halo support arrived in v0.1.40 and v0.1.40.2 — real, but validated by users rather than by the author’s own machines.
- It is a single-model engine. If you need to run several model families, a general engine like llama.cpp is the better base. Strata trades breadth for a very specific kind of depth.
What this means
Strata is the clearest evidence yet that the “you need a data center for a 125B model” framing was a constraint of how engines allocate memory, not of the architecture itself. A model that activates ten experts per token and pays for scale through lookups can live on a desktop, because a desktop has 64 GB of RAM and an SSD that a single GPU does not.
For anyone running agents locally, the practical takeaway is that the GPU is no longer the only thing worth budgeting for. RAM capacity and an NVMe drive are part of the inference path, not storage. A 12 GB card with 64 GB of RAM and a fast SSD is a working 125B system. On this machine, with hardware that is several years old, it produced 1,453 tokens per second on prompt processing and roughly 100 on generation — enough that the difference between local and cloud stops being a speed argument and becomes a privacy and cost argument.
Hardware that actually runs this
If you are building a box for an engine like this, the memory and storage matter as much as the card. The minimum viable card is a 12 GB GPU — a PNY RTX 5070 12 GB is the cheapest current card in the tier Strata benchmarks, around $570. The component that actually decides whether 24,576 experts fit is system RAM, and it has to be the right form factor: a desktop kit, not laptop modules. A Corsair Vengeance LPX 64 GB (2x32GB) DDR4-3200 CL16 desktop kit is the configuration this article’s measurements came from, and it costs less than most people spend on a GPU upgrade while doing more for this workload.
The n-gram lookup table lives on the SSD and is read a few rows per token, so drive speed is part of the inference path rather than a convenience. A WD_BLACK SN850X 2 TB Gen4 NVMe, around $170, covers the roughly 70 GB of weights plus headroom. Plan for 80 GB of storage and 32 GB of RAM as the floor, and expect the first boot to take several minutes while the engine allocates up to 55 GB of RAM.


Pingback: 800V DC and the Cooling Wall: How 120 kW Racks Rewrote Data Center Power - GenX AI Tools
Pingback: HBM Is Eating Your RAM: What a 3x DRAM Price Means for Local LLM Builds - GenX AI Tools