Kolibri bird perched on glowing AI chip with data streams

Aleph Alpha Kolibri: 78B Open-Weight MoE with 1M Token Context Built in Europe

Aleph Alpha Kolibri: 78B Open-Weight MoE with 1M Token Context Built in Europe

On October 3, 2026, Heidelberg-based Aleph Alpha released Kolibri-1, a 78.1-billion-parameter mixture-of-experts model with 3.46 billion active parameters per token, validated context up to 1,048,576 tokens, and full weights under Apache 2.0. The release is positioned as a sovereign German-English reasoning model for regulated work, trained in Germany and Finland, and intended to run on customer-controlled infrastructure rather than third-party APIs.

What Kolibri Is and Why It Ships Now

Kolibri targets organizations that cannot send internal documents to external inference services: public authorities, banks, manufacturers, and aerospace suppliers. The model card lists German as 21.3 percent of pre-training tokens with a bilingual 128,000-entry vocabulary designed to preserve German compounds. It ships with four reasoning effort settings — none, low, medium, high — and tool calling support.

The headline is the context window. Native training was 262,144 tokens, with continued training at 65,536 and initial training at 16,384. Aleph Alpha reports quality and serving tests at 1,048,576 tokens without YaRN or RoPE scaling tricks. Positional encoding is limited to sliding-window attention layers, which allows the extension without the typical extrapolation artifacts.

Architecture: 78B Total, 3.46B Active

Kolibri is a MoE transformer with 50 layers and 384 experts per layer, 6 experts active per token. Total parameters are 78,103,074,560. Only about 3.46 billion are active per forward pass, which keeps inference compute bounded while keeping the full model resident in memory.

Weights are released in FP8 with selected components in bfloat16. The model card states that all 78 billion parameters must be held in memory even though only a fraction is active. That implies roughly 78 GB of weight memory in FP8 plus overhead for KV cache at long context. Aleph Alpha says a single B200 or H200 can serve the model, and the company emphasizes on-premises deployment for data residency.

Training and Data Sovereignty

Aleph Alpha trained the model in Germany and Finland with curated data. The company controls data curation, training, evaluation, weights, and deployment, which it markets as compliance-ready for European sovereignty requirements. The 189-page technical report covers architecture, training methods, and evaluation.

Performance Claims

Aleph Alpha publishes internal comparisons. On AIME 2025 the company reports 96.9 for Kolibri versus 84.6 for Qwen3.6-35B and 79.8 for Mistral Small 4. On LiveCodeBench v6 the score is 85.9 versus 82.5 and 71.2. The model tops similar-size MoE models in Aleph Alpha’s evaluation while trailing dense Qwen3.8 27B in some German benchmarks, which the company attributes to the MoE activation budget.

The tokenizer and long context are tuned for retrieval over long German documents: laws, contracts, manuals. The model is designed to abstain rather than hallucinate, which Aleph Alpha frames as a safety feature for regulated use.

Running Kolibri Locally: Hardware Reality

Because the full 78B weights must be resident, local deployment is not a consumer GPU exercise. A workstation with 80GB+ of VRAM or a large CPU RAM pool with offloading is required. For prototyping, the 3.46B active compute fits on a high-end GPU, but the full model needs the memory pool.

For local LLM workstations, the practical entry point remains 24-32GB VRAM GPUs with fast system RAM and NVMe storage. Those machines handle 27B-class models comfortably and can evaluate the 9B variants of similar families. The same hardware class is used for testing MoE models with partial offloading.

What to Watch For

Quantization shifts calibration. The FP8 release is tuned for inference on Hopper-class hardware. Aggressive further quantization can change probability distributions and abstention behavior. If you run locally, validate reasoning effort settings on your own German-English document sets before production use. Also note that context extension beyond 262K is validated but not natively trained; long prompts can increase KV cache pressure linearly.

Why a German-English MoE Matters

Most open-weight models optimize for English. Kolibri is explicitly bilingual with German prioritization. The 1M token window is aimed at RAG over entire codebases or legal corpora without chunking. Combined with Apache 2.0 licensing, it allows European enterprises to keep prompts and weights in-house.

The release coincides with a broader push for sovereign AI in Europe. Unlike API-only models, Kolibri can be audited, air-gapped, and fine-tuned on private data without external telemetry.

Gear for Local LLM and Sovereign Workstations

If you want a high-end workstation for testing 27B to 78B class models locally, the hardware that handles Qwen3.8-27B comfortably is the same class that supports partial MoE offloading. A 32GB VRAM GPU gives headroom for full-precision weights and adapters.

Kolibri’s 78B weights do not fit on a single consumer card, so the value of a 32GB card here is bandwidth rather than capacity: it is enough to hold the 3.46B active experts plus a long context window without offloading on every token. The ASUS ROG Astral RTX 5090 White OC Edition 32GB GDDR7 is the card that gets a 1M-token context resident while the weights stay in system RAM.

For a turnkey build sized for a resident MoE, the memory pool matters more than the card. The Corsair Vengeance i5200 RTX 5090 with Intel Core Ultra 9 285K, 64GB DDR5 and 2+2TB NVMe is the one prebuilt in this class that already ships with enough DDR5 to hold 78B weights in RAM, which is why it is worth naming despite the price.

The AMD side of the same market is the HP OMEN MAX 45L RTX 5090 with Ryzen 9 9900X3D, 32GB DDR5 and 2TB SSD. Its 128MB L3 cache is genuinely useful for an offloaded MoE, where routing decisions and KV cache management run on the CPU. The 32GB of VRAM is not enough for Kolibri’s full weights on its own; it is enough for the 9B variants of comparable families.

What This Means

Aleph Alpha’s Kolibri adds a European-built, Apache 2.0 licensed, 78B MoE with 1M token context to the open-weight landscape. It is not a chat model replacement for consumer use; it is a tool for sovereign, long-context German-English reasoning with explicit control over data. The hardware requirements are substantial, but the weights are free and the licensing is permissive. For organizations that need German language reasoning and long document retrieval on-prem, Kolibri is now the reference implementation.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *