Geeknewz exclusive timeline and hardware math, built from Aleph Alpha's Kolibri launch post, the Kolibri product page, and independent reporting from RuntimeWire. No invented scoops, quotes, or benchmark numbers.
Aleph Alpha picked German Reunification Day to ship Kolibri, an English-German mixture-of-experts model with open weights under Apache 2.0. The marketing word of the week is sovereignty. The quieter story in the company's own tables is how fast the team moved from an internal Origin checkpoint to a public release, and what you still need hanging in a rack before any of that legal framing matters.

If you care about European open weights, the Oct 3 drop is real. If you care about running those weights without sending documents to someone else's API, the GPU row on the model card is the part that decides whether this is a pilot or a poster.
Three months from Origin to Kolibri
| Milestone | Kolibri Origin | Kolibri |
|---|---|---|
| Pre-training finished | June 11, 2026 | September 11, 2026 |
| Public release | No public release | October 3, 2026 |
| Total parameters | 30.6B | 78.1B |
| Active parameters per token | 3.27B | 3.46B |
| Pre-training tokens | 7.51T | 20T |
| Longest trained context | 65,536 (64k) | 262,144 (256k) |
| Reasoning modes | One mode | None, low, medium, high |
| Experts (total / active) | 128 / 8 | 384 / 6 |
| Knowledge cutoff | EN Sept 1, 2024; DE Aug 1, 2025 | EN/DE June 18, 2026 |
Aleph Alpha's launch post is blunt about the calendar. Origin finished pre-training on June 11. Kolibri finished on September 11. That is three months between those two dates, and the company says it used that window to jump from 30B total parameters to 78B, from a 65k longest trained length to 256k, and from 7.5T pre-training tokens to 20T. The public ship date is October 3, about three weeks after pre-training closed.
Our math on the scale jump, using only those company numbers: 78.1 ÷ 30.6 ≈ 2.55× total parameters, while active parameters barely move (3.46 ÷ 3.27 ≈ 1.06×). Pre-training tokens go 20 ÷ 7.51 ≈ 2.66×. That is the MoE pitch in one sentence. You stock a bigger expert library, keep roughly the same compute per token, and pay for memory and serving complexity somewhere else.
The company also says German made up about 21.3 percent of the 20T pre-training mix, or roughly 4.3T German tokens, and that the bilingual 128k tokenizer is meant to compress German compounds more cleanly than English-first vocabularies. Those claims come from Aleph Alpha's own write-up and tech report. Treat them as the vendor's evidence package, not as independent lab results.
What "sovereign" still costs in GPUs
The product page lists a model memory footprint of about 78 GB for FP8 weights. Minimum hardware is two A100 80 GB GPUs, two H100 SXM5, one H200, one B200, or one B300. Recommended serving starts at two H100s or better. Context is advertised up to 1,048,576 tokens, with 262,144 recommended for efficient complex work.
That is the catch behind the sovereignty slogan. Open weights under Apache 2.0 mean you can download Kolibri-1 from Hugging Face and serve it with Aleph Alpha's vLLM plugin. They do not mean a laptop demo. Two A100 80 GB cards is a data-center commitment before you count KV cache, concurrency, or the long-context overrides the launch post mentions for million-token serving.
RuntimeWire's same-day write-up makes the same point from the outside: active parameters measure compute per token, not whether the full 78B model fits on ordinary office hardware. The launch materials also note that some training-data prep used other models (Gemma, Mistral-Nemo, Qwen) for rephrasing and filtering. Aleph Alpha still frames sovereignty as jurisdiction, pipeline control, and local deployment freedom, not as a supply chain untouched by foreign tools.
Serving instructions in the launch post are concrete. Install aleph-alpha-inference, then run Kolibri-1 through vLLM with the Kolibri reasoning and tool-call parsers enabled. For contexts beyond 262,144 tokens, Aleph Alpha documents extra flags to raise max model length toward the million-token ceiling. Recommended sampling is temperature 1.0, top_p 0.97, and top_k 128. That is enough for an infra team to start a bake-off; it is not a one-click consumer app.
Geeknewz verdict
Geeknewz's view: Kolibri is interesting if you need German-English work in a regulated stack and you already budget for multi-GPU inference. The Origin-to-Kolibri timeline is the clearest proof that Aleph Alpha can iterate a training factory, not just publish a one-off checkpoint. Do not confuse Apache 2.0 weights with cheap local AI. If you cannot provision at least the two-A100 floor, you are reading a sovereignty brochure, not a deployment plan. Watch whether independent German-language evals match the company's tables, and whether customers actually publish on-prem serving stories with the million-token path enabled.
