All posts
Production Published 14 min

Aleph Alpha Kolibri-1: EU open-weight routing for German-English agents and on-prem RAG

On October 3, 2026 Aleph Alpha released Kolibri-1 under Apache 2.0: a 78B MoE with 3.46B active parameters, bilingual DE/EN specialization, reasoning, and tool calling. This guide explains how enterprises should place it on a routing board versus closed frontier coding models.

Jigar JoshiJigar JoshiAgentic AI Architect and Consultant
In this post (7 sections)

Introduction

Most October routing conversations are about cheaper closed frontier SKUs. Kolibri-1 is a different decision: whether a regulated German or English workload can leave the building at all. Aleph Alpha positions the model for public administration, industrials, and aerospace, with documented training choices and Apache 2.0 weights rather than a hosted-only API.

Primary sources: Kolibri has landed, the Kolibri product page, and the Hugging Face model card. Adjacent routing: GPT-6.1 Sol, Opus 5.5, Grok 4.7 and stop paying frontier prices.

What shipped

  • Release date October 3, 2026 (German Unity Day). Press follow-up October 5 confirmed general availability.
  • 78,103,074,560 total parameters; 3,457,573,120 active per token; 50 MoE layers; 384 experts with 6 routed plus 1 shared.
  • Languages: German and English only. Knowledge cutoff June 18, 2026 for both.
  • Reasoning modes none / low / medium / high. Tool calling supported.
  • Native context 262,144 tokens; Aleph Alpha validated serving to 1,048,576 and recommends ≤262,144 for complex and latency-sensitive work.
  • License: Apache 2.0 on the published weights and config. Aleph Alpha states the license does not extend to unpublished training code or methods.
  • Hardware: about 78 GB FP8 weights. Minimum examples include 2× A100 80 GB or 1× H200.

What is technically significant

Kolibri is not a multilingual generalist. Aleph Alpha chose depth in two languages, a German-aware tokenizer, and sliding-window attention with full attention every fifth layer so long context stays cheaper than a dense 70B-class model. The serving story is vLLM plus a first-party plugin (`aleph-alpha-inference`) with `--reasoning-parser kolibri1` and `--tool-call-parser kolibri1`. That is the difference between a downloadable checkpoint and a production agent worker.

Official Kolibri-1 card scores at high reasoning (Harbor for SWE-Bench; verify on your harness)
BenchmarkKolibri-1Use as
LiveCodeBench v685.9Coding-agent directional
SWE-Bench Verified66.4Repo-fix directional
HumanEval+92.7Short coding, not agent loops
BFCL v4 overall61.4Tool-call reliability
GPQA Diamond (EN / DE)84.3 / 81.3Hard Q&A, not planner default

What this means for developers

  • Install `aleph-alpha-inference>=1` and serve with the documented vLLM flags. Do not assume stock vLLM parsers understand Kolibri reasoning traces.
  • Pin sampling to the card defaults unless evals say otherwise: temperature 1.0, top_p 0.97, top_k 128.
  • For contexts above 262,144 tokens, Aleph Alpha documents `--max-model-len 1048576` plus an HF override. Treat that as a research path, not the production default.
  • Compare against Qwen3.8-27B and DeepSeek-V4.1-Flash on the same DE/EN RAG and tool-call suite before choosing an on-prem worker.

What this means for businesses

  • EU public-sector and industrial buyers get an Apache 2.0 DE/EN option with a published energy and compute account, which is a procurement artifact as much as a model artifact.
  • On-prem cost is GPU occupancy, not token list price. A 78 GB FP8 footprint is a cluster decision, not a SaaS line item.
  • Do not replace GPT-6.1 Sol or Opus 5.5 cloud planners because Kolibri is “open.” Those SKUs still win many English coding-agent evals and carry different cyber labels.
  • Legal should read the Apache 2.0 scope note: weights yes, unpublished training stack no.

Routing checklist

  1. 01
    Name the data-residency constraint
    If prompts cannot leave the EU or the customer VPC, Kolibri is a candidate. If they can, start with the closed-model routing board.
  2. 02
    Stand up a vLLM worker
    Use the official container or pip package. Enable reasoning and tool parsers. Log effort level per request.
  3. 03
    Run DE and EN evals separately
    German public-sector RAG, English coding agents, and bilingual tool calling will not share a winner.
  4. 04
    Cap context at 256K in production
    Keep 1M as an explicit override with latency SLOs.
  5. 05
    Keep a cloud verifier
    Route merge-critical or high-stakes English code review to Opus 5.5 or GPT-6.1 Sol until Kolibri wins that lane in-house.

Conclusion

Kolibri-1 is a serious on-prem DE/EN agent worker with honest serving limits and Apache 2.0 weights. It does not retire cloud frontier models. Teams with residency, German-language depth, or air-gapped RAG should bench it this week. Teams without those constraints should keep it on a watch list and spend eval time on GPT-6.1 Sol, Opus 5.5, and Grok 4.7 instead.

Sources: aleph-alpha.com — kolibri has landed a sovereign open weight model ; aleph-alpha.com — kolibri ; Hugging Face ; aleph-alpha.com — kolibri sovereign ai made in germany

The weekly take

Agentic AI patterns, delivered Thursdays

What I am shipping, watching, and pruning out of client stacks each week. One email. No fluff.

Shipping an agentic AI project this quarter?
Book a 30-min consult
Frequently asked

Questions readers ask about this post

Share this post
LinkedIn Facebook WhatsApp