Skip to main content
株式会社オブライト
AI2026-10-059 min read

Kolibri-1 Requirements

VRAM and GPU for Aleph Alpha's 78B MoE (about 78GB in FP8, 2026)

Kolibri-1 needs about 78GB in FP8, running on 1x H200 or 2x H100. Aleph Alpha's 78B/3.46B-active Apache 2.0 MoE: VRAM by precision, vLLM launch, German focus. Oct 2026.


Kolibri-1 requirements: about 78GB in FP8, so 1x H200 or 2x H100 and up

Kolibri-1 is an open-weight Mixture-of-Experts model (Apache 2.0) released by Germany's Aleph Alpha on October 3, 2026, with about 78B total and 3.46B active parameters per token, specialized for German and English. According to the model card, the official FP8 weights are about 78GB. The minimum setup is 2x A100 80GB, 2x H100, 1x H200, 1x B200 or 1x B300, and the recommended setup is 2x H100, 2x H200, B200 or B300. For 4-bit quantization, our estimate from parameters times bit width is roughly 39-45GB (this is not a model-card figure). Inference is light because only 3.46B parameters are active, but all experts must still sit in memory, so you need room for the full 78B. Figures as of October 2026.

What is Kolibri-1? Aleph Alpha's sovereign-AI, German-focused model

Aleph Alpha is a German AI company that has pitched sovereign deployment, meaning systems kept under the customer's own control, for regulated sectors such as public administration and industry. Press coverage places Kolibri-1 in that context: an open-weight model for self-hosting in public administration, manufacturing and aerospace. Its defining trait is a bilingual German-English design with a custom tokenizer (UniBPE, 128,000 vocabulary) tuned for German morphology such as compound words. The knowledge cutoff is June 18, 2026. It supports a reasoning mode with effort levels and Hermes-style tool calling.

In other words, this is not a general multilingual model. It is built so German-speaking companies and agencies can run it themselves without sending data out, and Japanese is not a main target. Keep that in mind when reading the hardware numbers.

Specifications

ItemDetail
DeveloperAleph Alpha (Germany)
Release dateOctober 3, 2026
Total parameters~78B (78,103,074,560)
Active parameters~3.46B per token
ArchitectureMoE, 50 layers, SWA:GQA attention at 4:1
Experts384 per layer (1 shared + 6 routed active)
Vocabulary128,000 (custom tokenizer for German morphology)
Context262,144 native / up to 1,048,576 (262K or less recommended)
LanguagesGerman and English
FeaturesReasoning modes (low/medium/high/none), tool calling
Weight precisionFP8 (128x128 blocks); embeddings, LM head, norms and MoE router in BF16
LicenseApache 2.0 (weights)

Requirements cheat sheet: memory and GPUs by precision

Only the FP8 row comes from the model card. The BF16, 4-bit and 3-bit rows are estimates from parameters times bit width. BF16 is 78.1B x 2 bytes, about 156GB. 4-bit is 78.1B x 0.5 bytes, about 39GB, plus the parts kept in BF16 and quantization scales, giving roughly 39-45GB. 3-bit is 78.1B x 0.375 bytes, about 29GB, plus the same overhead, roughly 30-34GB. KV cache and activations come on top and grow with context length.

Horizontal bar chart of Kolibri-1 weight memory by precision: BF16 about 156GB (estimate), FP8 about 78GB (official), 4-bit about 39-45GB (estimate), 3-bit about 30-34GB (estimate)
PrecisionWeight memoryBasisExample GPU setups
BF16~156GBEstimate2x H200 (282GB) or 3-4x H100
FP8 (official)~78GBOfficial1x H200, 1x B200, 2x H100, 2x A100 80GB
4-bit quant~39-45GBEstimate1x RTX 6000 Blackwell (96GB), or 1x 80GB A100/H100 (little KV headroom)
3-bit quant~30-34GBEstimate1x 48GB-class GPU; a 32GB RTX 5090 leaves no headroom

To estimate whether your own hardware fits, the free VRAM calculator lets you pick a model, quantization and context length. Kolibri-1 is included in its model list.

Official recommended setups and community measurements

The model card lists a minimum of 2x A100 80GB, 2x H100 SXM5, 1x H200, 1x B200 or 1x B300, and a recommendation of 2x H100 SXM5, 2x H200, 1x B200 or 1x B300. Since about 78GB of FP8 weights will not fit on a single 80GB card, 80GB-class GPUs need two (160GB total). A single H200 (141GB), B200 or B300 holds the weights with the rest going to KV cache.

A community report on NVIDIA's developer forum ran FP8 Kolibri-1 on a DGX Spark (GB10) with 128GB of memory. The weights alone took 73.5 GiB. On vLLM 0.29.0 with a 2,048-token prompt, it measured about 1,402 tok/s prefill and 20.4 tok/s decode at zero context, and about 19.6 tok/s decode at 32k context. Loading took 18 to 25 minutes. A later vLLM 0.30.1rc1 build with SM121 optimizations reportedly reached about 48 tok/s decode. This is a single user's report, not an official figure.

Can it run on a Mac or consumer RTX cards? (estimates)

What follows is an estimate from memory capacity, not an official validation. On Apple Silicon, a Mac Studio with 128GB or 192GB of unified memory has room for the 78GB FP8 weights or the 39-45GB 4-bit estimate. However, macOS caps how much memory the GPU can use, and the model card's vLLM FP8 recipe targets NVIDIA GPUs, so running on a Mac depends on community MLX or GGUF conversions, which we could not confirm at the time of writing.

On consumer RTX cards, a 4-bit build (about 39-45GB) would just fit across two 24GB cards (48GB total) and leave comfortable room across two 32GB RTX 5090s (64GB total). Because only about 3.46B parameters are active, offloading experts to system RAM may still give usable speed, but that too is speculation. Remember KV cache is needed on top of the weights, so leave headroom.

Launching with vLLM (from the model card)

The model card's vLLM command sets an FP8 KV cache plus Kolibri-specific reasoning and tool-call parsers.

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

To use contexts beyond 262,144 tokens, append the following. The official guidance is still to stay at or below 262,144 for efficiency and quality on complex tasks.

--max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'

Recommended sampling is temperature=1.0, top_p=0.97, top_k=128. Reasoning depth is chosen through the chat template as low, medium, high or none. In the DGX Spark report, thinking mode was on by default and reasoning appeared in a separate response field.

What it means for Japanese companies: Japanese is not a main target

Kolibri-1 targets German and English, and Japanese is not a main target. Its Japanese quality is not stated officially, so you should not assume it works for Japanese business documents or customer support, and testing on your own data is essential before adoption. For Japanese-centric work, looking at other open-weight models with strong Japanese first is the realistic route.

There are still cases worth considering: companies that handle exchanges with German-speaking partners or sites in-house, teams that need to process German and English documents locally and keep them confidential, or anyone studying how to design a sovereign, keep-data-in-house AI setup. The Apache 2.0 weights make commercial use easy, though the model card says the license covers the weights only, not the underlying code, architecture or training methods.

How it compares with other large open-weight MoE models

Next to other large MoE releases from 2026, Kolibri-1 sits on the small end for both parameters and memory. Figures for the other models come from our own requirements articles.

ModelTotal / activeMemory guideLicenseFocus
Kolibri-1 (this article)~78B / ~3.46BFP8 ~78GB, 4-bit ~39-45GB (est.)Apache 2.0German and English
Qwen3.8-Flash-Next125B / ~6B4-bit ~111GB, BF16 ~354GBQwen Community License 1.0Multilingual, multimodal
GLM-5.3-Flash320B / 18B4-bit ~190GB, BF16 ~740GB (est.)MITMultimodal, 1M context
MiMo-V2.6-Flash309B / 15BFP8 ~320-350GB, 4-bit ~110-190GB (est.)See article1M context

Kolibri-1 fits on two 80GB GPUs or a single H200, and a 4-bit build could come within reach of one large workstation GPU. If you need multilingual coverage including Japanese, models such as the Qwen line in the table are often the better fit.

Caveats before adopting

- Language coverage: mainly German and English; Japanese quality is not stated officially
- Sub-8-bit figures are estimates: the model card gives about 78GB for FP8 only; real quantized sizes vary by converter
- KV cache is extra: FP8 KV is recommended, and pushing to 262K or 1M adds a lot of memory beyond the weights
- Slow startup: the DGX Spark report took 18 to 25 minutes to load weights
- License scope: Apache 2.0 covers the weights, not the underlying code or training methods (per the model card)
- Benchmarks are vendor-reported: the card cites 75.5% overall in English and 70.8% in German; verify on your own workloads

FAQ

How much VRAM does Kolibri-1 need?

The official FP8 weights are about 78GB. The minimum setups are 2x A100 80GB, 2x H100, 1x H200, 1x B200 or 1x B300, with extra room needed for KV cache. A 4-bit build is estimated at roughly 39-45GB.

Does it run on an RTX 4090 or a Mac?

The official validation is on datacenter GPUs. By capacity alone, a 4-bit build would fit on two 24GB cards or a Mac with 128GB or more, but that is an estimate and a suitable quantized build must exist.

Can I use it in Japanese?

The official target languages are German and English, and Japanese quality is not stated. Test on your own data before using it for Japanese work.

What is the context length?

262,144 tokens natively, with 1,048,576 validated as the maximum. The official recommendation is 262,144 or less for efficiency and quality.

Is commercial use allowed?

The weights are Apache 2.0, so commercial use is easy. The model card says the underlying code, architecture and training methods are outside the license.

Why is memory large when only 3.46B parameters are active?

The experts used change from token to token, so all expert weights must sit in fast memory. Compute is light, but memory is sized for the full ~78B parameters.

Summary

Kolibri-1 is an Apache 2.0 open-weight MoE with about 78B total and 3.46B active parameters, specialized for German and English. Its FP8 weights take about 78GB and run on 1x H200 or 2x H100 and up. A 4-bit build is estimated at 39-45GB, but official validation covers datacenter GPUs only. Since Japanese is not a main target, Japanese companies should consider it for German-speaking use cases or as a sovereign-AI reference, and only after testing on their own data.

Feel free to contact us

Contact Us