Skip to main content
株式会社オブライト
AI2026-10-097 min read

Whistle Explained: Cactus's 16.9MB Speech-to-Text Model

Whistle is Cactus's 16.9MB open speech-to-text model for CPU, Apache 2.0. It supports 7 languages but not Japanese. Speed vs Whisper base, setup, best uses.


TL;DR — What Is Whistle

Whistle is an open speech-to-text (ASR) model released by Cactus Compute on October 2, 2026. Its defining trait is that the whole model is a single 16.9MB file that runs on a CPU with no dependencies. It is licensed under Apache 2.0, and the parameter count has not been disclosed. It runs on the same C++ engine, the same .cact container format, and the same quantization as the company's tiny agent LLM (see our Needle 2 explainer), so you can chain speech recognition to tool calling in a very small footprint. It hit #1 on the Hacker News front page on October 8, 2026, with about 768 points.

Important note: Whistle does not support Japanese or Chinese. It covers seven languages: English, German, French, Spanish, Italian, Dutch, and Polish. If you need Japanese transcription, choose a Whisper-family or other Japanese-capable model (see the comparison table below).

Spec Sheet at a Glance

ItemValue
DeveloperCactus Compute
Release dateOctober 2, 2026
Model size16.9MB (single file)
ParametersNot disclosed
RuntimeCPU only, no dependencies
LicenseApache 2.0 (open weights, free)
LanguagesEnglish, German, French, Spanish, Italian, Dutch, Polish (no Japanese or Chinese)
Input length per passUp to 30 s (16 kHz mono)
Target devicesMobile, wearables, robots, smart home, automotive, microcontrollers

What It Can Do

- Transcription: processes up to 30 seconds of 16 kHz mono audio per pass; language is auto-detected or set explicitly, e.g. language="de"
- Word timestamps: start, end, and probability for each word
- Speech embeddings: one encoder row per 80 ms frame, available without decoding
- Keyword biasing: nudge the model toward names or command words you specify
- Silence detection: returns an empty transcript before decoding when it detects silence

Silence detection helps with the classic problem of phantom text appearing over quiet audio. As covered below, though, community members still reported false output on TV audio.

How It Works (Architecture Highlights)

Input audio becomes 80 log-mel bins (25 ms window, 10 ms hop, 250–3500 Hz), and a convolutional stem compresses 3000 frames down to 375. The encoder is 8 non-causal Simple Attention blocks. The decoder is 8 Laddered Simple Attention blocks of width 512, with GQA (8 query heads, 2 KV heads), engram lookups at layers 3 and 7, and gated cross-attention in every decoder layer. Decoding uses beam search with 5 beams and up to 320 tokens; the vocabulary is 8,192 plus 7 language tokens. The --audio-depth option selects how many decoder layers to use at load time (minimum 2).

Whistle's pipeline: up to 30 s of 16 kHz mono audio → 80-bin log-mel → conv stem (3000→375 frames) → 8-block encoder → 8-block decoder (GQA) → text with word timestamps; encoder output is also available as one embedding per 80 ms.

Speed Comparison (Apple M4 Pro, CPU, 10 s Clip)

ModelSizeTTFTThroughput
Whistle16.9MB11.1 ms1,319 tok/s
Whisper base145.3MB73.2 ms266 tok/s
Moonshine tiny v241.9MB22.8 ms262 tok/s

All figures are vendor-reported. Whistle's TTFT is reported as 5.9 ms on a 5 s clip and 36.3 ms on a 30 s clip. For accuracy, the vendor compared WER across 86,174 utterances: Whistle comes out ahead of Whisper base and Moonshine tiny v2 on LibriSpeech clean/other, SPGISpeech, Earnings-22, and the FLEURS average, while Whisper base is ahead on TED-LIUM, AMI, and the MLS average. Exact numbers appear only in the vendor's chart, so we describe the trend only.

Platforms and Installation

Prebuilt engines are provided for 17 targets, including macOS, Linux, Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, the browser, and WASI; each ships a needle binary, libneedle.a, and needle.h. Weights are on Hugging Face, the engine is at the needle3 repo, and the source is on GitHub.

From Python, pip is the fastest route.

pip install cactus-needle

For microphone capture or audio that is not 16 kHz, install the mic extra (soxr, sounddevice).

pip install "cactus-needle[mic]"

The Shortest Way to Use It

In Python, transcription is one line. Options cover word timestamps, keywords, and language.

import needle

text = needle.transcribe("clip.wav")["text"]

# With options
r = needle.transcribe(
    "clip.wav",
    word_timestamps=True,
    keywords=["Cactus", "Needle"],
    language="de",
)

# Get speech embeddings (no decoding)
emb = needle.Whistle().embed(audio)

The CLI covers downloading weights, transcribing files, trying the mic, and comparing against Whisper and Moonshine.

needle download whistle
needle --model whistle.cact --audio clip.wav --audio-word-timestamps
needle whistle playground   # try with the mic
needle whistle compare      # compare with Whisper and Moonshine

To go straight from voice to tool calls, pass both Needle and Whistle models together.

needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav

For native deployment, fetch the platform engine first, then Whistle.

needle download macos-arm64
needle download whistle

How It Differs From Other Speech Recognition Models

ModelSizeJapaneseBest use
Whistle16.9MBNot supportedCommand recognition in English/European languages, tiny devices
Whisper base145.3MBSupportedLightweight multilingual transcription
Whisper large-v3-turboLarge (GPU recommended)SupportedAccuracy-focused transcription
Moonshine tiny v241.9MBTo be confirmed / limitedLightweight English recognition at the edge
Parakeet (NVIDIA)Large (GPU recommended)To be confirmed / limitedHigh-accuracy, fast, English-centric recognition

Whistle's strengths come down to size and first-response speed. For general-purpose transcription accuracy, larger models still have the edge. If you want local Japanese transcription, a Whisper-based app such as the one in our TypeWhisper explainer is the realistic choice. And if you also want Japanese speech output after recognition, see the Irodori-TTS v4 explainer.

Community Reactions (Hacker News)

The Hacker News thread mixed surprise at the size and speed with sober criticism. Many commenters felt raw accuracy falls short of Whisper Large, Whisper large-v3-turbo, and Parakeet. Hallucinations on silence and TV audio were reported, and the public demo has no streaming output. Some also said the name collides with an existing Android app called Whistle, which caused confusion.

Narrower use cases got a different verdict. One user tested 170 command phrases: raw Whistle got 70 right, versus 168 for Qwen ASR 1.7B. After training a small template-restricting network on 10,000 generated utterances, Whistle reached 164, and on another phrase set it scored 206 of 208 against 184 for Vosk small. The takeaway is that the strongest use case is fixed-command voice control on tiny devices (for example, Home Assistant on an Echo Show), not general dictation.

Where It Fits — and Where It Doesn't

- Fits: fixed-command recognition in English and European languages; voice control on wearables, robots, automotive, and microcontrollers; setups that go directly from voice to tool calls
- Doesn't fit: Japanese or Chinese transcription; long, high-accuracy jobs such as meeting minutes; anything that requires streaming output

If Japanese voice input is your goal, pick a Whisper-family model or one that explicitly lists Japanese support. Consider Whistle only for on-device command recognition in English and European languages.

FAQ

Does Whistle support Japanese?

No. It supports seven languages: English, German, French, Spanish, Italian, Dutch, and Polish. Japanese and Chinese are not supported, so use a Whisper-family or other Japanese-capable model for Japanese transcription.

Can it be used commercially?

The weights are released under Apache 2.0, so commercial use is permitted as long as the license terms are followed, and it is free to use.

Does it need a GPU?

No. The model is a single 16.9MB file that runs on a CPU with no dependencies.

Can it transcribe long recordings?

One pass handles up to 30 seconds of 16 kHz mono audio. Longer recordings need to be split into segments.

How does its accuracy compare with Whisper and Moonshine?

In the vendor's WER comparison, Whistle leads Whisper base and Moonshine tiny v2 on LibriSpeech, SPGISpeech, Earnings-22, and the FLEURS average, while Whisper base leads on TED-LIUM, AMI, and the MLS average. Many community members found it less accurate than Whisper Large-class models and Parakeet, so validate it against your own use case.

Feel free to contact us

Contact Us