Cactus Compute has released Whistle, an open speech recognition model small enough to fit in a 16.9 MB download, and says it runs on hardware as modest as a microcontroller. The size matters because a model that small does not need a server. Audio can be transcribed where it is recorded.
That changes three things for anyone building voice features. Nothing has to be uploaded, so a recording never leaves the device, which removes a privacy question before it is asked. No per-minute cloud bill arrives, because the work happens on a processor the customer already paid for. And no network round trip occurs, so a spoken command is not delayed by a bad connection or a dead zone.
Whistle runs on an ordinary processor. It does not need a graphics chip, which is the usual requirement for modern AI models and the reason most of them stay in data centers. Cactus, writing on its own company blog on October 2, says the model shares its engine with Needle, a separate Cactus model that reads a request and a list of available functions, then returns the function call to make. Load both and a single program can take a spoken clip such as “turn off the kitchen lights” and return the matching command, with no transcript handled by the developer in between.
The language support is wider than the file size suggests. Whistle handles three Romance languages (French, Spanish, and Italian), three Germanic ones (English, German, and Dutch), and Polish as its only Slavic entry. It detects the language on its own unless the developer names one. It also returns the time each word was spoken and can bias its output toward names or terms the developer supplies, such as hard-to-spell surnames.
The target hardware is where the pitch gets ambitious. Cactus lists phones and wearables, home devices and cars, and robots and microcontrollers. The company says its engine ships prebuilt for seventeen platforms, among them the browser, Windows on ARM, RISC-V chips, and the Apple and Android mobile systems. Those are Cactus’s claims about compatibility. The blog post does not show the model running on the smallest of those chips.
The performance figures come from Cactus itself. The company compared Whistle with Whisper base, OpenAI’s small open speech model, and Moonshine tiny v2, a compact English-only model. On ten seconds of audio, running on an Apple M4 Pro processor, Whistle produced its first word in 11.1 milliseconds. Whisper base took 73.2 and Moonshine took 22.8. Whisper base is also roughly eight and a half times larger, at 145.3 MB, and Moonshine comes in at 41.9 MB.
Accuracy is a more mixed picture, and Cactus reports it candidly. Whistle posts a lower word error rate on LibriSpeech, SPGISpeech, Earnings-22, and the FLEURS average. Whisper base wins on TED-LIUM, AMI, and the MLS average. Cactus measured Whistle’s own scores across 86,174 utterances and took the rivals’ numbers from the figures their authors published. The post names no independent benchmark, and the comparison was run on a laptop-class chip, not the low-power hardware it is meant for.
Cactus also says it screened for leakage. It matched audio checksums and speaker IDs from each reported test set against the material Whistle learned from, and found no overlap. Outsiders can check that claim for themselves, since the weights are public on Hugging Face and the code is on GitHub.
The practical question for a product team is whether the accuracy holds on real microphones, accents, and background noise, none of which a clean benchmark captures. Teams shipping voice control on devices with weak connectivity or strict privacy rules can test that against their own recordings this month, before deciding whether a cloud transcription contract is still worth paying for.
Cactus Compute, company blog post “Whistle: Speech to Text in 16.9 MB” by Jakub Mroz and Henry Ndubuaku, October 2, 2026.