TL;DR
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
Cactus Compute released Whistle, a 16.9 MB speech-recognition model designed to run locally on a CPU, supporting seven languages and word timestamps. The company reports strong results against Whisper base and Moonshine tiny v2 on several benchmarks, while Whisper leads on others; independent validation and details such as licensing are not provided in the source.
Cactus Compute announced Whistle, a speech-recognition model distributed as a 16.9 MB file and designed to transcribe audio locally on a CPU. The company says it supports seven languages, returns word-level timestamps and can run in the same C++ engine as its Needle model, a design aimed at devices with limited storage or connectivity.
Whistle accepts 16 kHz mono audio of up to 30 seconds per pass. It supports English, German, French, Spanish, Italian, Dutch and Polish, with automatic language detection unless a language is specified. Cactus says transcription, timestamp generation and speech-embedding output all run on the device; its in-browser demonstration downloads the model on the first use and says audio does not leave the device.
The model can return word start and end times with probability scores, or encoder embeddings arranged at one row per 80 milliseconds. The company also describes keyword biasing, which gives supplied phrases additional weight during decoding. Its decoder uses five-beam search, and transcripts are capped at 320 tokens. A silence check can return an empty transcript without running beam search.
Cactus reports a first-token time of 11.1 milliseconds and a decode rate of 1,319 tokens per second for ten seconds of audio on an Apple M4 Pro CPU. In its comparison, Whistle is 16.9 MB, against 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2. These are company-reported measurements using each model’s stated official runtime and defaults, not a guarantee of performance on other hardware.
Local Transcription on Smaller Devices
A small model that can transcribe speech without sending audio to a server could suit phones, wearables, robots and smart-home devices, uses Cactus lists as intended applications. Local processing may reduce reliance on network access and avoid transmitting recordings to a cloud service, although privacy outcomes also depend on how an application stores and handles audio and transcripts.
The reported size and speed could make speech recognition more practical where memory, latency or connectivity is constrained. But the benchmark table is not a single overall ranking: Whistle scores better on some datasets, while Whisper base leads on others. Readers and developers will need to match the reported results to their language, audio type and device before drawing conclusions about accuracy or suitability.
on-device speech recognition software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Whistle Fits Cactus’s Needle Engine
Cactus describes Whistle as sharing components with Needle, its existing model, and says both load into the same C++ engine with the same container and quantisation. The company’s architecture description says the encoder uses eight attention blocks, while the decoder has eight layers and reads encoded audio through gated cross-attention. Developers can select a shallower decoder depth at load time; the encoder still runs all eight blocks.
The comparison covers word error rates on datasets including LibriSpeech, SPGISpeech, Earnings-22, TED-LIUM, AMI, MLS and FLEURS. Cactus says Whistle is ahead of the compared systems on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. Whisper base leads on TED-LIUM, AMI and the MLS average. The report flags gaps in published results: Moonshine tiny v2 is English-only, and Whisper has no reported figures for several listed tests. It also warns that Whisper’s AMI result uses AMI-IHM, a different subset from the AMI figures for the other models.
““audio never leaves your device.””
— Cactus Compute
small speech-to-text model for mobile
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Independent Testing and Deployment Details
The available report is from Cactus Compute, the model’s developer. It does not provide independent benchmark verification, and the supplied material does not include enough detail to assess how results change across CPU models, memory limits, background noise or real-world recording conditions. The quoted speed measurements apply to a ten-second clip on an Apple M4 Pro CPU and to specific runtime settings.
The source does not state Whistle’s software license, commercial-use terms, supported operating systems, minimum memory requirements or a broader release channel beyond the described model and demonstration. It is also unclear whether the reported privacy behavior applies to all integrations or only the browser sandbox. The benchmark gaps and the differing AMI subsets limit direct comparisons across the full table.
offline voice transcription device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Availability and Testing to Watch
Cactus has announced Whistle and provided an in-browser demonstration that downloads the model for local use. The source does not give a separate release timetable, package distribution plan or roadmap for additional languages. The next useful evidence will be independent testing across devices and recording conditions, alongside clearer documentation of licensing, deployment requirements and privacy behavior in third-party applications.
Developers considering the model can compare it with alternatives on their own audio and target hardware, paying attention to accuracy as well as model size and latency. Until those checks are available, Cactus’s benchmark and speed figures should be treated as vendor-reported results, not universal performance measures.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Whistle?
Whistle is a speech-recognition model from Cactus Compute. The company says it runs on a CPU and is distributed as a 16.9 MB file.
Which languages does it support?
Cactus lists English, German, French, Spanish, Italian, Dutch and Polish. The model can detect the language automatically or use a language specified by the user.
Does Whistle process audio locally?
Cactus says transcription, word timestamps and embeddings run on the device. Its web demonstration says audio stays on the device, but the source does not establish that every third-party integration will handle data the same way.
Is Whistle more accurate than Whisper?
Not across every test in Cactus’s comparison. The company reports Whistle ahead on several datasets, while Whisper base leads on TED-LIUM, AMI and the MLS average. Some results are missing, and the AMI subsets are not identical.
How fast is Whistle?
Cactus reports 11.1 milliseconds to the first token for ten seconds of audio on an Apple M4 Pro CPU. The company says latency varies with clip length; results on other processors may differ.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
