TL;DR
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
Cactus Compute released Whistle, a 16.9 MB speech-recognition model designed to run locally on a CPU. The company says it supports seven languages and reports fast inference and competitive word-error rates, though results vary across benchmarks and independent validation is not provided in the report.
Cactus Compute has released Whistle, a 16.9 MB speech-recognition model that the company says can transcribe audio locally on a CPU in seven languages. The launch targets devices including phones, wearables, robots, smart-home products and vehicles, where a compact model may reduce reliance on sending voice recordings to cloud services.
According to the company’s October 2 announcement, Whistle accepts 16 kHz mono audio in clips of up to 30 seconds and supports English, German, French, Spanish, Italian, Dutch and Polish. It can detect the language automatically or use a language specified by the caller. Cactus Compute says the model also returns word-level start and end times with probabilities, and can produce speech embeddings without generating a transcript.
The company says Whistle runs in its C++ CPU engine with no dependencies and shares that engine with Needle, Cactus Compute’s existing model. Its description says the encoder processes audio into frames and the decoder generates text using five-beam search. A configurable decoder depth allows different numbers of decoder layers to be selected, while all eight encoder blocks remain active. The transcript limit is 320 tokens.
Cactus Compute’s published comparison reports a 11.1 millisecond time to first token for a 10-second audio clip on an Apple M4 Pro CPU. It lists 73.2 milliseconds for Whisper base and 22.8 milliseconds for Moonshine tiny v2 under the company’s test setup. The report also gives Whistle a decode rate of 1,319 tokens per second, compared with 266 for Whisper base and 262 for Moonshine tiny v2. These are vendor-reported measurements, not an independent evaluation.
Local Transcription on Smaller Devices
A model file of 16.9 MB could make speech recognition easier to deploy on devices with limited storage or processing resources. If it performs as described, local inference may also help products operate without a constant network connection and reduce the need to transmit recorded speech to a remote service. Those benefits depend on how a model is integrated and on the device running it; the release material does not quantify battery use, memory requirements or privacy protections beyond saying audio stays on-device in the browser demonstration.
The reported results also show why compact-model comparisons need more than a single speed or accuracy figure. Cactus Compute says Whistle leads the listed models on several datasets, including LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. It also reports that Whisper base scores better on TED-LIUM, AMI and the MLS average. The trade-off between size, speed and recognition quality will matter to developers choosing a model for a particular setting, especially where errors in names, commands or noisy speech carry consequences.
on-device speech recognition software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Cactus Compute Tested Whistle
The release compares Whistle with Whisper base and Moonshine tiny v2, listing model sizes of 16.9 MB, 145.3 MB and 41.9 MB, respectively. The reported runtime tests used each model’s official runtime at its defaults: Whistle’s C++ engine with five beams, OpenAI Whisper, and Moonshine Voice in non-streaming mode over the whole audio. The speed figures refer to ten seconds of audio on an Apple M4 Pro CPU.
The company defines time to first token as the interval from audio input to the first generated token. Its decode-rate figure counts tokens divided by elapsed time after that point, so it excludes encoding time. The report says Whistle’s time to first token increases with clip length: 5.9 milliseconds for five seconds of audio, 11.1 milliseconds for ten seconds and 36.3 milliseconds for 30 seconds. It says Whisper pads every input to 30 seconds, making its first-token timing flat across clip lengths in this comparison.
Dataset coverage is not identical across the models. Cactus Compute says Moonshine is English-only and that Whisper has no published results for SPGISpeech, Earnings-22 or the cleaned AMI dataset used for the other models. The report identifies Whisper’s AMI result as AMI-IHM, a different subset. These gaps limit direct comparisons across the full chart.
““one 16.9 MB file, runs on the CPU with no dependencies””
— Cactus Compute, in its release announcement
As an affiliate, we earn on qualifying purchases.
Independent Results and Device Costs
The published report is from Cactus Compute; it does not provide independent test results or enough detail to assess performance across a broad range of hardware. The company’s measurements use an Apple M4 Pro CPU, and actual speed, memory use and power consumption may differ on phones, embedded systems or microcontrollers. The release calls those devices target platforms but does not specify a tested device list.
It is also unclear from the report how Whistle performs in real-world conditions such as background noise, overlapping speakers, accents and longer conversations beyond its 30-second clip limit. The benchmark chart gives word-error-rate comparisons, but the supplied material does not include the full per-dataset scores or the testing details needed to reproduce them. The local-processing statement describes the browser demo; the report does not set out deployment-specific data-handling practices for products that adopt the model.
voice transcription app for smartphones
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Availability and Adoption Details
Cactus Compute’s announcement presents Whistle as available through a browser demonstration, where users can test clips of up to 30 seconds in the supported languages. The company says the first use downloads the 16.9 MB model. The supplied release does not specify licensing terms, package distribution details, supported hardware requirements or a roadmap for updates.
Developers considering the model will need those details, as well as tests on their intended devices and with their own audio. Further independently reproduced accuracy and performance measurements would help establish whether the reported speed and word-error results carry over beyond the company’s comparison setup.
privacy-focused speech recognition hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Whistle?
Whistle is a speech-recognition model released by Cactus Compute. The company says it runs locally on a CPU and has a model size of 16.9 MB.
Which languages does Whistle support?
Cactus Compute lists English, German, French, Spanish, Italian, Dutch and Polish. It says Whistle can detect the language automatically or accept a specified language.
Does Whistle send audio to the cloud?
Cactus Compute says audio in its browser demonstration stays on the device. The announcement does not describe data handling for every possible third-party integration or deployment.
How does Whistle compare with Whisper and Moonshine?
In Cactus Compute’s tests on an Apple M4 Pro, Whistle had a reported 11.1 millisecond time to first token for ten seconds of audio, versus 73.2 milliseconds for Whisper base and 22.8 milliseconds for Moonshine tiny v2. The company reports that accuracy varies by dataset, with Whisper ahead on some measures. These results have not been independently validated in the supplied material.
What are Whistle’s clip and transcript limits?
The release describes transcription for 16 kHz mono clips up to 30 seconds, with transcripts capped at 320 tokens. Cactus Compute also says the model can return word timestamps and speech embeddings.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
