Cactus Compute Releases Whistle: A 16.9MB ASR Model for Ultra-Small Devices
Key point
Whistle achieves a 4.31 WER on LibriSpeech test-clean, outperforming Whisper base while using 9x less storage.
Details
Cactus Compute has released Whistle, an automatic speech recognition (ASR) model designed for ultra-small devices such as smartwatches, smart home gadgets, and microcontrollers. The model is 16.9MB in size, representing a 9x reduction in file size compared to Whisper base, while offering 6x faster inference speeds.
Performance Metrics
Whistle is a 55M parameter model (36M active) quantized to CQ2bit. It demonstrates competitive accuracy against Whisper base across several benchmarks:
- LibriSpeech test-clean: 4.31 WER (Whisper base: 4.9)
- LibriSpeech test-other: 10.49 WER (Whisper base: 11.0)
- FLEURS average: 21.4 (Whisper base: 24.5)
- SPGISpeech: 7.65
- Earnings-22: 19.01
Architecture and Features
The model utilizes a log-mel front end and convolution stem feeding into an audio encoder, paired with a Simple Attention + Hadamard MLP decoder that uses gated cross attention. Key features include:
- Laddered Decoder: Similar to Needle, allowing deployment at any depth from 2 layers up for flexible resource management.
- Keyword Biasing: Enhances recognition of specific names during beam search.
- Word Timestamps: Generated via decoder attention, enabling precise seeking and highlighting.
Platform Support
Whistle supports 17 platforms, including macOS, Linux (x86-64, ARM64, ARMv7, RISC-V, MIPS32), Windows (x64, ARM), Android, iOS, watchOS, tvOS, and browser environments via WebAssembly and WASI components.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.