AI Briefing
KOSign in

Cactus Compute Releases Whistle: A 16.9MB ASR Model for Ultra-Small Devices

·2026.10.06 02:27

Key point

Whistle achieves a 4.31 WER on LibriSpeech test-clean, outperforming Whisper base while using 9x less storage.

Details

Cactus Compute has released Whistle, an automatic speech recognition (ASR) model designed for ultra-small devices such as smartwatches, smart home gadgets, and microcontrollers. The model is 16.9MB in size, representing a 9x reduction in file size compared to Whisper base, while offering 6x faster inference speeds.

Performance Metrics

Whistle is a 55M parameter model (36M active) quantized to CQ2bit. It demonstrates competitive accuracy against Whisper base across several benchmarks:

  • LibriSpeech test-clean: 4.31 WER (Whisper base: 4.9)
  • LibriSpeech test-other: 10.49 WER (Whisper base: 11.0)
  • FLEURS average: 21.4 (Whisper base: 24.5)
  • SPGISpeech: 7.65
  • Earnings-22: 19.01

Architecture and Features

The model utilizes a log-mel front end and convolution stem feeding into an audio encoder, paired with a Simple Attention + Hadamard MLP decoder that uses gated cross attention. Key features include:

  • Laddered Decoder: Similar to Needle, allowing deployment at any depth from 2 layers up for flexible resource management.
  • Keyword Biasing: Enhances recognition of specific names during beam search.
  • Word Timestamps: Generated via decoder attention, enabling precise seeking and highlighting.

Platform Support

Whistle supports 17 platforms, including macOS, Linux (x86-64, ARM64, ARMv7, RISC-V, MIPS32), Windows (x64, ARM), Android, iOS, watchOS, tvOS, and browser environments via WebAssembly and WASI components.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.