AI Briefing
KOSign in

Cactus Compute Releases Whistle, a 16.9 MB On-Device Speech Recognition Model

·2026.10.02 09:00

Key point

The model achieves an 11.1 ms time-to-first-token on an Apple M4 Pro CPU and supports seven languages.

Details

Cactus Compute has released Whistle, an open speech recognition model designed for mobiles, wearables, robots, and microcontrollers. The model is a single 16.9 MB file that runs on the CPU with no dependencies, sharing the same C++ engine and quantization as the company's previous model, Needle.

Core Capabilities and Architecture

Whistle performs three primary functions on-device: transcription, word timestamping, and speech embedding. It supports seven languages (English, German, French, Spanish, Italian, Dutch, and Polish) with automatic language detection. The architecture utilizes a non-causal encoder with eight Simple Attention blocks and a decoder with eight Laddered Simple Attention blocks. A key architectural feature is the gated cross attention in each decoder layer, which reads encoder outputs once per clip, allowing five-beam search to share the same audio processing rather than re-encoding the audio for each beam.

Performance Benchmarks

On an Apple M4 Pro CPU, Whistle demonstrates significant speed advantages over competitors like Whisper base and Moonshine tiny v2:

  • Time to first token: 11.1 ms (compared to 73.2 ms for Whisper base and 22.8 ms for Moonshine tiny v2).
  • Decode speed: 1,319 tokens per second (compared to 266/s for Whisper base and 262/s for Moonshine tiny v2).
  • Model size: 16.9 MB (compared to 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2).

Whistle achieves lower word error rates than Whisper base on LibriSpeech, SPGISpeech, Earnings-22, and the FLEURS average. However, Whisper base performs better on TED-LIUM, AMI, and the MLS average.

Integration and Deployment

Whistle integrates directly with the Needle engine, allowing a single binary to handle both speech and text tasks. The needle_complete function can take an audio clip, transcribe it, and execute tool calls in one step, returning a JSON object with the transcript and function calls. The model is available via pip install cactus-needle and supports prebuilt binaries for seventeen targets, including macOS, Linux, Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, and the browser.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.