Open & local AI

Whistle brings 16.9 MB on-device speech recognition to edge devices

Cactus Compute releases Whistle, a tiny open-weight speech model that runs on CPUs without dependencies and integrates with the Needle engine for direct tool calling.

A small microchip representing Whistle next to a large stone representing Whisper, illustrating size difference.
Illustration generated for this article

Cactus Compute has released Whistle, a new open-weight speech recognition model designed specifically for resource-constrained environments. Launched on October 2, 2026, this 16.9 MB model runs entirely on the CPU with no external dependencies, targeting mobiles, wearables, robots, and microcontrollers. It shares its underlying C++ engine with Needle, allowing developers to load both text and speech capabilities from a single binary.

What happened

Whistle represents a significant shift toward ultra-lightweight, on-device audio processing. Unlike larger models that require cloud connectivity or substantial local storage, Whistle operates as a single file that loads directly into memory. The model supports transcription for up to 30 seconds of audio in seven languages: English, German, French, Spanish, Italian, Dutch, and Polish. Language detection is automatic unless a specific language is specified by the developer.

The release emphasizes privacy and efficiency. Audio data never leaves the device during processing. Beyond simple transcription, Whistle provides word-level timestamps with start, end, and probability metrics derived from the decoder's attention mechanism. It also offers speech embedding capabilities, outputting one row per 80 ms frame from the encoder without generating a full transcript. This flexibility allows engineers to use the model for various tasks, from command recognition to semantic audio analysis.

Integration with the existing Needle ecosystem is a core feature. The same needle_load function can read Whistle models, Needle text models, or both simultaneously. This unified approach means a single application binary can handle voice input, process it into text, and immediately execute corresponding tool calls without intermediate handling by the host application.

How it works

The architecture relies on shared components with the Needle model to minimize code duplication and memory footprint. The front end processes 16 kHz mono audio using a 25 ms window and 10 ms hop, creating 80 log-mel bins band-limited to 250-3500 Hz. A convolutional stem reduces the frame count significantly, resulting in 375 frames for a 30-second clip, with each frame representing 80 ms of audio.

The encoder consists of eight Simple Attention blocks using mHC residual lanes and a Monarch Hadamard MLP instead of a traditional feed-forward network. Crucially, the attention mechanism is not causal, allowing frames to attend to future context within the clip. The decoder uses eight Laddered Simple Attention blocks with gated cross-attention to read from the encoder. This design allows the projections to run once per clip, meaning beam search operations only cache short transcripts rather than re-processing audio.

Decoding uses five beams scored by length-normalized log probability. An Aho-Corasick automaton runs alongside the beams to support keyword biasing, increasing the likelihood of specific phrases during search. The model includes a silence detection mechanism that measures loudness range before decoding begins, returning empty results for quiet inputs to save compute. The decoder depth is selectable at load time via --audio-depth, allowing further trade-offs between accuracy and speed while keeping the encoder fixed.

Key details

  • Model size: 16.9 MB, significantly smaller than Whisper base (145.3 MB) and Moonshine tiny v2 (41.9 MB).
  • Performance: Achieves 11.1 ms time-to-first-token and 1,319 tokens per second on an Apple M4 Pro CPU.
  • Languages: Supports English, German, French, Spanish, Italian, Dutch, and Polish with automatic detection.
  • Outputs: Provides transcription, word-level timestamps with probabilities, and speech embeddings.
  • Deployment: Prebuilt binaries available for seventeen targets including macOS, Linux, Android, iOS, RISC-V, and WASM.
  • Integration: Uses the same .cact container format as Needle, enabling combined speech and text workflows in one engine.

Why it matters

For developers building embedded systems or mobile applications, Whistle removes the need for complex cloud APIs or heavy local libraries. The ability to run speech recognition on a microcontroller or wearable without internet connectivity opens new possibilities for privacy-focused and offline-first products. The small footprint means it can coexist with other application logic without consuming excessive memory or battery life.

The integration with Needle simplifies the development of voice-controlled agents. Instead of building separate pipelines for speech-to-text and intent recognition, developers can use a single engine to transcribe audio and trigger function calls directly. This reduces latency and engineering overhead, as the application does not need to parse intermediate text strings. The unified binary approach also simplifies deployment and updates across diverse hardware platforms.

What you can do

  • Install the package via pip using pip install cactus-needle to start experimenting with the Python API.
  • Test real-time transcription in your terminal using the needle whistle playground command.
  • Compare performance against other models using needle whistle compare to see latency and accuracy differences side-by-side.
  • Implement keyword biasing by passing a list of specific terms to improve recognition of names or technical jargon.
  • Deploy prebuilt binaries to target devices such as Android, iOS, or RISC-V boards using the provided needle download tools.
  • Use the needle.Whistle() object in Python to generate speech embeddings for audio classification tasks without full transcription.

Tools from the Bytechap store

$89

DocBento

Self-hosted document management that reads every scan and answers with page citations.

Live demo

Keep reading

All stories