edge ai

How Whistle Fits Offline Speech Recognition Into 16.9 MB

How Whistle Fits Offline Speech Recognition Into 16.9 MB

Imagine saying “turn off the kitchen lights” to a watch and having the command reach the light without a round trip to a server. That small moment hides a difficult engineering problem: speech recognition usually wants plenty of memory, compute, and a reliable connection.

Whistle takes a different route. It packages speech-to-text into a 16.9 MB model that runs on a central processing unit (CPU), accepts up to 30 seconds of audio, and supports English, German, French, Spanish, Italian, Dutch, and Polish. The same model also provides word timestamps and speech embeddings, making it useful for apps that need more than a block of text. (cactuscompute.com)

First, shrink the sound

A microphone produces a waveform: a long stream of measurements describing air pressure over time. Feeding that raw stream directly into a neural network would be wasteful, so Whistle first converts it into a compact audio representation.

The model expects 16 kHz mono audio, meaning 16,000 samples per second from one channel. It examines the signal through overlapping 25-millisecond windows, moving forward by 10 milliseconds at a time. Each window becomes 80 log-mel values. A log-mel representation records how much energy appears in different frequency bands, using a pitch scale that roughly follows human hearing.

Thirty seconds of audio produces about 3,000 frames. A convolutional stem then reduces that sequence three times, leaving 375 frames, each representing roughly 80 milliseconds. A convolutional layer applies the same small pattern detector across neighboring inputs; here, it acts like a careful audio downsampler. Fewer frames mean less work for every later attention operation.

The encoder hears the whole clip

The encoder turns those reduced frames into a context-rich description of the recording. Whistle uses eight Simple Attention blocks. Attention is a weighted lookup: each frame decides which other frames matter, then combines information from them. Because this encoder is not causal, a sound near the beginning can use evidence that arrives later in the clip. That is useful when the final consonant of a word changes what the earlier syllables meant.

The blocks also use several compact architectural choices borrowed from Needle. Four residual lanes provide parallel paths for carrying updates through the network, while a structured mixer replaces the large dense feed-forward layer found in a conventional transformer. The goal is not to imitate a desktop-sized model with fewer numbers; it is to spend each operation where a small device can afford it.

The decoder reads audio without replaying it

The decoder generates text one token at a time. A token is a small piece of text, such as part of a word, rather than necessarily a complete word. To choose each token, the decoder uses cross-attention, a mechanism that lets a text-generating network read representations produced by an audio network.

Whistle adds a learned gate to that connection. In simplified form, each decoder layer behaves like this:

new_state = old_state + sigmoid(gate) * attended_audio

The sigmoid function turns the gate into a value between zero and one, allowing each layer to control how strongly it uses the audio context. The encoder’s key and value tensors—the pieces cross-attention uses to find relevant audio information—are projected once when the clip arrives and then cached.

That cache matters during beam search. Beam search keeps several promising transcript candidates instead of committing to the first likely word. Whistle uses five beams, but it does not run the audio encoder five times. Each candidate gets its own small text-generation cache while all of them reuse the prepared audio context. The result is a much more sensible cost profile for CPU inference. (cactuscompute.com)

The decoder is also laddered. The --audio-depth setting selects a decoder depth at load time, while the full eight-layer encoder still runs. A smaller depth can reduce memory and latency for a constrained device; a deeper one gives the decoder more room to resolve difficult speech. The same runtime can therefore target more than one hardware profile without requiring a completely separate model design.

One model, three useful outputs

Whistle is not limited to plain transcription:

  • Text: the recognized sentence and detected language.
  • Word timestamps: the start time, end time, and probability for each word, useful for subtitles, search, highlighting, and audio editing.
  • Speech embeddings: a vector for each 80-millisecond frame, allowing an app to compare or retrieve sounds without decoding a transcript.

There is also a practical silence check. Before beam search begins, the engine measures the clip’s loudness range. If the signal falls below its threshold, it returns an empty result instead of trying to invent speech. That small guard can prevent a voice interface from reacting to a fan, an empty room, or a microphone that was never opened.

A short Python path

For a 16 kHz WAV file, the basic Python interface is deliberately small:

pip install cactus-needle
import needle

result = needle.transcribe(
 'clip.wav',
 word_timestamps=True,
 keywords=['Kraków', 'Needle'],
)

print(result['text'])
print(result['language'])

Keyword biasing gives selected names, places, or product terms extra weight during decoding. That is useful when an ordinary language model would prefer a more common spelling. Microphone capture and other sample rates require the package’s optional microphone extras; a standard 16 kHz WAV needs no additional audio stack. (cactuscompute.com)

From spoken words to tool calls

The most interesting part appears when Whistle loads beside Needle, the compact model designed for structured tool calls. A tool call is a machine-readable instruction such as set_lights(room='kitchen', on=False). Instead of asking an application to transcribe audio, clean the text, pass it to another model, and parse the response, the shared engine can accept the clip and return one structured result.

needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav

The response can contain the function calls alongside fields such as audio_text and audio_language. The transcript stays inside the engine until the final structured response is produced. Whistle and Needle share the same C++ runtime, model container, and quantization—the practice of storing model weights at reduced numerical precision to save space—so an application does not need a second inference stack for speech. (cactuscompute.com)

What the benchmark numbers mean

On an Apple M4 Pro CPU with ten seconds of audio, Whistle’s reported size was 16.9 MB, compared with 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2. Its time to the first generated token was 11.1 milliseconds, versus 73.2 milliseconds and 22.8 milliseconds for those comparison runs. After the first token, the measured decode rates were 1,319, 266, and 262 tokens per second respectively. These results use each project’s official runtime and default settings, so they show a useful direction rather than a promise for every phone or microcontroller.

Accuracy is less one-sided, which is exactly what a fair comparison should reveal. Whistle leads on several reported LibriSpeech splits, SPGISpeech, Earnings-22, and the FLEURS average. Whisper base leads on TED-LIUM, AMI, and the MLS average. Word error rate, the fraction of words recognized incorrectly, depends heavily on the speakers, recording conditions, language mix, and scoring rules in each benchmark.

The small-model lesson

Whistle’s achievement is not only the 16.9 MB file. The real design is the combination of a reduced audio representation, a compact encoder, cached cross-attention, selectable decoder depth, and a runtime that can turn speech into local actions. Speech recognition becomes an input layer for an embedded application rather than a cloud service bolted onto the side. On a watch, robot, car, or microcontroller, that difference can decide whether a voice feature feels like part of the device or like a fragile network request.

ahsan

ahsan

Hello! I am Mr Ahsan, the writer of the Website. I am from Netherland. I like to write about technology and the news around it.

Comments (0)

No comments yet. Be the first to respond!

Leave a Comment

Your comment will be visible after review.