Whistle runs a local speech-to-text model on CPU
Cactus Compute
Whistle released a local CPU speech-to-text model that transcribes up to 30 seconds of audio for multiple languages and returns word-level timestamps and speech embeddings on-device. The model downloads as a single 16.9 MB file. As a result, speech audio no longer needs to leave the device and can be turned directly into tool calls using the same C++ engine binary as Needle.
Why it matters
Whistle runs a small speech-to-text model locally on a CPU. The newsletter frames it as keeping inference local rather than relying on remote services.