Whistle: Speech to Text in 16.9 MB (cactuscompute.com)

🤖 AI Summary
Today, a new speech recognition model named Whistle was launched, designed for a variety of devices including mobiles, wearables, and microcontrollers. This compact model, at just 16.9 MB, operates directly on the CPU with no external dependencies, using the same C++ engine as its predecessor, Needle. Whistle provides efficient transcription in multiple languages (English, German, French, Spanish, Italian, Dutch, and Polish) with features like word timestamps and speech embeddings. Its architecture includes advanced techniques such as Simple Attention and gated cross attention, which enhance real-time processing capabilities. The significance of Whistle lies in its lightweight design and efficient performance, which is a major advancement in on-device speech recognition technology. The model achieves lower latency and superior word error rates compared to existing alternatives like Whisper, particularly notable when decoding short audio clips. With the ability to process audio in real-time while delivering accurate transcriptions, Whistle sets a new benchmark for speech recognition in constrained environments, opening up opportunities for improved interaction in smart devices and applications where computing resources are limited.
Loading comments...
loading comments...