🤖 AI Summary
Wispr Advanced Interfaces Lab has unveiled Canto, a groundbreaking speech recognition model optimized for the complexities of real-world dictation. Unlike traditional models that excel in controlled environments, Canto is designed to operate amidst background noise and various audio conditions, achieving the lowest word error rate (WER) in its category when tested with over 2,300 unique speakers. The innovative training approach involved supervised fine-tuning and reinforcement learning, allowing Canto to adapt to challenging dictation scenarios, such as music and low-volume speech.
The significance of Canto lies in its ability to outperform established models from tech giants like Google and OpenAI, particularly in real-time applications where low latency is crucial. By incorporating a unique training methodology that evaluates the quality of complete transcripts and utilizes specific user corrections, Canto improves accuracy even in unpredictable conditions. This model represents the first step in Wispr's broader quest for advanced speech recognition solutions, with future iterations aimed at enhancing performance in multi-speaker recognition and contextual vocabulary usage. Canto's introduction marks a pivotal moment for speech technology, emphasizing the need for models that can adapt to the nuances of everyday communication.
Loading comments...
login to comment
loading comments...
no comments yet