Fine tuning STT models to help with my Lisp (vivekkairi.com)

🤖 AI Summary
A recent development in fine-tuning speech-to-text (STT) models has demonstrated significant advancements in performance and usability for individual users. After initially fine-tuning OpenAI's Whisper Turbo—achieving a 6.5% word error rate (WER) but facing speed issues—a user opted for the Parakeet model, which combines a Conformer encoder with a small LSTM decoder. Parakeet showed promise, initially yielding a WER of 22-25%, but struggled with real-world audio scenarios due to a lack of silence in the training data. An innovative solution involved augmenting the dataset by adding silence to audio clips, leading to a dramatic reduction in WER to 2.96% on training samples and 6.14% on the test set. This development is significant for the AI/ML community as it highlights the importance of realistic training data in model performance, particularly in personalized speech recognition tasks. The successful application of transfer learning techniques, such as full encoder fine-tuning and thoughtful dataset augmentation, showcases how user-centric adjustments can lead to practical improvements. Moreover, the model's efficient use of resources, with a notable reduction in size from 2.4GB to 622MB through ONNX conversion, demonstrates the potential for deploying lightweight, high-performance models in real-world applications, making advanced AI more accessible to daily users.
Loading comments...
loading comments...