Inflect-Micro-v2: complete voice in 9.36M parameters (huggingface.co)

🤖 AI Summary
Inflect has officially released Inflect-Micro-v2, a compact text-to-waveform speech synthesis model boasting just 9.36 million parameters. This model prioritizes high-quality, fixed-voice English text-to-speech (TTS) synthesis while enabling CPU or CUDA inference. Notably, Inflect-Micro-v2 is designed to effectively handle long text inputs with a deterministic output by using fixed seeds for reproducibility, marking a significant advance in local TTS technology that is both lightweight and efficient. This release is particularly significant for the AI and ML community as it represents a major step towards optimizing speech synthesis models for quality without requiring extensive computational resources. Inflect-Micro-v2 marks an improvement over its predecessor by demonstrating a 66.2% preference rate in community studies against established TTS models like KittenTTS and Piper, while maintaining compact architecture and rapid processing speeds. With a streamlined architecture based on the VITS family, the system integrates advanced features like punctuation-aware segmentation and a controlled output of 24 kHz mono waveforms, making it a promising tool for developers seeking high-quality, locally deployable TTS solutions.
Loading comments...
loading comments...