Gemini 3.8 TTS Playground (simonwillison.net)

🤖 AI Summary
Google has unveiled two new text-to-speech models, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, which significantly enhance the capabilities of AI-driven audio synthesis. With an extensive library of over 2,000 voices, users can create custom voices using just a 30-second sample of any voice they own the rights to, making voice synthesis more accessible for a wider audience. The integration of a "bring-your-own-key" playground interface with GPT-6 Astra takes advantage of the open CORS policy of the Gemini API, enabling seamless user experimentation. This release is significant for the AI/ML community as it simplifies the process of generating multi-character dialogues, allowing developers to define conversations with distinct voices and styles effortlessly. For instance, a demo featuring a playful debate between two pelicans showcased the efficiency of the system, producing over 1.5 minutes of audio in approximately 20 seconds for a mere cost of 2.74 cents. The introduction of these models underscores the ongoing advancements in AI voice generation and highlights competing innovations in the space, setting a new standard for dynamic, interactive audio experiences.
Loading comments...
loading comments...