🤖 AI Summary
MichiAI has released a new full-duplex speech language model with 530 million parameters, boasting approximately 80ms latency. This model represents a significant advancement in speech naturalness and quality, overcoming previous audio artifacts like a metallic buzz through pure adversarial training. By embracing natural audio characteristics such as breaths and laughter, MichiAI generates a more realistic conversational experience, moving away from the overly smoothed output typical in traditional models.
Key innovations include a zero-coherence loss training regimen that preserves language capabilities while learning from conversational data, and a seamless inference codebase optimized for real-time streaming. Unlike existing systems that rely on rigid turn-taking rules, MichiAI intensely mimics human conversational dynamics, producing natural backchanneling responses and maintaining dialogue flow. Additionally, it can express subtle human-like sounds such as laughter and sighs, enhancing its conversational realism. The model is designed for live interaction, with the next step involving the development of a web client for users to interact with MichiAI directly.
Loading comments...
login to comment
loading comments...
no comments yet