How AI hears accents: An audible visualization of accent clusters (accent-explorer.boldvoice.com)

🤖 AI Summary
BoldVoice fine-tuned a HuBERT audio foundation model on a massive in-house dataset of accented L2 English (they sampled 30 million recordings ≈25,000 hours) to build an accent identifier and then inspected its 768‑dimensional latent space with a 3D UMAP visualization. The model uses only raw audio and accent labels (no transcripts), was trained with all layers unfrozen on A100 GPUs for about a week, and the team published an interactive map (accentoracle.com) where each plotted recording—filtered so predicted and target accents match—is playable via an anonymizing, accent‑preserving voice conversion. UMAP compresses the embeddings to three dimensions, preserving global structure but not all fine-grained info; clicking points plays a standardized version of the sample so listeners can focus on accent patterns rather than speaker identity or noise. Technically and socially significant, the visualization shows that the model’s learned accent groupings often reflect geography, migration, and colonial history more than linguistic taxonomy: e.g., Australian samples sit near Vietnamese ones (suggesting hybrid accents), French, Nigerian and Ghanaian group together, southern vs. northern Indian subcontinent accents form distinct subclusters, and Mongolian lies close to Korean. While distances aren’t objective phonetic measures, these emergent patterns indicate the model captures real phonetic features and sociohistorical diffusion, offering a tool for accent research, model interpretability, bias analysis, and accent training—while reminding us that latent spaces mirror both speech signal structure and real-world contact patterns.
Loading comments...
loading comments...