Meta AI Unveils Voicebox: Revolutionary Generative AI Model Sets New Standards in Speech Synthesis and Multilingual Capabilities

Katerina Petrova News
Meta AI Unveils Voicebox: Revolutionary Generative AI Model Sets New Standards in Speech Synthesis and Multilingual Capabilities

Meta AI develops Voicebox, the first generative AI model for speech that can generalize across tasks with state-of-the-art performance. Voicebox creates high-quality audio clips and can synthesize speech across six languages, perform noise removal, content editing, style conversion, and diverse sample generation. Voicebox outperforms the current state-of-the-art English model VALL-E and YourTTS on word error rate and achieves new state-of-the-art results on audio style similarity metrics on English and multilingual benchmarks. Voicebox is based on a new approach called Flow Matching and was trained with more than 50,000 hours of recorded speech and transcripts from public domain audiobooks in six languages. Although the Voicebox model or code isn't publicly available due to potential risks of misuse, Meta AI is sharing audio samples and a research paper detailing their approach and results. Voicebox could usher in a new era of generative AI for speech with its versatility for tasks such as in-context text-to-speech synthesis, cross-lingual style transfer, speech denoising and editing, and diverse speech sampling.

More in News

All briefings

Read more in AiShorts