Seriously impressive #AI text to speech from Microsoft.

"VALL-E is a transformer-based TTS model that can generate speech in any voice after only hearing a three-second sample of that voice. This is a significant improvement over previous models, which required a much longer training period in order to generate a new voice."

VALL-E: Microsoft’s new zero-shot text-to-speech model can duplicate everyone’s voice in three seconds
| mpost.io/vall-e-microsofts-new

#tts #microsoft

Follow

The contrast in quality between this and Apple’s new book narration voices is astounding. I mean, it’s great for fraud and deepfakes that they can train new voices on such tiny samples, but the results seem uncanny in a bad way.

What on earth makes them think minimal training is the right goal, rather than overall quality?

@andrew

@pwinn I mentioned in another reply thread here but I can see at least one potential reason for such a goal -- voice prosthesis. Most people don't have hours of their voice committed to a recording somewhere. Being able to train an AI with a few moments from an old voicemail could give people who would never hear their own voice again a second shot.

I'm sure there are myriad other similar applications. Key will be limiting its use.

Fair point! While my mind ran more toward “My voice is my passport, verify me,” the idea of recreating lost voices from brief recordings does seem like a reasonable goal.

@andrew

@pwinn For sure, whether it outweighs the security and privacy (and whatever else) risks ... remains to be seen.

I think we've seen with other things that broadly fall under the umbrella of "technology" that a strategy of simply keeping it out of the hands of bad folks doesn't really work.

tl;dr: You ain't wrong either! 😬

Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.