Seriously impressive #AI text to speech from Microsoft.
"VALL-E is a transformer-based TTS model that can generate speech in any voice after only hearing a three-second sample of that voice. This is a significant improvement over previous models, which required a much longer training period in order to generate a new voice."
VALL-E: Microsoft’s new zero-shot text-to-speech model can duplicate everyone’s voice in three seconds
| https://mpost.io/vall-e-microsofts-new-zero-shot-text-to-speech-model-can-duplicate-everyones-voice-in-three-seconds/
The contrast in quality between this and Apple’s new book narration voices is astounding. I mean, it’s great for fraud and deepfakes that they can train new voices on such tiny samples, but the results seem uncanny in a bad way.
What on earth makes them think minimal training is the right goal, rather than overall quality?
@pwinn For sure, whether it outweighs the security and privacy (and whatever else) risks ... remains to be seen.
I think we've seen with other things that broadly fall under the umbrella of "technology" that a strategy of simply keeping it out of the hands of bad folks doesn't really work.
tl;dr: You ain't wrong either! 😬