The contrast in quality between this and Apple’s new book narration voices is astounding. I mean, it’s great for fraud and deepfakes that they can train new voices on such tiny samples, but the results seem uncanny in a bad way.
What on earth makes them think minimal training is the right goal, rather than overall quality?
Fair point! While my mind ran more toward “My voice is my passport, verify me,” the idea of recreating lost voices from brief recordings does seem like a reasonable goal.
@pwinn For sure, whether it outweighs the security and privacy (and whatever else) risks ... remains to be seen.
I think we've seen with other things that broadly fall under the umbrella of "technology" that a strategy of simply keeping it out of the hands of bad folks doesn't really work.
tl;dr: You ain't wrong either! 😬
@pwinn I mentioned in another reply thread here but I can see at least one potential reason for such a goal -- voice prosthesis. Most people don't have hours of their voice committed to a recording somewhere. Being able to train an AI with a few moments from an old voicemail could give people who would never hear their own voice again a second shot.
I'm sure there are myriad other similar applications. Key will be limiting its use.