Nothing to see here. Move along.

"Microsoft researchers announced a new text-to-speech AI model called VALL-E that can closely simulate a person's voice when given a three-second audio sample. Once it learns a specific voice, VALL-E can synthesize audio of that person saying anything—and do it in a way that attempts to preserve the speaker's emotional tone."

arstechnica.com/information-te

@kimzetter I was really impressed by the ethics statement at the end. Loosely transcribed:

"This tool can easily be used to forge audio of anyone saying anything objectionable that you want, but do not expect us to take any responsibility for giving you this tool."

Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.