OpenAI says it can clone a voice from just 15 seconds of audio
engadget.com
engadget.com
In Dutch for instance, there's 3 ways to pronounce the R, and everyone has a certain combination of when they use which kind.
In the US this is highly determined by location, to the point where the New York Times even built a cool quiz [0] that can guess where you're from pretty accurately based on a few dozen vocab questions. So there's a dataset out there that would allow you to handle most of this variance by simply plugging in a ZIP code for the person you're spoofing!
A lot of these generators rely heavily on the base model and are worse at generating e.g. Dutch accented English and/or Dutch accented Dutch. But it works great for e.g. Californian accented English with unique voice.
This learned to be a Shazam for human voices instead of songs. And just like you can figure out by ear a snippet of music was made on a Yamaha DX7, seems 15 seconds is enough to narrow it down to a reasonably small set of vocoders that can recreate the sample given.
It feels like OpenAI is mostly concerned with developing proofs of the untrustability of every digital medium, creating a convincing case either for doing everything in person again, or necessitating cryptographic signatures on absolutely everything.
Agreed. I'm still waiting for the part of AI where it's supposed to benefit mankind.
So far, it's 10% entrainment, 20% employment elimination, and 70% crypto-grade hype.
The next AI winter can't come soon enough so the tech industry can get back to doing useful things.
How is this not benefiting mankind?
If you're concerned with unequal distribution of the generated wealth then that's a political problem, nothing to do with AI.
Often the same people will cry over-regulation about ML tech rollout potentially being suppressed and controlled via policy, even though that is literally the implication of their position.
Which, to me, makes sense. Once the underlying technology exists, a malicious actor would not think twice before developing tools of deceptionlike those. It makes sense that OpenAI would work on that "in the open" to demonstrate that we now need to be skeptical of audios.
Of course, the road from here to there may be bumpy...
Ultimately you should reason and investigate deeper to hold somewhat accurate opinion on anything.
Most media and news out there already is where the agenda comes first and then it is about compiling cherry picked content from wide variety of content as evidence for that agenda. And you can find content to support any agenda already.
It is a simple algorithm of:
1. We want to prove X.
2. From millions of datapoints we pick the 100 that support X the most.
3. We write an article on it.
4. Most of our readers already agree with X so they will be happy about it. No need to go deeper.
5. Anyone who doesn't agree would probably not read our platform anyway.
No it doesn't, this is ridiculous. All of you guys with your "this is just like the printing press" arguments are so intellectually lazy, think a little more about the problems of scale and speed and the way information spreads in the modern world.
Obviously not everywhere: I don't need a reddit worldnews comment or a recipe for blueberry muffins signed.
Am I alone in thinking that this would not be a bad thing?
Should they then let anyone register their "voice id" to prevent others frok generating similar voices?
What if your voice happens to resemble the voice of a "prominent figure?"
It is not the same damage.
Not to mention, what constitutes a political leader? Just the upper echelon? What about local civil servants? Mayors, cops, judges? Are they gonna have a database of every public or political figure in the country? No they won't. This is absurd.
Barack Obama started out as a "community organizer."
AI can clone someone making a public speech at that level, then store it forever until they become the president.
This generation is making the same can-kicking mistakes as their parents and nobody wants to admit it.
Society is not ready for all of these "intelligences", heck, we can't even figure out what to do about all the drawbacks of social media and we've had it for decades.
It's being developed because for decades, tech companies have touted the supposed benefits of new technology and hidden the costs. It's been an avalanche effect of greater and greater technology for greater and greater costs. Only now, the costs are so great that people are starting to wake up.
When we look back in time to see the devastation wrough by technology, we won't look at AI as the starting point. It will be the smartphones, the 4G internet, and the 8K video that we so blindly accepted without ever considering the immense changes to society that they implied. AI may be the start of the end of reasonable society, but it is only the APEX of a phenomenon that has been happening ever since we accepted fossil fuel use without considering the implicates of climate change. It's an entire societal attitude.
Voice cloning has a good benefit to drawback ratio compared to many technologies. Threat actors with enough resources could always clone your voice. Everyone can hire voice actors. Costs more time and money sure, but it was always possible.
Beneficial use cases:
* Accessibility: Give people that lost their voice their voice back.
* Personalized Digital Assistants: More humane voice.
* Voiceovers: Sick content creator can let the AI do the voice, filmmakers can fix mistakes in post and so on and so forth.
* Language Learning: Learn an accent by mimicking how the AI voices yourself in the target accent.
* Audiobook Narration: Make a good audiobook out of a textbook.
* Preservation of Cultural Heritage: Preserve an accent that's about to die out.
* Corporate Training: Believe it or not, a robotic voice sucks to listen to after a while.
* Interactive Gaming: Entire virtual DND campaign based on a prompt with AI generated world, story, and characters drawn and voiced by AI with quality control by AI.
Now compare that to guns, tanks, flamethrowers, chemical weapons, atomic bombs, hydrogen bombs, killer drones and bombers. Yeah uh... death. Made to kill people.
Voice cloning has always been technically possible and so has fake videos but lowering the barrier of entry to voice cloning and deepfakes has very real implications for society that we're not remotely equipped to deal with.
Nothing is believable anymore. Stuff like this will only become more commonplace:
https://edition.cnn.com/2024/02/04/asia/deepfake-cfo-scam-ho...
https://news.cgtn.com/news/2024-03-03/AI-deepfake-scams-tric...
https://www.trendmicro.com/vinfo/es/security/news/cybercrime...
https://www.newyorker.com/science/annals-of-artificial-intel...
They require training data longer than 15 seconds, which could lead the out out to resemble more the actual voice.
I've seen weird behaviors where the AI voice forces a British accent to pronounce certain words which I don't have.
Descript also uses voice synthesis to regenerate edited portions of conversations with a noticeable cut to smoothen the transition, which is pretty useful.
Some more discussion on official post: https://news.ycombinator.com/item?id=39866493
> OpenAI says they see this technology being useful for reading assistance, language translation and helping those who suffer from sudden or degenerative speech conditions.
Of course they would. They are manipulating others by appealing to accessibility. Who can argue with that? Accessibility has become the "but think of the children" appeal of the AI world. Yes, it can help those who suffer from degenetative speech conditions, but it can also deceive and push propaganda to whole new heights.
This sort of manipulation is characteristic of antisocial personality disorder. In fact, if you look up the symptoms of that disorder, technology companies like OpenAI exhibit quite a lot of them...
(1) First, most large companies are essentially tech companies in some way, at least in that they exist to advance technology. Therefore, your observation is good reason to abolish large companies.
(2) (Direct) tech companies often exhibit one symptom that other companies don't, at least in not such a strong way. Tech companies lie more often directly to the public: car companies, drink companies, etc. don't lie about making the world a better place, they just try to appeal to basal desires. Tech companies actually are deluded enough to think they are making the world a better place, and use manipulation to convince us that this is true.
https://www.hollywoodreporter.com/business/business-news/ai-...!
I'd argue voice cloning isn't really crucial to accessibility. I mean, this is just cosmetics. I'd like to hear opinions of people who actually can benefit from this and if it is really such a "game changer" for them.
https://www.aarp.org/home-family/personal-technology/info-20...
But that conditioning is not actually true and perhaps you are exhibiting it: we actually can change things. I can already think of some things we could do:
(1) Form an economic coalition that bans the use of AI, and squeezes out AI. Actually, the general public attitude towards AI is already ambivalent enough so that this might work.
(2) Form a revolution against modern tech-controlled society. This would obviously involve a lot of things, but it could involve mass public activity against AI companies.
Literally what is the point of this tech? To eliminate voice actors? I don't buy for a second their supposed use-cases of accessibility or assisting people with disabilities. We're gonna enable people to fake voices realistically off of 15 seconds of for example a voicemail just so media companies can save a few bucks? This is actually madness.
St Pauls Letter to the Corinthians 10:23