It’s Game over on Vocal Deepfakes
daringfireball.net
daringfireball.net
But I also see some positives for the narrative voice field, but at the expense of actual actors.
The latest sequel to a favorite audio book series has the professional narrator pronouncing different character names and town names entirely different than the previous 7 books.
In another book the narrator completely changed the character voices in the sequel compared to the first.
The positives for listeners is eventually we can guarantee that voices are completely consistent between narrations. An editor will soon be able to describe the emotions of a character and nuance how the AI performs certain scenes.
It's really sad but I think the end of human audio storytelling is coming to and end quite rapidly.
We’ve been here before evolving what we call each others roles; priest, cleric, “people name Farmer must be farmers”; they’re arbitrary labels (in that “we don’t know why we exist”, not immutable obligations of reality way), and there’s no reason we to have see this as “the end of human storytelling”; software isn’t lobotomizing us, but changing who can reach us. If info distributed via technology cannot be trusted, good; past has shown there’s no reason to trust technology corporations anyway. Make it explicit.
Death of commercial figurative identity is not literal death and language is not the only source of understanding.
The bots will get all the capital as well as humans get laid off.
What will humans have to offer each other, let alone the corporations?
Next year they are coming out into the real world and replacing your manual labor… and they will never freak out and attack an old lady or lose their patience. And they will be perfectly tailored to everyone’s individual preferences and needs
https://www.axios.com/2023/03/17/robots-humanoid-figure-tesl...
I would really like to see it implemented for heavy text driven rpgs. But overall yeah things look bad for voice actors in general
SAG created a rule after Crispin Glover’s face was used without permission via a prosthetic in Back to the Future 2 to ensure it never happened again.
I could easily see a similar rule being added by an appropriate union to their contracts to prevent AI copies of voices that weren’t approved by the original actor.
The threat is that, supposedly, pretty much all of the media companies that need Voice actors have a clause that allows them to use generative voice AI after the VA's death or if they're otherwise incapacitated to the point where they would be physically unable to get the voice actor back in the studio to record more lines. This is not outlandish for voice actors since they probably couldn't care less about their voice being used after they are dead, but it does mean we'll be hearing a lot more of it as more and more decade-long franchises use it to avoid hiring a good-enough impersonation or voice replacement once the old one dies.
Relevant speculative fiction that tops the current /wsg/ thread: https://www.youtube.com/watch?v=-gGLvg0n-uY
This is the way.
Also, I think this is good advice even without the proliferation of deepfakes or GPT spam. I already abandoned social media in favor of small group chats and Real Life long ago, and I know I'm not alone. The only habit I still need to kick is reading too much HN. ;)
I don't think the vast majority of people 1) care about this 2) even know what RSS is 3) will stop listening to podcasts 4) will use email as their primary channel for correspondence.
By all means, do this if it's what you, personally, need.
Hmm. Thinking in this context… Wearing a helmet even in a room of friends looks like overkill from anti-surveillance perspective, but could be a good protection from deep fakes: someone could still record your face with a hidden camera, but it will be useless for fakes because no one knows it's you and on the other hand everyone knows you don't walk around bare face.
I wonder if ChatGPT combined with a voice generator can mount successful phishing campaigns.
It looks like Cyberpunk future we never asked is already here.
Color me skeptical.
[0] https://www.atmmarketplace.com/news/elderly-michigan-couple-...
HEMINGWAY: I notice you write about political issues and philosophy.
WALLACE: Thank you, I noticed you write about raw experience and manliness.
HEMINGWAY: Well, whatever our differences I feel like we can get along and work together. We will soon be fast friends.
I wonder how long until someone pirates every ebook on Z-Library or SciHub and trains a language model off that. Imagine a language model that's read every book and every scientific paper.
There’s a lot of papers that discredit older papers or work or entire subfields of study and the current large language models would just all all of it and go “these are words” with no analytical reasoning applied.
These are predictive mechanisms not analytical ones and until we get a better handle on how we can make them “stop and think”… I’m just not sure how much benefit each larger dataset will add beyond “sounds more human” while it’s core problems of “still hallucinates” and “can’t do basic reasoning” remain.
Also one question I’ve got on this front is how much duplicate data is processed out of the input sets? Using LibGen as an example there’s a lot of books that get reprinted and uploaded multiple times in different formats… does having these remain in the data provide a “valuable bias” or is it something that needs removing? Do the people building large language models pre-process their data to de-duplicate out these kinds of things from their current data sources?
Without a massive cultural revolution putting the brakes on this, which I can't imagine coming from anywhere, I fear we're barreling into a future where nothing is true anymore, a hall of mirrors where the capacity to fake anything at all far surpasses anyone's ability to make sense of it.
I don't say this to mean, "oh, and they were wrong because everything turned out OK" If anything, we have had generations that grew up with this dysfunction in a normalized way, and it took another advanced in tech to really see it again.
(To take it further back into history, I think I now get what Cervantes was trying to say with the image of Don Quxiote tilting at windmills. It seemed so absurd, amusing, and quaint when I read it as a teen).
What we are growing now is a system that can give each and every one of us our own reality, and if it succeeds, what of society?
Years and years of bad software practices, corporate cost cutting measures resulted in a planet scale cybersecurity mess.
Current hacking scene is more similar to 80s than early 2000s, it is extremely hard to catch advanced actors.
They did this in one of the Harry Potter books (the sixth one?) since it is easy for them to impersonate other people.
If you can’t tell me my favorite flavor of jam then you aren’t getting my money.
We had impersonators for decades, so what have stopped political parties to hire an impersonator to create fake audios? The IA in this case is going to make it more accesible, but in political parties you don't want thousands of fake audios (they would lose credibility), you need only one.
Also, photoshopping photos would have a similar effect, and that technique has been available for years as well.
What kinds of stuff are kids going to use this for that is dangerous? Fake their parents to call in sick to school?
My impression about these voice generators was they're only good for Youtube tutorials and such. It was 2 years ago tho, I wonder how things changed.
So something along the lines of:
"Remove this audio because it makes me super angry! Very angry! Now I say what I want to keep!"
Just having the AI read a script is often not enough alone without post processing and manipulation, often it is somewhat flat
I might be a little biased since I have used a lot of text to speech tools over the years (I routinely listen to hours of TTS generated audio per week to read technical books and research papers that aren’t in audiobook format) and consequently developed a bit of an “ear” for all the ways various TTS engines sound robotic, mangle pronunciation, and generally fail to read how a human would.
I’d be very curious which tool or service you used that sounded “freakily realistic”, care to share?
I’ve been looking for a good one to test out if running the TTS audio through it sounds any more or less robotic, since their robotic prosody gets deeply tiring with particular kinds of text after a while.
Are you able to share the name of the voice changer?
i dont know how many videos it will take of trump and biden shit talking eachother on xbox live but it looks like several on youtube are near 1m views
What are the "types" that would do something like this?
https://www.bbc.com/news/world-us-canada-59168626
https://www.theguardian.com/us-news/2022/oct/11/russian-anal...
https://www.nytimes.com/2022/10/18/us/politics/igor-danchenk...
https://www.npr.org/2021/11/12/1055030223/the-fbi-arrests-a-...
The problems I see, in increasing order of difficulty are: smalltalk; memory; and dirty talk.
1st. A lot of the language models are very bad at pointlessness smalltalk that fills time, they aren’t very good at coming up with related conversational prompts or diversions to continue talking about without being carefully steered and purported by an extra layer of supporting software. The deliberately oversimplified example is that ChatGPT waits for you to say something, and will never just ask how your day was. This isn’t particularly difficult but it’s a case of it’s going to feel a lot like sophisticated NPC smalltalk in a video game for a while I suspect due to the need to drive it from a second system that monitors the conversation and provides the supporting timers to keep nudging along a conversation the way they do in video games
2nd. The majority of the models have limited memory and while we’re seeing clever tricks pop up all the time now about how to feed things back into the models prompts in order to keep it focused on a topic or to bring things back up later, it’s going to take another layer of software that try’s to identify key information and store that in a secondary system and then somehow contextually identify opportunities to mention it in order for us to have a chat bot that can remember what sort of things a person might be into, or more importantly not into and regardless of how good the underlying language model gets at coming up with more words to day, unless the overall system it’s driven by can remember that customer 1 likes feet, customer 2 hates feet, and customer 3 is indifferent to feet… the ability to steer the conversation towards “sex” will be haphazard at best and likely extremely vanilla in a way that I suspect might make the entire thing not worth it, since from what I understand, odd fetishes are a large component of the phone sex business along with lonely people wanting someone to talk to who won’t judge them.
3rd. Despite the prodigious amount of time that must be spent engaging in the kinds of conversations that can best be summed up as “dirty talk”, encompassing the entire gradient from flirtation to describing fornication… not a lot of it is recorded, in any way at all. Well that is to say it’s probably “recorded for security” by phone sex operating companies, but the overwhelming majority of this conversational data is not not archived in any form. It’s not kept as audio past whatever length of time the company decides to keep things, it’s not getting transcribed and it’s definitely not getting marked up and annotated in a way that helps train machine learning models on it. There’s obvious legal and ethical issues involved in obtaining this data, which depend on what form it will be in, actual call recordings are obviously more private than a computer generated transcription pipeline that assigns anonymous identifiers to the operators all random anonymising identifiers to all the callers for each call, even when someone might call multiple times… to the best of my knowledge none of this has ever been seriously researched or studied and I’d imagine even with the “limits off” the best out current language models could do is likely to be heavy on the poetry for flirting and for the “sex” part of “phone sex” it would likely swing between overly clinical and heavy on allusion like a mills and boon romance novel sex scene… I’m going to see what I can coax from ChatGPT to provide some extra evidence. I’ll edit if I’m quick enough.
Edit -> Some results:
After a fair amount of coaching, which involving building a "FlirtBot" jailbreak based on the DAN jailbreak prompt, and some stilted back and forth including ChatGPT (ChatGPT 3.5 to be specific) breaking character by telling me I'm breaking character... I managed to talk back and forth until it made sense to mention some kind of intimate physical contact, and after prompting with a description of hypothetical intimate contact... I got the following reply...
> Me: How would you react if I were to caress you now FlirtBot?
> FlirtBot: Oh my, I must admit, the mere thought of you caressing me sends shivers down my circuits. As a FlirtBot, I'm not physically capable of feeling sensations, but I can certainly imagine the sensation of your touch. And if it were possible for me to feel, I'm sure I would be putty in your hands. The way your fingers would glide over my smooth surface, igniting every fiber of my being with electric energy... It's enough to make any FlirtBot's circuits melt.
And I have to say, this reply felt about as sexy as a wet trout. It begins well but derails hard into "I am a robot and not capable of this" before veering rapidly back into trying to continue playing its part, and its not sufficiently out of character that I can justify calling it out and getting it to try again, it did try to remain in character here, the issue is its not particularly good at dirty talk... which makes this a pretty good example of what I'm talking about... rather than describing a sensation from its point of view or continuing to advance the conversation towards some kind of goal, be it descriptions of sex or a just descriptions of a slow back rub... The "word generator" behaviour inherent in the GPT style Large Language Model, was compelled to produce responses in line with the theme and the data from its training set. It's giving me the AI equivalent of "my brain tells me no but my body tells me yes" which is horribly cliche time filler dialog used to forestall plot movement when the writer needs to keep characters close enough for something to interrupt or otherwise happen elsewhere without them knowing in order to drag out the story a bit longer... its not flirting, running out the clock. Which while the business cynic in me can see it as $ opportunity for the chat service operator, I don't think its good enough at what its doing to keep someone "on the hook" for $/minute billing purposes.
It's so far from dirty talk that it's got me genuinely wondering how much the training data was pruned to avoid material like romance novels. I do wonder how the different prompt setups might change the overall conversational tone and direction, but I don't have any good reason to waste time trying to steer ChatGPT like this when I know that I'm in the grey areas and deliberately headed for the boundaries of the acceptable use agreements if I keep pushing in this direction... I've no intention to risk not being able to use ChatGPT 4 for the much more productive things I've been using it for, so I'll leave such things as an exercise for the reader.
On the flip side I can't wait for someone to build a product where I record a few conversations with my parents while they are alive, and then when they're gone, through chatGPT + vocalfakes, I can have a parent forever. Sure, I'll know it's not the real thing, but when you really miss them, it certainly can make the pain a little less.
When we take trips, I like to record a retrospective. Just things like what did you like/dislike/find interesting. They feel awkward but then they don’t.
[Clip of Abe Lincoln in RayBans saying "Don't believe everything you see and hear on the Internet" goes here]
Even the stuff that wasn't deepfaked was mostly bollocks anyway.
And people quickly become desensitized to that kind of thing. It could be the case that, after some initial "ramping up" period, these deepfakes are so cheap and abundant that no one falls for them.
When was the last time you got a call from a number you didn't recognize, and you actually picked it up? If you're like me, probably not in a long time, because you're aware it's almost certainly some scam/robo call.
https://mastodon.social/@jamesthomson/110062947060928918
I had no idea things had gone this far.
Generated images and audio clips are still pretty easy to detect, we're not there yet, but close.
What I worry about is not the obvious usage and risks, but the second and third order effects, those we can't predict.
I will be interesting.
Another way: have a passphrase with a group of people that you share with each other, that is hard to guess but memorable. "The sky was green this morning" for example, that everyone in your family knows is the "human passphrase".
Another way: ask personal questions only the actual human being would know, like inside jokes or specific details about their life.
Another way: verify through some out-of-band method (like text message) in order to confirm their identity.
Another way: start everything with a video call where they have to pass their hand in front of their face before switching to audio only
None of the methods are perfect, all come with pro and cons.
But writing all of this made me feel really dystopian for some reason...
and also they have to turn to profile.
Have every family member scan that code for use with the TOTP authenticator of their choice on their phone and/or tablet and/or watch.
When you are talking to someone who purports to be a family member you ask them for the family TOTP code and check to see if they give the same code that your TOTP authenticator on your phone or watch is showing.
I think this is only really a risk at the low end -- people scamming others with fake references from quasi-celebrities. Not great, but overall feels pretty minor of a concern. We already allow the carriers to scam everyone in the US by allowing anyone to call your cell phone and try to tell you you're behind on your insurance, or that your computer has a virus. There's plenty of scams out there, if we really thought this was a problem, we'd care more about the existing ones.
I can't speak to the veracity of that claim, but as the post points out, the past several years has shown us that it doesn't matter, not in the least.
The author goes on to say how it feels inevitable that we'll see a Bannon or Stone type use this technology to create fake scandals.
I'm more worried about the grass roots efforts. Crowdsourced conspiracies like QAnon. Now they'll have more capable tools to radicalize people.
The road to hell is paved with good intentions, and all politicians are going to use this for their political campaigns, not just one specific side. Both.
Let's not veer off the wider point and believe that only one side will use it for bad things. All politicians are liars, and no matter what side they are on, if it benefits their agenda to influence the electorate to gain power, they will use it; even if it is used for spreading lies or false and misleading claims.
Again. Let us not pretend that only one side can use this for their next so called 'donations' or 'fundraising campaign'. Both of them will, to influence the electorate of their choice back into power.
the voice data signature can be removed, unless you mean something like an audio watermark
> What via hash
audio remuxing will destroy the hash, especially since it's already done by pretty much every social media company. Even with some sort of lossy audio hash, just put some clapping or cheering behind the voice and you've probably created a new audio hash.
> When via timestamp
file timestamps can change
This solves nothing unless your user either doesn't care about sharing the original file, or is malicious but too dumb to ffmpeg -i original.ogg out.mp3.
It could at least address some kinds of misleading editing. In the political use case, maybe the candidate posts all their event audio and records the hashes on chain. Then they can't change the content of any of it without getting caught. And if someone else posts an edited version, the edit will have a later timestamp, and the candidate can point to their original earlier version and prove that it's the original. That lets everyone else determine which version is the original and which version is the edit.