Faux Rogan
fakejoerogan.com
fakejoerogan.com
I voted on each clip without listening to other clips first. I misidentified two that were real as being fake. Wonder if those might have been clips where he was reading from a script with a flatter voice than usual like another top-level comment ITT talks about.
If audio of this quality was being presented in a context where I didn’t know it might be fake, I probably would have chalked it up to issues with the recording mainly, and would have believed that the audio was real.
I suspect that question strikes to things like gaining the tools to transpose speaker affect independent from voice characteristics etc. Imagine a remix culture where a future popular podcast voice is a mashup of different older popular speakers and things like that ...
The 3 Dessa team members did not in 3 months of work create anything innovative probably. Rayhane Mamah, one of the Dessa team members, had previously published a Tacotron 2 (Google's 2017 research) implementation (https://github.com/Rayhane-mamah/Tacotron-2) that has similar noise/distortion and intonation/prosody issues as their "RealTalk model".
Following on the above, Google's TTS research already demonstrated human-parity as measured by MOS score in early 2018. That research was deployed as Google Duplex in mid 2018.
Google's TTS research also showed the deficiencies of this technology. Without the invention of AGI, the TTS models do not understand the underlying text; therefore, it'll be unable to do more "complex things with intonation/prosody". Furthermore, the models suffer from overfitting. The model performance degrades significantly when performing TTS on text not typically seen in the training data.
However, I call shenanigans. Some of the real clips were him reading from a script for the advertisements at the beginning of his show. Joe has a really flat/unnatural delivery when he's reading from a piece of paper.
In addition, if they were including the advertisement reading in the training material, that would definitely mess up the final model. Advertisement Joe doesn't like Joe.
Someone should test this on criminal psychopaths and see if they are significantly better or worse at it than the general pop. I would bet there will be two major populations of psychopath in solidly each category.
Yup, it's fairly obvious there is a difference between a somewhat more robotic intonation vs someone who puts emphasis on certain parts of each word. Still, they nailed the sound/uniqueness of his voice pretty well.
Somewhat in my defense, although I've seen a few clips with Rogan, I'm not very familiar with him. I agree with the GP's comment, I thought the real Joe Rogan reading ads sounded very staccato and robotic. Near the end I got a little better with the "OK, robotic sounding Joe is the real one."
That said, I'm an introvert and I don't enjoy social situations with people I don't know, so while perhaps not a psychopath it's plausible the my poor ability to infer emotional intonation is at play both here and in social situations.
However, as I mentioned in this comment[1] I don't think you have much to worry about. The correct criteria for diagnosing fake vs real is actually not being robotic. I knew immediately what was real because there was emphasis and affectation, as opposed to the monotone counterpart.
Finally, a minor quib: "psychopathy" is not about an ability to infer intonation (quite the opposite). I'm an introvert like you (Hell, I'm probably considered a hermit) and I was easily able to identify the doppelganger.
What does introversion and being a hermit have to do with psychopathy?
I also don't want to make people feel like they're necessarily terrible people if they can't. Sorry if that's the way it came across.
Thinkint about this: It seems that culturally we are comfortable telling people feel like they're dumb (ie SATs, grades, uni admissions) or unathletic (competitive sports is a big part of growing up)... But we don't say anything about moral character.
I never had a coach tell me I am not athletic or a teacher tell me I am dumb.
Explained musically: the fake one goes from C to C# instantly, and the real one sweeps through all the frequencies in between in order to switch tones.
Those psychopaths who are "bad" at determining authenticity versus those who are "good" at it (mostly-wrong as opposed to mostly-correct, respectively)?
I'm hesitant to agree with that; from my understanding, a "psychopath" (current medical definition is Antisocial Personality Disorder, ASPD) is one who is outside of social norms and mores. Much like a precise machine, they are able to attenuate their focus to the nuances most neurotypical people take for granted.
Although far from a medical professional, I would hazard to guess an autistic person would be one of the groups you might attribute to supposed psychopathy (specifically: those unable to diagnose fake from authentic). However, even that claim should be taken with an ounce of salt. Point being, I get triggered when people use psychopath willy-nilly!
I understand people commonly intend empathy to convey 'well intentioned' or 'harboring feelings to understand and help others,' but I thought the actual meaning is the first sentence of your comment?
But again, I shudder to think how technology like this could be used to malign people.
Or rather, would it make maligning people impossible because you would never be able to tell if it's fake or real?
If you haven't listened to it, we just released a longer clip of the RealTalk model[1]. In my opinion it's even more compelling than these clips.
One of fascinating parts of building this has been the questions we received while showing it to people. I'll note a few anecdotes specifically:
"What is the difference between this and a real voice?"
"Can I learn to discern fakes over time?"
"Would we relate differently to a generative voice model posthumously, compared to current media forms like videos?"
These aren't questions that we necessarily have answers to yet, but they're important discussions to have.
[1] https://www.youtube.com/watch?v=DWK_iYBl8cA&feature=youtu.be
I think it is completely irresponsible to advance the state of the art in this field without simultaneously developing techniques to demonstrate that the generated work is artificial.
Please develop validation tests while developing your generative techniques.
The job of the adversarial network is to tell apart real from fake. The job of the neural network is to fool the adverserial network. Both are trained in tandem.
One could imagine training another adverserial network that isn't used to train the network itself, and so will pick up on nuances that the original adversary doesn't pick up on. Anyone could do that, I don't think it's the author's responsibility.
Somewhat related:
https://keenlab.tencent.com/en/2019/03/29/Tencent-Keen-Secur...
Generative-adversarial models have had a lot of success in image generation; however, the same cannot be said for speech synthesis.
Unless they have figured out a new technique, they are probably using Tacotron 2 (https://ai.googleblog.com/2017/12/tacotron-2-generating-huma...). Google's Tacotron 2 already achieved human-parity TTS without adversarial training as measured by MOS.
I imagine eventually, you could have some type of transcript that's annotated with speech synthesis markup language (SSML). Then, you have a CI/CD pipeline that would run this text-to-speech engine and regenerate the audio. I could then pair this up with the video. I honestly wonder if we are a year or two away from this being possible.
> I honestly wonder if we are a year or two away from this being possible.
We've launched a service for adding voiceovers 2 months ago... https://wellsaidlabs.com/
Techcrunch: https://techcrunch.com/2019/03/07/wellsaid-aims-to-make-natu...
GeekWire: https://www.geekwire.com/2019/ai2s-incubator-gives-birth-wel...
> How can I do this with my voice?
To prevent abuse of our technology, we need to review your use-case before creating you a custom voice.
> The google WaveNet stuff is pretty good but still not there yet [1].
Google WaveNet is not built for high-quality voice-overs but rather for cheap and fast text-to-speech.
Seems like an ideal goal for them might be to have audible.com aquire them for a boatload of money.
Some of the British voices are close to perfect. If you want to listen at 2x then it's often better than a human because at speed, there is no value in dramatic intonation, instead impeccably consistent pronunciation and pace are what helps intake which would otherwise be grating at standard speed.
What is really nice is that when you start feel "beaten up" by a high pace reading, you can switch voices and brighten up the experience. I find switching back and forth between male and female voices helps avoid ear fatigue and keep up the pace. I immediately become more receptive to the content when a new voice begins.
I would expect that pretty soon, we'll start seeing a lot of fake online/dating profiles with generated pictures and text. Probably a lot of chatbots posing as attractive women trying to scam people. Seems it should also be possible to condition a GAN to generate pictures that look like someone specific. These bots could try to pass themselves off as people you know.
People also seem to be wanting to get more drugged up, plastic, and bio-engineered.
Maybe in the future the average person is the meeting somewhere in the middle.
I'm worried we're going the way of Japan and its 'herbivore' culture. A world where people have given up on dating because it's too complicated. Other human beings are uncomfortable to interact with, so everyone retreats into their own little bubbles and gets more and more of their needs met by machines... On an individual basis, it a seductive idea to have an idealized robotic partner that does exactly what you want when you want it and doesn't have any needs of its own, will never leave you, but what will that do to humanity as a whole?
There's probably virtually infinitely many (just as life has infinitely many ways to kill an individual or group in pre-modern times).
Also, just as in prehistory, it would have taken visionary and moral heroes to build for their future's survival. But today, the role of those who have those concerns is not important in civilization anymore...
I cannot say what overall effect it will have just yet, but law enforcement will definitely have a harder time collecting evidence.
EDIT: Still impressive tho. I imagine the typical use case for this sort of tech is not just trying to fool people who have been explicitly forewarned.
Does anyone in the field know more?
More likely is that you're helping them write the results section of their research paper, with a sentence like this, "In a blind trial of 5,000 internet users, over 93% of people were unable to tell the difference between generated audio and real audio at a statistically significant level (P<=0.035)."
Note: if we assume a binomial distribution, you need to get 7 or 8 correct to reach the magical P<=.05 barrier. If you assume that there are a fixed 4 generated clips and 4 real clips, then you need to get them all right (https://en.wikipedia.org/wiki/Fisher%27s_exact_test). I think it's fair to use the binomial distribution because the website does not tell you the number of real and generated clips before you take the test.
However - I guess the answer correctly with all but one - “You are much less likely to injure yourself if you do it correctly” - which I can now hear the different (or perhaps confirmation bias).
There’s something artificial about the timing with the pauses between sentences on the ML based ones that was the main giveaway for me, also while his voice is damn near spot on - there are some vocal inflections or perhaps “emotion” in a few words that also hinted to me.
Hope that helps.
I think the real tell is that Real Rogan has a lot of variance, Faux Rogan sounds too much like the average Rogan.
There's already businesses that sell camera apps (or cameras? my memory is a bit dim) that save a photo along with a cryptographic hash to prove authenticity. Their customers are for example insurance companies, which require their clients to take pictures of damaged property etc. for claim filings that way.
And if the fake confirms someone's worldview or serves their interests, confirmation bias will take over and no "fancy math" is going to change their mind.
What I'm hoping is that we'll find means to actually build trust, e.g. with signing as we discussed here. I'm not entirely confident in that, but wouldn't it be nice if we engineers had a hand in building some useful tools for society? :)
There have been instances of of deceit, but in general the ability to give some credence to leaked political tapes has been a positive thing for society.
With this avenue for political exposure disappearing and totalitarianism making a comeback, it is a bit worrisome how we will continue to blow the whistle on politicians when being present with a camera isn't even enough.
> Photo manipulation is a very old practice.
But not modern audio or video. We're running out of options.
> There's already businesses that sell camera apps (or cameras? my memory is a bit dim) that save a photo along with a cryptographic hash to prove authenticity
This only works if the photo is of yourself and you want to prove its authenticity. If the photo is of someone else, it's still your word against theirs regardless of what hashes you have.
Yeah but would you figure it out if you just listened to it expecting Joe Rogan?
https://github.com/syang1993/gst-tacotron
I haven't tried it yet, but if you're looking to do something similar, this repo (and the papers on which the implementation is based) might be a good starting point.
I wonder how long it would take until Joe brings those guys to his podcast that is something I’ll be waiting for.
If it weren't for that, they sound pretty natural, so it would be difficult.
I only got a 4 out of 8.
I’m assuming the longer the audio clips the easier it would be to detect the fake/AI?
Unfortunately the UI/UX just isn't there yet
Who are you, my father? Am I the only one who finds this trend really cringey?