Use deep fake tech to say stuff with your favorite characters
fakeyou.com
fakeyou.com
If I ever end up paralysed and unable to speak, you bet your ass I'm talking like Ned Flanders to everyone.
That's funny. What was the sentence you typed if you don't mind me asking? I had subpar results with 50 cent, chosen randomly, and very NSFW language, but also Obama with "My fellow Americans" was strange.
I wonder how they're training and improving the models. I work on a product that could help them train/track/package/deploy/monitor models effectively. Maybe they're not up to date. There also is an issue with the language.
- simulate tone and timbre
- simulate cadence and mannerisms.
Now I wonder if AI can do this or if this is a perfect example of where AI falls short?
I think the AI is now capable to reply to "how likely this will mean this other thing if you say it this way", that applies to text, pictures, videos, whatever, you name it. Is not intelligent in any way, but still gives meaningful responses back, which is useful for the use cases to which it is applied, search, games, etc. That been said, I have no doubts, that with time, it will reach understanding resolving basically almost any problem... but.. we always have doubts, and you can't just divine the future, look at the James Webb Telescope, we launch it to get some answers, so doesn't matter how intelligent a system may be, we would need more research, and the system will need it too, even if it's an AI, (because its needs to know things, to learn from them, in case that wasn't obvious)
[1] https://15.ai
Funny, I just wrote about this type of encounter with exceptional talented people with similar results, and while I didn't detail it in the response, several of those were from those I met on 4chan from 2005-2015 (I checked out after gamergate as it got toxic and less fun).
I hadn't seen 15.ai but its pretty accurate from what I played with. I wonder if he'd be open to see how he trains his algo (deepthroat) with the data sets he gets. He seems to not care about making money from this, or IP and hates NFTs so this would be a good indication he might be up for it.
Sidenote: Also, did you just out yourself as pony*** (brony)?
In all fairness, if you looked up my Mastodon account I'd be outed in moments :p
It's not that big of a deal anyways. I still consider it less degenerate than the folks who burn a decade of their life working on an SAAS that they hate. At least I got to go to a few cool conventions.
I must be getting old.
Good luck enforcing that. If someone makes a fandom crossover video and the voices are only available from separate services, that’s only good retroactively if it gets popular on YouTube and someone then tattles to the 15.ai C&D team. Even then, that doesn’t take down the mirrors.
https://news.ycombinator.com/item?id=23965787
The code repos used are listed in their credits section, and it looks like a mixture of (customised?) Tacotron2, Glow-TTS, HifGan, and others. Videos are generated using Wav2Lip.
Text-To-Speech (TTS) has improved greatly over the past several years, but there's still a lot of metallic sounds in "pure" TTS implementations. I've started exploring voice style conversion, otherwise known as "voice cloning", and there are some interesting repos out there with decent results. These work differently from TTS, in that you don't type out the text to be spoken, but rather pass in an audio file of what you want the cloned speaker to say, and the system outputs an audio file with the same sounds (words, intonation) but with a different speaker identity.
This may be easier to get the right cadence and emotion in the generated audio, as text doesn't capture proper emotion and intonation. I suspect game character audio will use more of voice-style conversion instead of pure TTS simply to get the right emotional cadence of the lines being delivered.
Some interesting voice style conversion repos (in no order, just a random selection if anyone is interested in exploring):
https://github.com/yl4579/StarGANv2-VC
https://github.com/ebadawy/voice_conversion
https://github.com/RussellSB/tt-vae-gan
https://github.com/auspicious3000/autovc
https://github.com/edresson/yourtts
Papers With Code has interesting repos there as well: https://paperswithcode.com/task/voice-conversion/latest
I wish there was some kind of notation to help generate intended inflection and emphasis. Like <sarcasm> or /s tags
Here is my bojack. What are youuuu doing here is read flat and funny! suck a D dumb S sounds decent though.
https://fakeyou.com/tts/result/TR:n45c3yyjwrg3fqbdcxfpn3xrac...
It would basically involve a two-step approach where the first model extracts text and intonation and the second model synthesises the target voice.
I also read some weird teardown of all of the chips and technology used to synthesize that voice a while back and it was very interesting. I'd like to read it again and archive it.
three weeks later
You would need more than just ten -- you wouldn't want it to read digits, you would probably prefer 'sixteen hours later' over 'one-six hours later'.
but yeah that's the general idea I'm getting at.
https://fakeyou.com/tts/result/TR:twwgqfh2432z2sq1e1k1ek4340...
It could use some smoothing of the AI artifacts (no idea how you’d do that though). It’s like they’re talking over a broken mic.
[1] https://fakeyou.com/tts/result/TR:sxshpsntvje985ymtknpyskr04...
https://fakeyou.com/tts/result/TR:6tjeqs8j4x8324v8jq0bgshvem...
https://fakeyou.com/tts/result/TR:0cbfqd1478ms1rnwg7g9d6r6qs...
It's pretty dumb to suggest that I shouldn't be allowed to create something on my computer privately that I can just imagine on my own, imo.
Now imagine people getting sued for being themselves.
I think there should be, at most, a protection covered by libel laws as written today and nothing more. Otherwise we risk getting into that mess.
Not really. You can't force people to say stuff they don't want to say, and it was always possible (although with a higher barrier to entry) to come up with recordings of people who sound like other real people. Impersonators aren't a new concept.
Just because you have a recording that sounds like person x saying thing y doesn't mean that person x actually said thing y and, critically, it never did. Nothing has changed here.
> manipulating my face and voice to say what I did not and never would say is a misuse of my property
Your appearance and the way that you sound a) aren't property, and b) aren't yours.
> We'll be happy to remove any of the voices featured here for any reason.
Maybe they've requested to be removed from the site. Impossible to know though, and they are unlikely to acknowledge it if you ask.
Alan Rickman: https://fakeyou.com/tts/result/TR:hsrgb9haeqff63s966e42m7dnq...
Bernie Sanders: https://fakeyou.com/tts/result/TR:e7b02rgxqzrmfkavrn0pr39jsg...
Snoop Dogg: https://fakeyou.com/tts/result/TR:kn2yaam78wq03d2w53xfd0kcvq...
Morgan Freeman: https://fakeyou.com/tts/result/TR:msxvghpkzs7942vkdfhmxjg2bs...
Which characters seem to work well?
This is the perfect application for this technology.
https://fakeyou.com/tts/result/TR:ev54r4q5b1txx28k32e1s5nw6t...
But really it seems like fair use IMO. these short clips seem very benign.