Nvidia Vid2vid: High-resolution photorealistic video-to-video translation
github.com
github.com
From a technical standpoint I think this is very impressive, and I'm also interested in creative/artsy use of this. Their "replace trees by houses" example is pretty dull but gives a good glimpse at what can be done.
Around March/April this year, I actually did download the TensorFlow toolkit for face transfer that was used by /r/deepfakes people and tested it out (there were samples of photos of politicians included); the results were, at best, worse that I could do in 2 minutes in Gimp. Maybe they could get better if I had an expensive GPU farm at my disposal, but I'm pretty doubtful - given that the news died down pretty quickly, and no reasonable-quality faked pictures or videos were reported ever since.
Yes, Adobe do have some remarkable algorithms that would be difficult to replicate (e.g. heal brush and content aware fill) but these are a small minority of Adobe's software advantage.
The one that irritates me the most is vector drawing programs: open source programs (and even paid competitors) just can't touch Adobe Illustrator for the sort of work I do. I'm sure at least 50 percent of it is familiarity and muscle memory, but I've desperately tried switching to a few different options like Inkscape or Affinity and left wildly disappointed.
I would be interested in a Lightroom alternative if anyone can recommend one though.
In this specific case I think it might be beneficial if you can spend a lot of money on gathering training material and tweaking the network. But then again someone here mentioned the quality 4chans fake porn has reached, so maybe I'm wrong after all.
People are motivated by money, governments pay little and there’s no way to get rich working there so it doesn’t make that much sense.
Around 1960 space flight, supersonic flight.
More recently hypersonic flight, and we don’t really know becase people back in WWII had no idea.
And this sticking to the spreadsheet concept, which is very limiting.
---
Contrast for example Tableau -- it's a great idea and generated a lot of enthusiasm for a while, but never quite took off as an office package one needs to have. The normal awkwardness of its first versions is still there; they don't have the deep <whatever> that the Excel team has.
In comparison, Excel can do that too (just worse), but it can also solve equations, do your company's bookkeeping, and pretty much every other task that relies mostly on numbers.
I would argue Open/LibreOffice Calc comes fairly close to Excel if you ignore the worse user interface (which is fair in the original assertion that it's "mostly polish and small incremental improvements")
considering that's a major part of "better" that's big ask!
Especially for public figures with lots training data available.
There's been r/deepfakes where around 30% of the content was porn with swapped faces. It was banned though.
Fake celebrity porn has been on the internet since 1996 at the very least. It's always been crummy; but porn in general requires a thick suspension of disbelief and an intense focus in a partial object of desire (what Lacan calls the objet petit a) that blurs everything else.
Initially, every Tom, Dick & Harry script kiddie will try to blackmail people whilst the public and justice system is not really aware of the technology possible.
When eventually enough of the public is made aware to not trust photos and videos anymore then these types of blackmail attempts would be less effective. But there will still be less informed gullible targets.
The blackmail angle I expect will work for a while. People falsely accused of sexual crimes can have their lives ruined by the accusation.
But as you say the media (real and social), political propaganda, etc will exploit the uninformed masses for all it is worth, so blackmail will still work. For a while.
This stuff even has the catchy name "deep fake".
That's not all you would need for verification, but it is a big help.
The only way I see it working is if the key is "burnt-in" to in the camera hardware and any applications cannot MitM it.
> the app itself could take a deepfake video and have the OS sign it
Note how I said "it asks the user via the OS to approve the result". I would expect a modal OS dialog to let the user review and approve the content before being signed and passed back to the app.
Thinking about it, there's actually nothing stopping this from happening on today's hardware using just application sandboxing. Substitute "OS" above with "Signing App" that does the same thing (accepts media signature requests from other apps, and opens dialog to request approval from user with a preview).
https://twitter.com/PiratePartyINT/status/104296466807811686...
"Starting now, we cannot trust video or audio evidence. The ramifications for our legal & political systems will not be known for many years"
Publicly available research is usually the cutting edge in most areas of knowledge.
The type of neural network they use (GAN) works by having two networks battling each other, one tries to generate fake, and another (discriminator) tries to identify fake, it's a constant arms race. As the generator gets better, the discriminator also gets better. Which means, if the fake video is this good, there must be a discriminator network that identifies fakes just as good.
We did a similar project using GAN, generating images from a text description. You can see the progression of generator and discriminator battling each other, and both get better with time.
Because the discriminator (D) and generator (G) usually compete in a minimax game, the equilibrium probability of D correctly classifying an image as fake tends to 1/2 (ignoring distributional factors). If the competing networks have enough capacity and can be stably trained, then in theory they will reach equilibrium as the data distribution from G converges to the actual data distribution. If this is the case, then the discriminator correctly identifies fake videos with a probability of 1/2.
They may not reach equilibrium (making D > 0.5), but it's not clear that the discriminator itself is a panacea for identifying fake videos/images.
We'll have systems arguing with each other and no way to tell which is correct. If, for example, someone is able to get a copy of the forensic AI model, they can train their decoder/generator to work against it until the results pass as legitimate. With no human ability to argue with the results of the forensics AI, we'll just trust it and pass it off as truth.
If you can clone the jury and conduct a quadrillion private trials, your chance of success in court is going to increase substantially.
Things gonna get creepy.
As I don't have a DGX1 here, training the 2K resolution net for 10 days on a p3.16xlarge instance (also has 8 V100 GPUs) would cost USD 5875 on AWS. (USD24.48 per hour on-demand pricing * 24 hours/day * 10 days)
V100 has ~50% higher memory bandwidth than 2080Ti, so you probably wouldn't get 80% of the performance. Also, only two 2080Ti can be connected via nvlink.
If you want to train on your own dataset, that price does not seem unreasonable to me.
They could even generate a new "cast" for each market, after only shooting the show once.
- Porn, yeah, first application you can think of, there are already some startups doing it.
- Doubling actors, and applied to sound, maybe you could translate from one language to another but kind of keeping the accent and tone.
- Propaganda and misinformation. Now you can get your enemy to say and do whatever you want, on video.
- Photo-realistic games. Create a rough 3D model of an scenario and train the AI for it. Instead of photo-realistic rendering with math, render it with the AI based on a rough render, in real-time.
According to last month's nvidia rtx presentation/launch event [1], they are going to do something similar quite soon. Games will ship with DNN pre-trained offline on extremely high quality renderings. Game itself renders at lower resolution (limited by performance needed for proper raytracing) and uses DNN to upscale it.
Edit : found the answer on Internet, apparently the RT (raytracing) cores are different and separated from the Tensor (NN) cores on the RTX
https://www.theverge.com/2018/8/21/17763278/deepfake-porn-cu...
This could open the market for video propaganda-detector applications.
Look closely, while it does generate videos of a passing similarity, they aren't "photorealistic" in the slightest. They are good locally across time and space domains, but globally they are as far from realism as Doom 2 was.
The only explanation for the attitude I see in this submission is that most IT people trained themselves to spot CGI by looking at local artifacts, assuming that global artifacts won't happen because stuff on the scene is reasonable. There is no "stuff on the scene" with those videos, it's just mindless vector manipulation with no underlying world model. Cars wave around, trees grow a feet from each other and behave in a way incompatible with 3D perspective.
Relax, it'll require at least another AI/ML revolution (or even several) to achieve photorealism.
Heck with impersonating the POTUS. What about a lost friend, sibling or parent?
Heaven on Earth?
If so, that is amazing.
And if so, how do I turn a video I have into a simple/line version, to be able to then put a different 'skin' on it?
Example with video here https://www.youtube.com/watch?v=1Ndxtb0q76c
Obviously would need some tweaking but could be a good starting point
You have the code right there on Github, just install it on some PC with powerful GPUs (or rent one), tune some parameters, train the network and you can do the same things.
https://www.youtube.com/watch?v=MMbgvXde-YA
of course this being Nvidia they didnt implement it universally, you need to sign up for API access to black box gameworx like scam programs in order to implement it in your game.
gpg --sign video.mp4b) what the latency impact on the NN side to build the images (e.g. how many ms are we talking about?)
Thank a lot!
Can we expect CUDA to be the x86 on PC and Servers? Literally all works are defaulting to CUDA and Nvidia's library. I don't even see a contender trying compete. I don't even see AMD's ROCm being used or even mentioned anywhere.
He would say it's fake, but who would believe him?
How exactly can computer scientists explain deepfakes to laymen?
I tried pointing out to some friends some really bad artifacts in a video we watched, and they just could not grasp it. They couldn't see what I was seeing as it didn't look out of place to them. It isn't for lack of intelligence, they just didn't care enough to understand. That pretty much describes vast swathes of the population.
You show a video using the above technique to anyone with strongly held political/ideologic beliefs and an inclination to accept 'alternative facts' over actual facts and videos using these techniques will be like a wildfire and almost impossible to refute!
The real danger here, if anything, is impersonating common citizens or lower ranking government officials.