Total Moving Face Reconstruction
grail.cs.washington.edu
grail.cs.washington.edu
It's terrifying to think that in the next 5 to 10 years we won't be able to distinguish a forged, high definition video of pretty much anybody.
In 5 to 10 years, hopefully we will have learned to never ever ever take anything we see as fact, because we absolutely will not be able to distinguish rendered video from the real thing.
Indeed, as there are more and more cameras around (including autonomous ones of increasingly tiny size) imagery of the videographer will probably become a major authentication factor.
However, I'll still be optimistic and hope that with the increasing number of cameras, people will be less likely to engage in activities where they shouldn't. The opposite argument would be that with the increasing technology to fake such an activity, the amount of 'my twitter account was hacked' incidents will rise.
I guess only time will tell?
I can see this being used primarily against those in power though, so you aren't wrong.
A deeply closeted gay man (with a wife and kids) taking their first steps out of the closet and kissing a man could ruin his career with those things. Is that OK?
Consensus is not sufficient. This is why rights-based approaches are developed. It has to be OK to live an unpopular lifestyle that's not harmful to others.
I'm not suggesting a mechanism that operates on the 51% rule. I'm suggesting we 'elect' certain people to serve in roles that allow quick consensus to be formed when there is disagreement. Put those judgments in a safe place and you have 'laws' that can be referenced in the future.
In a world where entirely realistic body images can be created of anyone, it becomes possible to threaten to "leak nudes" of anyone, even if they've never taken any.
It is also the world where nude photos of anyone have exactly zero value, because everyone who wants them can generate those themselves and there's no easy way to distinguish between real, doctored and fake ones.
- We have mechanisms to prove that video is taken after a certain time (show a newspaper on the video, or get the video cryptographically timestamped via a trusted server).
- We have mechanisms to detect if something is taken before a certain time, by doing something interactive and unlikely with live viewers.
- We have mechanisms to detect if something has been modified from its original form (signing).
You might be able to make a CCD chip that signs every frame with a private key, and then ships the frame off to a public signing server too. Producing that CCD along with the video taken might be proof. But then you could defeat that with a display hooked to the camera, feeding the doctored image to the trusted camera.
I remember learning about Stalin's photo retouching, and reading about the systematized photo retouching and censorship in 1984. At the time I thought it was completely impractical, there's just too much work to do and not enough people to do it. Let's hear it for automation, putting the auto in Autocracy! :)
How does that defeat the cryptographic timestamping?
But you still wouldn't be able to guarantee that that actually happened.. if we have enough compute and good enough algorithms, you could spoof a scene in real time on a small display strapped to the front of the trusted camera. Even though the trusted camera is recording what it sees, what it sees isn't really what is happening.
So we still don't get back to being able to use video as evidence.
What did surprise me was the quality of the renderings, according to this article, not even their own QA department can tell the difference between their photos and renderings any more: http://www.cgsociety.org/index.php/CGSFeatures/CGSFeatureSpe... (so they put their photographers on a 3D modelling course, and the 3D modellers on a photography course, to blur the distinction even further--really cool article, IMO).
3D Reconstruction of human faces is literally on the edge of mainstream. I'm betting on it, personally.
Our system is similar as theirs, but more general: we laser scanned 300,000 real people and then associated each laser scan with dozens of photos of that person taken from different angles, lighting conditions and expressions. That data set was then used for a neural net training - actually a pipeline of neural nets.
We can accept 1 photo and get back a good quality 3D model, or a series of photos and get better quality, or HD quality video and get back frame by frame, in expression reconstructions just like their solution. In fact, our system is able to recover 36 people per video feed in real time, as well as handle 4 video feeds at once. We don't need as much reference information as they do, because we trained our system to generally understand the human facial form, rather than their solution that operates in isolation for a single reconstruction operation.
Our current system is targeted as a WebAPI for games and serious simulations - enabling 3rd parties to implement "put yourself in the game" functionality. As such we have 3 different geometry outputs aimed at game/simulation developers. We also do facial recognition, and we have a special "forensic" output for that.
Our current "best output" is purposely "Pixar like" rather than realistic. Taking them realistic tends to freak people out - especially women (seems like our culture has trained women to have an idealized self image, and when presented with their non-mirror true form, they don't like it.)
You can learn more at these links: https://3d-avatar-store.com/Web-API-Features-May-2014 https://3d-avatar-store.com/3D-Avatar-Creation-walkthru https://3d-avatar-store.com/New-Face-Finder
Your work is very nice as well. Like yours, our video version requires no manual interaction. It's primarily used by government agencies, and we've not exposed it to the public yet.
I have never really seen CG portrayal of people in live action films that didn't break suspension of disbelief, it always falls right into uncanny valley stiffness. The exception being Avatar, which for whatever reason doesn't seem to have had the problems other films have.
By not CGI-ing humans. That's enough for you to stop judging what you're seeing by the same criteria as you do real life. They wisely chose to stay well to the left-side of uncanny valley instead of trying to jump over to the right-side.
Look it up. It's also why Wall-E feels so human.
Here's some super impressive CG from the first Captain America movie back in 2011. Crazy insane stuff they're doing it. And that's for a very, very extreme case. You better believe they are doing slick stuff in non-extreme cases! http://www.fxguide.com/featured/case-study-how-to-make-a-cap...
For Iron Man it's not quite the same because it's a suit but it's just as impressive. Since Iron Man 2 there is no full suit worn by an actor. There are no legs and at this point there are barely even arms. There's a chest piece and open mask face piece and that's about it. http://movies.stackexchange.com/questions/2198/how-are-the-i...
In The Social Network there are a lot of scenes with the Winklevoss twins. Spoiler: They didn't use twin actors. They used two actors and CG'd the face of one onto the other body. No way in hell did you call this out on first viewing. (number 3) http://www.totalfilm.com/features/50-cgi-scenes-you-didn-t-n...
"The last movie I remember having a CG face was Clu" By strict definition if you remember it having a CG face then it wasn't good. It's not possible for you to remember a good CG face because if it's good you wouldn't know it was CG!
The next big test might by Paul Walker in Fast and Furious 7 or Philip Seymore Hoffman in Hunger Games 3. Both big budget movies where the actor died and CG will be used in at least some places. That's the ultimate. I think we can already get away with CG if the actor isn't known. But for actors who have faces we know and mannerisms we subconsciously recogonize it's another step up in difficulty.
Example?
The paper claims that it takes about 105 seconds to render a single frame. So one second of 30 fps video would take about 52 minutes to render. I would have to read more in depth to see what kind of savings can be had by sharing information across frames. (The paper also doesn't mention the use of GPU acceleration.)
This makes it seem like the mesh produced for each frame is specially deformed to render optimally for that camera angle, and it wouldn't necessarily look perfect from other angles.
I love the attached video in the link because it isolates perfectly the face. If you look closely you can see these tiny minute combinations within the face as each person talks; the eyes shifting, the face rotating, looking in various directions, the forehead crunching, the eyebrows raising, smiling, etc. All of these "cues" combine to create a message that we interpret instantly.
The face has inspired me lately to read more into this subject as it seems, at least on the surface, to be an extremely complex innate human ability; facial recognition.
Billions of people in this world, we all have a very similar facial structure with two eyes, a nose and a mouth, and yet you can recognize that small blurry face in a fraction of a second.
I imagine it's similar with animals. You can have thousands of birds in a flock, and they can recognize their mate instantly. To us, we'd have to carefully analyze the birds for days or weeks to make that same match.
This is nonsense, citation needed.
> For example, You have a one special neuron in your brain that fires when you see Bill Clinton's face.
This is New Age science, i.e. not science at all. It's a myth. One neuron is equivalent to one bit in a computer. One bit of information is insufficient to distinguish between faces.
> And that exact neuron fires no matter what picture of Bill Clinton you happened to see.
Maybe you could learn a tiny bit of neuroscience before spreading this kind of nonsense.
http://www.klab.caltech.edu/koch/
Sorry my post evoked such a harsh response from you.
"Seen <particular face>" is a binary state, so one output neuron that fires on recognizing a particular face is actually quite reasonable.
Obviously, the inputs that contribute to that recognition must be more complicated.
EDIT: As dragonwriter says below, lower levels of neural networks are responsible for general facial recognition but that triggers more specific neurons once the recognition is done from generic face -> specific person. Even in artificial neural networks single bit is sufficient at the end of the classification.
A neuron is interconnected to many others, but this doesn't mean one neuron is many neurons, any more than one binary bit is 64 binary bits by virtue of its position in a binary word.
In 2005, a UCLA and Caltech study found evidence of different cells that fire in response to particular people, such as Bill Clinton or Jennifer Aniston. A neuron for Halle Berry, for example, might respond "to the concept, the abstract entity, of Halle Berry", and would fire not only for images of Halle Berry, but also to the actual name "Halle Berry".
EDIT: Here is another link:
http://www.newscientist.com/article/dn7567-why-your-brain-ha...
> However, there is no suggestion in that study that only the cell being monitored responded to that concept, nor was it suggested that no other actress would cause that cell to respond (although several other presented images of actresses did not cause it to respond). The researchers believe that they have found evidence for sparseness, rather than for grandmother cells.
And yes, I realize it's even possible now, but with all the new algos and software coming out it will be easy enough for somebody to just mess with people's lives for fun.
I trust that, due to jurisprudence model, the legal framework will always have some lag[1] after game-changing technology becomes mainstream... sometimes this can be significant.
[1] disclaimer: active legislation efforts can move faster than tech - especially due to lobbying.
But this is not to say that our system for handling forensic evidence works great - it doesn't sadly, and part of the problem is that juries are often reduced to trying to weigh the credibility of competing expert testators without rigorous standards of measurement or terminology. You'll probably be interested in this: http://www8.nationalacademies.org/onpinews/newsitem.aspx?Rec...
For those of us who grew up reading Neuromancer, Snow Crash, and other cyberpunk yarns: that is today, we are living in that world at this moment.
The plus side of this is that people will be able to correct casting mistakes that ruin what would be otherwise great films
Keanu Reeves in Speed
Keanu Reeves in Sweet November
Keanu Reeves in The Lake House
Keanu Reeves in Devil's Advocate
Keanu Reeves in everything except Bill & Ted
Or perhaps these same algos will also provide utility in detecting / decoding "fakes", sort of like edge-tracking / error level analysis etc today.
...but still, outstanding work even if it requires manual calibration like curation of the input set.
I think this is what happened to the recent decapitation videos, they were reconstructed from home videos.
IMO the videos with the people dead in the floor are true, but the videos where they talk are staged.
Today we know there was a CIA team whose job was faking videos of Osama Bin laden: http://blog.washingtonpost.com/spy-talk/2010/05/cia_group_ha...
Remember Osama Bin Ladem appeared and disappeared according to US army interest at the time, finally ending in very strange circumstances(and being buried on the ocean, not letting anyone else interantionally to confirm(by DNA) he was Osama).
For me it is staged because current technology could synthesize a voice only if there are not strong emotions. The same happens with the voice.
With strong emotions it becomes very easy for familiars and friends to notice as people do specific gestures and most of them are not recorded in video.
That people are perfectly calm before dying I could understand, but that they do while saying exactly what their captors want I can't.
Also, before the videos most of the population in UK did not want to go to war, after the videos(with a UK native), most of them support war, quo prodis?