DeepFaceDrawing Generates Photorealistic Portraits from Freehand Sketches
syncedreview.com
syncedreview.com
One of the big problems at the moment is that photorealistic texturing basically requires finding and mapping a real life actor, which is all sorts of expensive.
This looks like a way to get past that in a super-efficient way: just mock up what "look" you're going for with a character, hit render and out pops all the data needed for a texture (I suspect we're not too many iterations way from also getting bodies, meshes and animation skeletons).
This would make it a huge field-leveller for catching up for HD content production.
As a hobby game dev myself I find this to not be a problem at all. I think human skin is largely a solved problem. It largely boils down to creating a proper material/shader, the material will have many layers of textures. These days, with the proper tools, this is easy, and generally more of an artistic affair than technical.
In fact, I will claim that texturing for games, in general, is largely a solved problem.
Look into both Substance suite and Quixel's offerings, watch/read tutorials. Both Algorithmic and Quixel revolutionized texturing and made AAA quality possible for indies, just like Unreal Engine made AAA quality available for indies.
Can you post an example? All the fake skin I’ve seen is still marching into the uncanny valley (video at least)
What an artist does with that implementation is always another matter, and it will always be down to the artist to create a pleasing final result. Specularity is one thing I think artists gets wrong all the time. At the artistic level, games often deliberately go for a more cartoony look, too.
And the uncanny valley thing is very subjective, no?
https://docs.unrealengine.com/en-US/Resources/Showcases/Digi...
I think this is more than good enough for real-time games. Of course, you will probably say, "no that sucks, uncanny valley!"
One thing to notice here is that the albedo (diffuse) map (texture) used can be very crude/simple, as it is simply one small piece of the puzzle.
If you have looked at a video, or even screenshots, like the digital human ones, where the focus is solely on the subject, often close-up, that is very different to see it running in a game, where you are less likely to notice small faults.
These days, I think facial rigs and animation is the cause of uncanny valley, to a higher degree than skin shaders.
Skin (and other stuff) still looks pretty weird to me.
The way how these AIs work is that they have memorized aspects of celebrity photos and then recombine them as needed. That means if your sketch looks in any way like a celebrity from the training data set, then the AI will likely reuse parts of those photos, which would make your generated texture a derived work, meaning you'd have to pay royalties to the celebrities in the dataset.
Previous research has shown that if you search the entire training dataset for the image most similar to one generated by a generative model like this, the images actually look pretty visually different. To a non-expert, they'd say the images are similar, but neither copied the other.
There are two inputs here: a collection of photos and a model. It's not being derived solely from the original photos. The scientist had a contribution as well, which cost him money and time to develop. As long as it's not wholesale copying and there is significant creative input (selection, changes, recombinations, etc) then I think it should be ok.
In legal terms, if your AI model is "sufficiently transformative," you don't need to worry about anything like that. But if your AI model is overfitted and is just memorizing, then yes, you're right.
Generally, a material is made up of _several_ textures that serve different purposes. Common textures (often called 'maps') used to build materials are albedo maps, surface normal maps, metallic/roughness maps, subsurface scattering maps (this one is important for realistic looking human skin!). There's others, and sometimes aesthetics/requirements require shader authors to _make up maps_.
Consider a texture map just like, a precomputed data cache. You can encode pretty much anything in them. Why, in the gritty gore and carrion filled trenches, I've created systems where artists can use maps to annotate parts of models they think suck. That map was used in a camera AI system that tried to avoid looking at parts artists are ashamed of, adjust depth of field threshholds depending on the depth of bad stuff... That kind of thing. (That texture seam is too gross for the sizzle cinematic, there's NO TIME to fix it. Just smudge the camera, or have it lose focus as it sweeps through! shameMap.png to the rescue).
The limit to the kind of data you can _use_ is only your imaaaagination..
For people, for most aesthetics, at minimum you need a color or diffuse or albedo map.
In the image labelled "Illustration of the model’s deep learning framework architecture", the input face has a strange line drawn underneath the chin. It seems like an odd thing for a human drawer to put in, and makes the person look like they have a double chin.
Yet in the output shown at the end of the pipeline, it appears as a shadow. I didn't go into the article suspicious, but this immediately made me wonder if for some of these sketches, a face to line drawing network was used for some sort of reverse process.
The image does appear in a part of the article discussing their learning methods, though, so I'm probably missing something important. But given that they "are working to release their code" it doesn't really help with confidence.
In a few years, deep learning is going to make any sort of development of real skill feel as archaic as assembler. Learn guitar? What's the point. The little magic black box soon-to-be-smaller-than-your-smartphone can make just about any song based on minimal inputs (i.e. a beat-boxed backing track). You'll be able to generate unique and stylized paintings of your relatives and pets in seconds. Probably you'll be able to generate printable 3D objects from descriptions. Engineers will be able to sketch parts from one perspective and have the details automatically fleshed-out from best-practices learned across millions of similar parts.
You'll never get away with an illegal U-turn ever again because the city will pull footage from peoples' internet-of-crap dashcams and the machine learning algorithms will comb the feeds and send fines directly to your mailbox with basically no human intervention.
It's one thing to have endless amounts of texture, music, content etc. But you also need to combine them into a playable game for example where all those different parts need to match. And the game still needs to feel fresh and make fun. Will a computer alone be able to do that? Can a hobbyist do that?
In contrary, I believe tools like this will increase the required skills someone needs to bring with him to make something worthwhile. That's why tools like Game Engines, 3D Modelling Software or Music DAWs get more complex year by year.
Case in point: do you believe a hobbyist will be able to do something like that [1]? No, it will be a team of dozens and dozens of specialists who will raise to bar even higher and higher. But they will profit the most of artificially generated art, which they can manually adjust and provide the final touches, to turn something generic into something great.
This is not a "in a few years" thing in some places: https://youtu.be/taZJblMAuko?t=1536
This is extremely impressive. Extrapolating for possible applications, I could imagine that techniques like this could one day become invaluable tools, say, for asset creators in the game industry. The video speeds up the process, but of course in a couple of years, this will be actual real-time performance. Extended to more than just human portraits, this could be a fantastic design tool.
Your conviction rate will go through the roof when the artist's sketches are a dead ringer for the suspect!
B) even so dmv photos are notoriously bad
No problem! Conviction/case-closure stats don't care if the person you convicted was local or not.
> B) even so dmv photos are notoriously bad
They're often unattractive but they usually identify people pretty well when they're not too old to do so.
For example, take the difference between 2:20 and 2:27 in the video. The upper half of the drawing hasn't changed, but the generated image has a lot more hair and different ears. While the technology looks impressive as it is, it seems to me that it would be better to leave areas the artist has barely defined as blurred rather than flickering between various high resolution features that are all roughly equally matching the sketch.
The whole thing works on statistical priors: if I have feature a at location x, there's a 90% I should have feature b at location y. So if the majority of pictures of beards in my dataset were also, say, wearing sunglasses, then naturally if I freehand draw a beard the net will probably output sunglasses even if I don't change the eyes!
The solution is to ensure that you sample the full data space that you wish to reproduce (not trivial). Neural nets do seem to interpolate but this is super high dimensional space so it's not always intuitive...there are many orders of magnitude more directions in which to move to get from point A to point B.
Looking at the virtual agents, it seems they are able to understand very crappy English (with all my attempts), how far are we from correcting it?
And the magical accent corrector wouldn't fix bad grammar.
I'm reminded of a former flatmate whose father chose not to raise her as bilingual in the mistaken belief a second language would impair her learning. Instead, when she chose to learned Spanish as an adult anyway, she picked up the slang and pronunciation of her Colombian relatives, but never quite reached native fluency. She pointed out the drawback to having a local sounding accent and name instead of being an obvious foreigner was that everybody who met her assumed her misunderstandings, pauses or the odd really ungrammatical phrase was because she was an unusually stupid Colombian.
Aside from the obvious idea of masking one's visual identity for example through VR avatars in lieu of face-to-face interviews (to hide for example gender, physical appearances and able-bodieness), one's voice would still reveal many factors that could be used to purposefully or accidentally skew any neutral position one might have during interviews.
I know for a fact, that if I had the Indian accent to accompany my surname, I would be ranked lower or rejected altogether during interviews. I sadly have a slight rally Finnish accent, so even if my grammar were to be perfect, I'm not a good hire compared with native English speakers even if I'm just as capable for the same position.
He immediately thought "we should use to to make porn without actors, that would make money" this seems closer to possible everyday.
Why is it interesting? Well for one, it raises an eyebrow as to the motivations behind either the technology or that of the promotional material produced for it.
I'd find it similarly interesting if, for an example, an entirely Russian team produced deepfake tech and produced promotional material for it entirely consisting of black people.
Especially in an era where we already acknowledge the prevalence of nation state cyber psyops / propaganda / manufactured news and "facts".
0. Portrait painted by trained artist in consultation with a witness;
1. "Identikit";
2. "PhotoFIT";
3. "DeepFaceDrawing" - [0] powered by ML & AI (which are trained on [1] & [2]?).
Trace over some famous cartoon characters and see what it outputs.
this is why I loved so much the pix2pix cat drawing demo https://affinelayer.com/pixsrv/ I hoped it would make a turning point for demoes but alas this is still unique
I can imagine it's much easier to train on type of face, but this could lead to later bias.