Making Blurry Faces Photorealistic Goes Only So Far
spectrum.ieee.org
spectrum.ieee.org
While this is obvious to anyone familiar with the technology, it's difficult to explain to casual observers. The image output looks real. It feels like it could be real. It's free of traditional scaling artifacts that would trigger suspicion. Without additional explanation, it's easy to see why casual observers would assume the hallucinated upscaled version is an accurate representation of the original image.
Historically, police sketches and blurry surveillance images are obviously low quality enough that people inherently know they're approximations. The problem with these ML hallucinated upscaled images is that they look and feel real enough that they bypass people's suspicions. We can try to present them as "Here's what the suspect might look like", but when they look like a full-resolution photograph, people will simply assume that it's exactly what the suspect looks like.
The company that develops this tech will want to market themselves as some sort of oracle to the people and LEO. Junk forensic science sticks in courts forever, we should be careful before we assent to more.
You arbitrarily picked 16 possible photorealistic faces out of a total solution space of what? Millions?
Wouldn't the balance of probability be on someone in the general population of humans more closely resembling one of your 16 candidate images than any of your candidates resembling the Ground Truth image?
Doesn't the problem get both better and worse as you scale your N up from 16? That is, it would be better because one of your candidates is more likely to match the Ground Truth, but it would be worse because you've also widened your net for catching false positives?
Which makes me wonder, can we train a model on only famous people, weighted by their relative famousness?
Well, yes, but why would you want to?
Wouldn't the solution be to find a 4 different axes and pick faces that represent different endpoints. With a little psychology to help us identify what features are best to inform the public of we should be able to create a collection of photos that will be more likely to result in someone identifying the suspect than either the single photo option or the photo collection option. We would still need to test to see if it is better than the single lower quality photo option.
A more correct view could be “It produces output that was trained on minimizing differences in the pixel/feature-space”.
My point being: minimizing actual human perceived differences is seldom/never done but always rather with some proxies or complex loss function constructs that make sure that no scientist has to actually deal with human observers and their preferences ;-)
As a layperson? Magic, because that’s how these things are marketed by startups and journalists.
It produces no black faces. Even as a guess.
Though, things like black people being more poorly recognized by facial recognition can be a blessing in disguise depending on the circumstances
That, by itself would be biased for the ML, since there are fewer images. But it's not an indication of bias by the people programming it.
But this would imply you need huge data-sets for every minority, not matter how small a proportion of the population.
It might instead be necessary to teach the model about race as a concept, so it can categorize images, and then process them correctly. But of course that leads to a different can of worms since you are explicitly making a "race aware" ML.
If I want to work on a project like this, and the only appropriate training dataset I have is photographs white males, am I not allowed to work on generating faces until I've fleshed out the data set?
It's not like they are offering some service to the general public where they can de-pixelate faces. It's a research paper.
A final note: The only professor with her name on the paper(Cynthia Rudin) is very active in researching the intersection of machine learning and social justice. I'm not so sure she would put her name on a paper that can be flippantly described in 3 sentence comments on internet forums as having "some issues" wrt race.
Recently, however, there's been a lot of work pitching upsampling for "deblurring" faces, which seems like a great way for LEO to run the programme until they hit on a face that most likely looks like a person of interest.
Hell, you can even make it explicitly so that the net takes two inputs, the blurred image and a suspects image, and generate a plausible upsample that is similar to the suspect.
That sounds like bad faith science that no judge would accept, but perfectly legitimate technologies like DNA testing has historically been abused this way in the courts.
Least needed "proof" ever.
"Ok, zoom in... alright, now enhance that... now upsample using AI trained on criminal database... "
You can just as well take these pixels and apply any kind of blur filter - you still wouldn't retain any more information. If you go from say 1k x 1k pixels down to 100 x 100 pixels, you end up with 1% of the original information, no matter what you do.
It's not about turning a shitty image to a nice a clean one. It's about turning a lowres image into a hires version, i.e. you start with 100 x 100 pixels (doesn't matter how you obtained them - could be a section of a much bigger image, for example) and try to extrapolate a 1k x 1k pixel version from it.
The pixelation you see in the examples is just a representation of what little information you have to work with. It's NOT in any way shape or form related to where you get these pixels in from in the first place.
Just to give you a different context here: imagine this upscaling being used to "enhance" a single face in an image like this [1] - there's no way "Gaussian blur" or whatever filter you'd like gets you more information out of that.
In the example image I created, I applied a Gaussian blur to the "pixelated" (i.e. enlarged) version of the marked image section.
As you can see, enlarging the same section using super-sampling (i.e. similar to what you proposed), doesn't change the information content and one version can basically be transformed into the other.
Also I don't understand why the first Mona Lisa results in a picture that when pixelated again, wouldn't produce the original pixelated picture. It is as if the create a face inspirated by the original, but not one that could ever be the original.
It can still be useful for entertainment purposes.