Google's New AI Photo Upscaling Tech Is Jaw-Dropping
petapixel.com
petapixel.com
Without comparing a high-res original photograph to the high-res output photograph, we do not know if this fine technique is capable of producing nice-looking high-res imagery, or if it is capable of reproducing how an image of the subject would have looked like had it been taken in higher resolution.
In other words, does the output of the technique match the actual object in the photograph?
This happens a lot with junior peer reviewers, they focus more on what the paper isn't or could be instead of what's actually there.
However for the rest of us, evaluating whether the work is relevant in the surrounding world really is important. And for that, resemblance to originals is critical in some use cases. I.e. we're evaluating a fit for specific use cases.
Do you believe that is not a useful discussion?
I can recall seeing a conspiracy get traction on Twitter a few months back, where it was claimed that a photograph of a famous person was actually a body double. Someone used an ML upscaler to "enhance" the image, and their followers began scrutinizing the result: "The teeth are different!", "The nose shape is wrong!", "It's not the same person!"
The result is impressive, but the moment you started to use it on someone you're actually familiar, it becomes weird very quickly for the obvious reasons. The teech, for example, are never right.
This kind of "believable but not truthful" results are rampant in all these machine-learning based tools. It's not very harmful in case of upscaling a few photos I guess, but I've been bitten by it in an acclaimed translation service called DeepL. I use it to translate Japanese to English frequently, and have found that it often (nontrivially) made up sentences that don't exist in the original paragraphs, sometimes have the opposite meanings, or totally ignore part of the text to make the result "more fluent". And unlike traditional translation tools, they are very hard to notice if you know nothing about the original language. I have to from time to time use some more "primitive" translation tools, and compare the results side by side, to avoid such issues. It's frustrating.
Not to mention what happens when some good-intending but ill-informed agency release an enhanced photo of a suspect on the run.
The more interesting situation is when such "enhanced photos" are used in prosecutions. I suspect that ML "forensics" techniques and questionable expert witnesses will be used to elevate hunches to guilty verdicts and raise (false) conviction rates.
Sure, in the extractive sense. But in the generative sense you can.
The fact that the extra information is 'synthetic' and not 'natural' doesn't mean that it isn't extra info, just that it may or may not (probably not) correspond with ground truth.
Another way to think of it is that super-resolution is effectively a (possibly benign) man-in-the-middle. If what you're concerned about is the information flowing from Alice to Bob, Eve isn't adding any info, and may in fact be drowning the signal in more noise. But you can also see it as Eve communicating more info to Bob than Alice is to Eve. Whether what Eve is adding should be considered noise or signal is highly context dependent.
Right. All the processes being used for these purposes are stochastic.
Edit: Actually they are deterministic, in the same way a pseudo-random-number-generator is, and typically rely on a 'seed' that would have to come from a random source to be non-deterministic, and doesn't, so users get to have pseudo-random but reproducible results. But that's really getting into the weeds.
I'd be extremely sceptical of its use on medical images for this reason.
Probably a lot in some cases and a little bit in most others. I wonder how long before this gets used in court by an incompetent prosecutor.
[0] https://iterative-refinement.github.io/
[1] https://iterative-refinement.github.io/images/super_res_exam...
If you ever have something that you would be happy to substitute a very good painting for a blurry image then this is good. If you need to know what something actually looked like in high def (license plate numbers, micro tumors) this is useless, or worse than useless if it ever gets admitted in court.
Probably better to use the original link.
Because those details are generated by AI.
For example, the woman in the photo might have different teeth in reality. We can't learn anything about her teeth because the teeth in the generated photo are one of many possible solutions that match the input.
Actually, the photo now has less information for practical purpose as you don't know which details are real and which have been manufactured.
So about the only gain is to improve the photo for aesthetic reasons.
Sure, it is still a guess but a better one than humans can make.
It knows people have teeth, but the teeth are only estimated, not real teeth. Maybe they have been copied from some other face. The point is, it is not her teeth.
Imagine it was low res car photo with unreadable plates. "Enhancing" them this way would not bring back the plates. It could paste some plates, because we know this part of car usually has plates, but the real information (actual registration number) is already lost and can't be brought back this way.
In any photo some detail has been lost. This is trivially proven, as the amount of detail in the actual scene was many, many orders of magnitude more than in the photo file.
Detail that has been lost cannot be recovered, it can only be replaced with something that makes contextual sense which is what AI is doing here.
This can be dangerous. A lot of medical imaging deliberately avoids using any kind of lossy compression due to worries about artifacts in the image. Actually adding new pixels that are not in the raw image seems especially worrying.
Depending on these numbers it could be used as a screener test for example, where it is used before a more invasive test is done.
I'm not a doctor, but I am a physicist and former pro-photographer, what is noise and what is signal in an experiment whose output is an 'image' has nothing to do with what makes a photo look good to human eyes. Often the whole point of methods of visualization is to make the image look objectively bad so you can easily pick out the areas of interest by the fact they are an eyesore. Applying upscaling to those images will actively destroy vital data.
I'm not sure I trust people to maintain that discipline.
It may not be exactly as dangerous as the opposite (doctor looks at image thinks there is something suspicious, checks upscaled image to see if it's there as well), but it's still very dangerous.
Up-sample your Tinder photo? Sure.
Look for a sarcoma or bulging disk? No
https://petapixel.com/2020/08/17/gigapixel-ai-accidentally-a...
Imagine that thing being removed or enhanced by some algorithm.
Also, why the heck would medical images want to be upscaled?
One reason might be to reduce eyestrain.
When it comes to things like distinguishing a shadow on a scan, I think AI might actually be better 'detecting' whether something is a real shadow or just very similar to a shadow. I think it's just one of those things where AI up-scaling improves stuff ~80% of the time but is worse the other ~20%. The fundamental issue may become the same with self driving cars; people trust the AI too much and become inattentive themselves.
While you certainly can't add 'correct' information that doesn't already exist in an image, the upscaling could correctly make existing information more obvious. Assuming that the human brain functions pretty much like AI (or rather the opposite) then at some point AI will become as competent which means that eventually with enough training & tweaking it should be as good or better than having a second human perspective.
It could actually be useful for compression as a predictor, but only if you also store the residual so that the true original image can be reconstructed.
Pretty scary stuff.
I've mentioned this before[0], so quoting myself:
"for example, a research team may decide to not spend money on expensive scientific cameras for monitoring experiment, and instead opt to buy an expensive - but still much cheaper - DSLR sold to photographers, or strap a couple of iPhones 15 they found in the drawer (it's the future, they're all using iPhones 17, which is two generations behind the newest one). That's using COTS equipment. COTS is typically sold to less sophisticated users, but is often useful for less sophisticated needs of more sophisticated users too. But if COTS cameras start to accrue built-in algorithms that literally fake data, it may be a while before such researchers realize they're looking at photos where most of the pixels don't correspond to observable reality, in a complicated way they didn't expect."
--
Not that different mind you, but humans are obviously super sensitive to tiny changes on faces - it doesn't take that much to make it look like a different person altogether. For the non-face images it was much harder for me to really detect many differences, and they certainly didn't bother me.
I would be interested to see what it does with Doom guy, as mentioned in the OP comments.
It's adding information either way, I agree. The difference is that the old algos used information from the image itself, and this one uses information from a lot of other images.
Exactly what I was getting at, FWIW.
2. There are certain portions of the image that clearly do not contain enough resolution to be reconstructed satisfactorily. E.g. teeth, skin imperfections. I wonder how well a person would react if their teeth were either messed up or "fixed" by "the AI".
For enhancing images Remini works very well for human face enhancing - sharpening and filling missing details of noisy blurred images.
[1] https://www.pixelmator.com/pro/ [2] https://photographylife.com/reviews/adobe-super-resolution [3] https://github.com/idealo/image-super-resolution
Looks like a great way to save bandwidth for video conferencing calls
Will the resulting image appear even more realistic than the painting?
Was wondering about that too. It certainly produces realistic-looking high-res images, but especially when the article talks about potential uses ranging "from restoring old family photos to improving medical imaging", it seems like accuracy may be more valued than "looking realistic".
Spot on. This is not "zoom, enhance" — it is "fabricate plausible detail based on a training set". Using it for anything other than making pictures look nice would be disastrous.
Other people in these comments are talking about using it for law enforcement. Train it on a bunch of pictures of black people holding guns and now suddenly it will "reveal" guns in the hands of all black people in blurry CCTV footage. (This specific example is likely a little simplistic to be a problem in reality, but it demonstrates the problem of thinking it actually reveals some hidden detail.)
Probably better to have linked to Google's original blog post and used their title "High Fidelity Image Generation Using Diffusion Models": https://ai.googleblog.com/2021/07/high-fidelity-image-genera...
I wonder, if my assumption above is correct, how this would behave if the image was low-res to begin with due to whatever reason. Would it perform at the same level?
If you give it a low quality image, it doesn't do anything.
So it's quite limited.
To spell it out: This is not re-creating what was actually present in the original image. That is, and will always be, impossible due to fundamental limits of information in the source seed. What it is doing is using AI Hallucinations to create a believable-looking fake.
This system makes an image that looks better than a low res image, but it doesn't necessarily make it look more like the original image.