Adobe Photoshop's 'super resolution' made my jaw hit the floor
petapixel.com
petapixel.com
These methods are already in extensive use (most smartphone images use extensive noise-reduction techniques), but we must be ever-cognizant that image-processing techniques can add yet another layer of nuance and uncertainty when we try to understand an image.
Thoughts?
1x1 0-0-0 (original)
1x2 0-0-1 (made up)
1x3 0-0-2 (original)
The point is worth belaboring because people have a tendency to take the output from these systems as Truth, and while they can be interesting and useful, they should not be used for things for which the truth has consequences without understanding their limitations.
You're right to compare this to how our brains reconstruct our own memories, and the implications that has for eyewitness testimony should inform how we consider the outputs from these systems.
This might seem like a quibble, but once you dive a little deeper into it, you realise that there's enormous latitude and subjectivity in the way you do that interpretation.
What's even crazier is that this didn't come with digital photography. Analogue film photography has the same problem. The silver on the film doesn't become an image until it's interpreted by someone in the darkroom.
There is no such thing as an objective photograph. It's always a subjective interpretation of an ambiguous record.
The nice thing about this was that you could hand the E-6 off to a magazine and end up with a photograph printed in the magazine that was very close to the original film. Any color shifts or changes in contrast you could see just with your eyes. You could drop the film in a scanner and visually confirm that the scan looks identical to the original. (You cannot do this with C-41.)
This was not used for forensic photography, though. The point of using E-6 was for the photographer to make artistic decisions and capture them on film, so they can get back to taking photos. My understanding is that crime scene photography was largely C-41, once it was relatively cheap.
With ML-enhanced photos, you might have a distanced face that is "enhanced" by the model, to become a face that wasn't there. Or a fingerprint, a birthmark, a mole, etc.
1. However good the guess is, it's still just that: a guess. Taking the standard of "evidence in a murder case", the OCR can and probably should be used to point investigators in the right direction so they can go and collect more data, but it should not be considered sufficient as evidence itself.
2. OCR is a relatively constrained solution space - success in those conditions doesn't mean the same level of accuracy can or will be reached outside of that constrained space.
To be clear, though - I'm making a primarily epistemic argument, not one based on utility. There are a lot of areas for which these kind of machine guessing systems are of enormous utility, we just shouldn't confuse what they're doing with actual data collection.
But the issue manifests as characters being incorrectly identified because of an algo t
Edit - re OCR do you mean e.g. from a picture of a blurred license plate we could rule in or out a subset of possible numbers, depending on how blurred, like a B could be a 8 but not a L? (And sorry if your example is unrelated). This is valid, and unrelated to super resolution, you can do this analysis with Nyquist and point spread functions.
I don't. Everyone knows this already and it seems like a lot of people are just saying it over and over to look clever.
I personally take a lot of high resolution art photos. One that is deeply in my memory is a picture I took of the Manhattan bridge from the Brooklyn side with a 4x5 camera. I can get out the negative and view it under magnification and read the street signs across the river. (I would link you, but Google downrez'd all my photos, so the negatives are all I have.) ML upscaling probably won't let you do that, but on the other hand, it's probably pointless. It's not something that has a commercial use, it's just neat. If you want to know what the street signs on the FDR say, you can just look at Google Street View.
(OK, maybe it does have some value. I used to work in an office that had pictures blown up to room-size used as wallpaper in conference rooms. It looked great, and satisfied my desire to get close and see every detail. But, you know you're taking that kind of picture in advance, and you use the right tools. You can rent a digital medium format camera. You can use film and get it drum scanned. But, for people that just need a picture for an article, fake upscaling is probably good enough. The picture isn't an art exhibit, or an attempt to collect visual data. It's just something to draw you into the article in the 3 milliseconds before you see a wall of text and bounce.)
Wow, Google ate your one digital copy? That's tragic.
What's the approximate resolution you could get out of a scan from these labs?
I was interested in getting into film cameras at one point, and I was disappointed with how low the scanning resolution is from most labs. For example mpix only advertises 18MB, which they say is good enough for a 12in by 18in print. North Coast Photo (what Ken Rockwell recommends) is even worse! What if you want something to put on the wall?
Granted, if the original film you shoot is perfect you can have a print done from the negatives, but that kind of defeats the point of having a high quality scan as a backup.
With my home setup, I can easily do 50-80 megapixels on a 4x5 negative. I use a flatbed photo scanner (the Epson V800) and wet-mount the film. It is Quite The Process involving a lot of parts (liquids, optical film to place on top of the mount, careful calibration of the focus point, etc.) but the results are excellent and relatively repeatable. But, all in, I'd estimate that it's probably a half hour of labor per photo, so you can see why labs charge so much. (Dry mounting doesn't save that much time, because of the amount of time you spend avoiding dust and optical artifacts intrinsic in using two extra sheets of glass.)
The real professionals use drum scanners. They are quite expensive, but offer incredibly high resolution and decent throughput for the operator. Looking around at prices, for $100 you can get a 320MP scan of a 4x5 negative yielding a 1.7GB file. http://www.drumscanning.com/rates.html For fine grained black and white films, you can certainly extract information that actually exists. As someone who mostly uses T-Max 400, though, that would be overkill. I can't imagine getting much more information out of my photos than I get with a flatbed.
In summary, you can see why even pixel peepers are content with their Sony A7R. Press button, get 50 megapixels. And no toxic chemicals being absorbed through your skin.
Looks like you can get those scanners used for pretty reasonable prices. Maybe if I've got a house one day and I think the odds of having to move within a few years are low I'll get into it and try setting up a lab.
> In summary, you can see why even pixel peepers are content with their Sony A7R. Press button, get 50 megapixels. And no toxic chemicals being absorbed through your skin.
Yep, and on top of that we're not limited by the sRGB gamut or bit depth issues of early digital cameras. Recent ones produce raw files that are extremely easy to develop and manipulate into something very nice looking.
The thing is, even on top of the enjoyment some people get out of working with film, if you're after a particular film-like look you might be able to save yourself a significant amount of post-processing time by just going with film. I've seen no one-click filter that can approximate it.
Never mind memories; there are parts of our eyes that aren’t responsive to light at all. We’re always hallucinating.
Are you referring to the blind spot, or something else?
Interestingly, the blind spot turns out not to be a design requirement, it is a contingent feature that cephalopods like octopuses (whose eyes evolved independently from vertebrates') don't have.
Technically, they were correct -- it wasn't her. It was an algorithm's best-guess reconstruction based on training data of other people's faces. Unfortunately, neither the original poster or anyone else in the thread seemed to grasp this concept.
It's nice for a casual Instagrammer, but then a lot of science and engineering also gets done using COTS equipment. I worry that at some point, a lot of money will be burned, a lot of time wasted, or even lives lost, because someone didn't notice they've based the conclusions of their scientific experiment/engineering analysis on such "computer best guesses". As a researcher, you'll see a weird pattern on some of the photos and will be left wondering, is that a real phenomenon, or is it just one of the black box, trade secret neural networks in the camera choking on input data it wasn't trained for?
> a small usually collapsible bed often of fabric stretched on a frame
But in this case, COTS is apparently...
> commercial, off-the-shelf
In other words "photo equipment" or "consumer/retail photo equipment."
Upon further reading[0] it seems odd to use the term here, but maybe I'm misunderstanding something. It's often used for software and has a key phrase...
> packaged solutions which are then adapted to satisfy the needs of the purchasing organization
But it's possible the term has been co-opted to mean something else now.
"Commercial off-the-shelf or commercially available off-the-shelf[1] (COTS) products are packaged solutions[buzzword] which are then adapted to satisfy the needs of the purchasing organization, rather than the commissioning of custom-made, or bespoke, solutions."
So for example, a research team may decide to not spend money on expensive scientific cameras for monitoring experiment, and instead opt to buy an expensive - but still much cheaper - DSLR sold to photographers, or strap a couple of iPhones 15 they found in the drawer (it's the future, they're all using iPhones 17, which is two generations behind the newest one). That's using COTS equipment. COTS is typically sold to less sophisticated users, but is often useful for less sophisticated needs of more sophisticated users too. But if COTS cameras start to accrue built-in algorithms that literally fake data, it may be a while before such researchers realize they're looking at photos where most of the pixels don't correspond to observable reality, in a complicated way they didn't expect.
In the novel, quantum computers (rather than ML per-se) are tasked with interpolating more and more detailed data from astronomical observations, to the point that tracking individual members of an alien species on a distant world, underground, is possible. Eventually it is noticed that cutting off the astronomical data entirely doesn't interrupt the interpolated data. Then things get weird.
I won't go into further plot details, as that would be spoilery, but it is a pretty good book, reminiscent to me of Greg Egan's oeuvre (the novel is actually by Robert Charles Wilson).
As an aside, the term of art is "make-or-buy" if you want to be able to Google it.
The discussion we are having is interesting because COTS are notorious for their hidden costs and how difficult they are to properly budget. Having to find a way to disable or reverse advance post-processing in a camera would be a fairly typical example of that. In this specific case it might mean having to commission a custom firmware from the camera manufacturer - something which is very much doable but might end up costing you as much as buying bespoke equipments for inferior results in the end.
Make-or-buy seems more a term for manufacturing industry/SCM. TIL
Most research papers are crap anyway, in a much more fundamental way and for much worse reasons/bad incentives with far more impact than "computational imaging".
This is probably the last thing I'd worry about when thinking about "millions/time/wasted" for some research.
https://www.theregister.com/2013/08/06/xerox_copier_flaw_mea...
It's not too dissimilar, I agree, but there are differences.
It’s much easier for the model to blow up uninteresting pieces of the photograph than interesting pieces.
When playing around funny things happen too: recursive upscale/sharpen and analogue artifacts begin resembling topography, molten metal etc.
Now imagine this being fitted into military drones, which it almost certainly is.
Even better comparisons are in the blog post for a competing product: https://www.pixelmator.com/blog/2019/12/17/all-about-the-new... (likely the same algorithm, but using a different training set, so results will be different from what Adobes product does).
It has comparisons with nearest neighbor, bilinear and Lanczos filters and uses a slider to make it easier to see the difference.
Papers on this task: https://paperswithcode.com/task/image-super-resolution
Aren't those comparisons misleading though? The ML sample is 4x the resolution, comparing to something that is supposed to be used to also make something 4x the resolution. They aren't upsampling the comparison, so I don't know what they're actually doing with it. It's so misleading I just assume the company is scummy.
In this one, i's are nearly illegible, but are corrected. It appears to be correct here, but could be wrong.
https://blog.adobe.com/hlx_fa6ee5a67d3f2e5b1945745176e6b955e...
This one's weirder. In the original, I can't read many of the letters. It looks like G[EL]NE[XR]A[IL]DI[ER][XK][IT]OR - lots of guesses. The enhanced version is GENERALDIREKTOR (I think). At any rate, it's much more confident in the spelling than I am, as a human.
https://blog.adobe.com/hlx_a337d462293b51f4d91182cbbebf68cf4...
https://www.youtube.com/watch?v=c0O6UXrOZJo
>On the scale of things too horrible to contemplate, "document-altering scanner" is right up there with "flesh-eating bacteria". Since 2006, Xerox scancopiers literally are making stuff up. They, for example, replace digits with others in scans.
Those images show an effect that looks similar to LCD subpixel rendering, which is an artifact of a scanner working at the limit of what it's sensor is capable of producing (typical CCDs have subpixel stripes (or arrays), just like LCD screens.
Scanners typically overcome this by oversampling and then downsampling the raw data to smooth out the effect. In theory you could also do this with less oversampling if you could manage to get the scanner to do subpixel offsets, and oversampling isn't needed at all if the CCD doesn't use striping or a Bayer array but instead layers the RGB detectors on top of each other, like the Foveon X3 CCD.
Anyway, it is pretty clear that the main benefit of the upsampling interpolation in these particular images is in correcting these subpixel color fringes. Downsampling back to the original resolution should still yield an improved image, which is quite intriguing.
Really glad to see they included Lanczos. It's extremely frustrating to those in the know to see comparisons that only use subpar algorithms. The worst only use a B-spline or nearest-neighbor upscale, and end up looking like one of those eyeglass prescription ads for seniors. Something like Lanczos is the minimum acceptable, I think.
Adobe's bicubic produces obvious severe artifacts, I assume it's something like a Catmull-Rom spline cubic.
2x upscaling isn't all that impressive to begin with (e.g. produce 4 pixels from 1) and can be done in fairly high quality using traditional non-learning algorithms.
I'm much more impressed by 4x and 8x super-resolution. I'm really not sure what the big deal is with 2x.
That alone makes it worthwhile, surely? No careful application required.
I was expecting the upscaled image to have extra "invented" detail from the supposed ML, as I've seen elsewhere.
But looking at these upscaled images, there isn't any at all. There's no extra texture, nothing.
I can't find any difference at all, like you say, from just some bicubic interpolation with sharpening.
No jaw dropping here, unfortunately.
...wait, THAT'S NOT MY FRIENDS!
It also has the rather large benefit of not being tied to Creative Cloud, which is one of the worst bits of software ever made.
https://github.com/AaronFeng753/Waifu2x-Extension-GUI
It has several models for processing animation and photos.
If you want it even tediously automates the scaling of videos and GIFs by exporting each frame, upscaling, and joining back together.
[1] https://www.pixelmator.com/blog/2019/12/17/all-about-the-new...
[2] https://www.pixelmator.com/blog/2020/09/15/pixelmator-photo-...
Edit: the photography bundle, which includes Photoshop, is $10/month, not Creative Cloud as a whole. Thanks to mastazi.
"As noted above, in 2021 I analytically derived the Fourier transform of the Magic Kernel in closed form, and found, incredulously, that it is simply the cube of the sinc function. This implies that the Magic Kernel is just the rectangular window function convolved with itself twice—which, in retrospect, is completely obvious. This observation, together with a precise definition of the requirement of the Sharp kernel, allowed me to obtain an analytical expression for the exact Sharp kernel, and hence also for the exact Magic Kernel Sharp kernel, which I recognized is just the third in a sequence of fundamental resizing kernels. These findings allowed me to explicitly show why Magic Kernel Sharp is superior to any of the Lanczos kernels. It also allowed me to derive further members of this fundamental sequence of kernels, in particular the sixth member, which has the same computational efficiency as Lanczos-3, but has far superior properties."
You can get on the order of sqrt(N) improvement in resolution from N images with an optimized system. I've done it in the past, with a hand held DSLR and Hugin (the panorama stitching program, in this case used to align the stack of almost identical pictures with subpixel accuracy)
This means that soon for many people digital cameras outside of their smartphone will become an even more niche product.
I've gone on hiking trips where my "challenge" was to only use my phone camera. It wasn't much of a challenge for landscapes.
Phone (imho) is for quick and dirty, not for a 'it's time to do proper photography'.
Well if you collapse the problem space to a single point that corresponds to a phone's standard field of view, then it won't be a problem...
But what if you wanted to catch a photo of a rare bird in flight at 500mm equivalent, or a surfer caught at 1/4000th of a second?
stuff like this - assuming it's a GAN under the hood it just tries to guess a 'plausible' possible interpolation, but if you're giving it very little information about what's in the original image, there will be a wide range of plausible images it could have arisen from, so the output can be very far from the truth.
And for artistic purposes, why does it matter how the final result was arrived at? If we have powerful and easy techniques for realising an artistic vision, that doesn't seem like a bad thing?
Off topic: I remember a few years ago, some students got very impressed by GIMP's Lanczos-3 upscaling that was much better than the photoshop version they had access at the time.
Often enough it will look worse than bicubic interpolation. Chroma noise get's crazy most of the time.
What's remarkable at: pattern cloth, straight lines with large contrast (electric wires against a blue sky, dark glasses mount against white skin, etc).
Nice to have option all in all.
When I use ML these days is it all hardcore data crunching by remote servers or is some of it running on my phone/laptop?
So any model that can be trained for generic use (Eg. Not trying to deep fake my specific face) can presumably be run on local machines.
Thanks!
Yes my jaw hit the floor when I saw this headline.
Plausible (i.e. "good looking" or "believable") results are not the same as actual data, which is why enhance wouldn't work on vehicle licence plates or faces for example.
Sure, the result might be a plausible looking face or text, but it's still not a valid representation of what was originally captured. That's the danger with using such methods for extracting actual information - it looks fine and is suitable for decorative purposes, but nothing else.
Let’s take the classic example of enhancing a blurry photo to get a license plate.
Humans may not be able to see much in the blur, but an AI trained on many different highly down-res’d images could at least give you plausible outcomes using far less data than a human brain would be able to say anything with confidence.
You wouldn’t hold it up as the absolute truth, but you’d run the potential plate and see if it matched some other data you have.
So yes, it wouldn’t magically add any more information to the image, but it could be far better at taking low information and giving plausible outcomes that are then necessary to verify.
GAN produces totally different results if you slightly change the input.
So, as others are also saying, these "enhances" are great for decoration and absolutely should be ignored as facts or truth (specially when it comes to face and license plate and others used by the law enforcement).
You’re fighting against “would it be reliable” but that isn’t the claim.
The claim is could it be better than human, and the answer is yes, it just depends on how well trained it is and the dataset.
But this is also entirely testable. I guarantee much like Go, if we set up a “human vs AI guess the blurry image” competition that AI will blow us out of the water. It’s simply a data * training issue, and humans don’t spend hours on end practicing enhancing images like they do playing Chess.
Again - it won’t be perfect, obviously. It will have false positives, of course.
Doesn’t mean it can’t be better than human.
Also GANs are pretty irrelevant, the model structure has nothing to do with the theory.
This is not really an issue that is new or limited to things that are called AI.
https://www.theregister.com/2013/08/06/xerox_copier_flaw_mea...
The difficulty will be making sure people treat it the same way, because it looks like a normal image.
Hmm. Mandating the use of style transfer to make it look like an artist's rendering would probably have the intended effect.
That's not the same as fabricating information, though. A blurry image still contains a whole bunch of information and correlation data that just isn't present in a handful of pixels.
This is not super-resolution, but something different entirely. Super-resolution would mean to produce a readable license plate from just a handful of pixels. That is an impossible task, since the pixels alone would necessarily match more than one plate.
The algorithm would therefore have to "guess" and the result will match something that is has been trained on (read: plausible), but by no means the correct one, no matter how many checks you run on a database.
To illustrate the point, I took an image of a random license plate, and scaled it down to 12x6 pixels. 4x super-resolution would bring it to 48x24 pixels and should produce perfectly readable results.
Here's how it looks (original, down-scaled to 48x24, and down-scaled to 12x6 pixels): https://pasteboard.co/JSu3WDU.png
The 48x24 pixel version could easily be upscaled to even make the state perfectly readable. A 4x super-resolution upscale of the 12x6 version, however, would be doomed to fail no matter what.
That's what I'm getting at.
Just for shits and giggles, here's the AI 4x super-resolution result: https://pasteboard.co/JSu7jkP.png
Edit: while I'm having fun with super-resolution, here's the upscaled result from the 48x24 pixel version: https://pasteboard.co/JSu9Qh6.png
I hope some day there will be an episode of a crime show where by chance two teams of detectives will independently work on the same case without noticing each other and by using standard police methods they will come to completely different incompatible conclusions and detain two different suspects who of course both confess after going through standard police interrogation. Screenwriters, use your powers for good!
(Actually Czech writer Karel Čapek (the same who invented the word robot) did practically the same thing in one of the short stories in Stories from Another Pocket, everybody should read it together with Stories from a Pocket)
edit: I was joking, but people pointing out that you still can't create something out of nothing etc might not be thinking big enough. I think this technology absolutely has the potential to help. police are literally still using artists impressions - photofits, to find perpetrators
An artist impression tells the audience that it is inaccurate. A realistic photo tells the audience that this is _exactly_ who we are looking for.
That aside, you'd still probably need a ML PhD to have a chance of correctly interpreting the results, given the myriad potential issues with current systems.
https://www.pixelmator.com/blog/2019/12/17/all-about-the-new...
Truth is that anything other than the full suite (and maybe the photographer plan) doesn't make sense financially. And then they killed the month by month subscription as you said.