Super Resolution
blog.adobe.com
blog.adobe.com
Super-resolution is only guessing. It's ok for art, not for critical tasks.
“Wouldn’t it have been easier to just look them up in the phone book?”
Pure genius
I worry this is going to be a case where the marketing is at direct odds with public education efforts.
https://www.cs.huji.ac.il/~peleg/papers/icpr90-SuperResoluti...
>To be clear, this isn’t a knock on the Gigapixel software. Vaarakallio tells PetaPixel that the software is “amazing” and he uses it all the time.
Professional photo retouching is art. It's okay to use Gigapixel for an artistic task, it's not OK to use it to enhance a photo that you're going to show to a jury. That's what GP means by 'critical': use cases where it matters whether or not the pixels being added map to an objective reality rather than an algorithmic guess about what would look good.
Originally super-resolution was a hardware technique, and not "guessing". If you can [edit: this was poorly worded "control an imager positioning"] control imaging with finer resolution than the sensor has, you can take multiple images and reconstruct a higher resolution image in a principled way for say 2x resolution gain (cf super-resolution microscopy), also some telescope systems. Some modern photographic systems actually do this directly (piezo motors?) on the sensor.
Of course this only works if what you are imaging is reasonably static over the time needed to take all the images.
You can do an approximate version of this with video, with caveats because you don't control the motion. The key thing is, though, you actually have more data to work with.
This idea ran in parallel with image processing people attempting to estimate higher resolution from a single image for a while, and unfortunately the terminology stuck in image processing also. Something like resolution extrapolation is probably better but that ship sailed ages ago.
User 'twic' posted a link to a very interesting article that describes this and also explains the difference between photosites and pixels:
https://chriseyrewalker.com/the-hi-res-mode-of-the-olympus-o...
Also, cameras have had stabilization systems for a while now; I would assume they need similar pixel-scale precision. Some cameras shift the lens, some shift the sensor, but either way they need to shift the image-on-sensor by a very small amount, and also do it very rapidly.
[1] For example: https://www.pro-lite.co.uk/File/psj_piezoelectric_nanopositi...
I think in general satellite imaging is a good place to look for such implementations, since they have a naturally and predictably moving imaging system.
https://chriseyrewalker.com/the-hi-res-mode-of-the-olympus-o...
I also have an Olympus E-M1 MkII, but I haven't tried the high resolution mode yet. You just gave me a TODO item!
https://en.wikipedia.org/wiki/Image_stabilization#Sensor-shi...
https://support.d-imaging.sony.co.jp/support/ilc/psms/ilce7r...
https://www.nikonimgsupport.com/na/NSG_article?articleNo=000...
https://www.canon.co.uk/pro/stories/8-stops-image-stabilizat...
There is called stacking, and the subpixel offset comes naturally from path distortion in the atmosphere itself.
I remember using registax when the first batch of consumer telezoom point and click camera came out and got a 30x long exposure of the moon, with, well, mediocre results.
Theoretically, with 30FPS cameras like Sony A1 and said gyroscope data, you can create super resolution images.
IIRC Olympus' some handheld super resolution modes use both shake and sensor shift to increase resolution.
https://ai.googleblog.com/2018/10/see-better-and-further-wit...
> In the early 2000s, Farsiu et al. [2006] and Gotoh and Okutomi [2004]formulated superresolution from arbitrary motion as an optimization problem that would be infeasible for interactive rates. Ben-Ezraet al. [2005] created a jitter camera prototype to do super-resolution using controlled subpixel detector shifts. This and other works inspired some commercial cameras (e.g.,Sony A6000,Pentax FF K1,Olympus OM-D E-M1orPanasonic Lumix DC-G9) to adopt multi-frame techniques, using controlled pixel shifting of the physical sensor. However, these approaches require the use of a tripod or a static scene.
Super-resolution was already the name for this field in this 1987 review! https://kmh-lanl.hansonhub.com/publications/imrecov.pdf But 1987 is pretty recent as far as this field is considered. A. W. Lohmann and D. P. Paris talked about super-resolution this way in 1964! https://www.osapublishing.org/ao/abstract.cfm?uri=ao-3-9-103...
So no. It's not annoying that super-resolution is the name of this field. The name predates the entire field of image processing, it predates the invention of the digital camera! And it is the original use of the title "super resolution", people that do this in hardware adopted the name later.
Because of this, conversations around it are often confused, and it would be clearer if there were a different terminology for the latter.
Isn't this also how our eyes work?
[1] http://imajeenyus.com/electronics/20120908_improving_adc_res...
[2] https://accidentalscientist.com/2014/12/why-movies-look-weir...
Note that 24 fps film cinema runs at 72 Hz flicker due to triple-exposing each frame. That's why they could get away with 24 fps: Since they are dealing with stored images, multiple exposures were possible to avoid flickering. Using as low fps as possible was desirable because film costs money, film transports get more finicky as fps are increased and exposure gets tricky (doubling or even tripling the frame rate leaves you with a much shorter exposure time, which you need to compensate either with a larger aperture (which may not be available, or if it is, cost a lot of money), or more sensitive film (which may not have been available, or has much worse quality), or by having more light in the scene -- all of these are either undesirable or expensive).
So they chose 24 fps because the motion perception isn't quite in the shitter at that fps, though it's pretty bad (nowadays idolized as "cinematic look"; even more hip and cool if you do it on Youtube and deliver 24p which gets converted to 60 fps by means of 3:2 alternation, resulting in a super janky look known as "That Pro Youtuber Look"). This is the reason why people did not like The Hobbit.
Meanwhile video people went on to invent interlacing, because video was fundamentally about a non-stored image. Here the reason why low fps are desirable are again cost, since higher fps needs higher video bandwidth, and that makes everything more expensive and lets you have fewer channels (likely not a concern in the early days). Interlacing lets you have a higher frame rate (a CRT at 25 Hz or even 30 Hz will trigger nausea in everyone) without incurring higher costs, just like multiple exposures in a cinema projector.
They're still useful tools to have in your toolbox as a photographer or designer, even for critical tasks, and I don't really see how this is different. There may be certain failure cases, but everything has failure cases.
It’s like how modern word-guessing while texting can have really weird results because the guess of what you meant to say has been turned into real words which creates new meaning.
Where previously you’d just have had a few mangled words it’s now been corrected into proper words, often with an unfortunate sexual innuendo.
It is especially something that some cameras can do by deliberately doing sensor shifting.
They should not call it super resolution or at best emulated super resolution or artificial super resolution.
>you may want to uncheck detect faces… unless you want Ryan Gosling popping up all over the place.
Sooo not really a case against super-resolution, just a funny result of having used the wrong settings
Machine learning is educated guessing based on previously seen data. As mentioned by others there are ways to do super resolution that only uses the data available. I can't think of any that can upscale a single image, although I have vague memories of having seen something about using moiré patterns to infer the higher resolution texture of some features.
Who knows how this evolves and what new applications people may devise? For today, I agree: it's just art.
- https://en.wikipedia.org/wiki/Super-resolution_microscopy
- Stimulated emission depletion microscopy (STED): https://en.wikipedia.org/wiki/STED_microscopy
- stochastic optical reconstruction (PALM/STORM)
- structured illumination microscopy (SIM)
Here is one of my favorite STED imaging papers, looking at the skeleton of neurons: https://www.sciencedirect.com/science/article/pii/S221112471...
The big thing you have to look out for is the light intensity killing the cells or bleaching your signals. One of our collaborators is actively working on on-the-fly SIM processing for live cell imaging.
A quick glance at pubmed it looks like 11Hz was do-able several years ago https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2895555/ and this (sorry paywalled, I've heard sci-hub has the paper...) https://pubmed.ncbi.nlm.nih.gov/30478322/
STED is promising for live imaging too. Lots of beautiful pictures out there!
The interpolation techniques (making a photo out of hardware input) manage to get 10-15x more resolution out of that sensor layout compared to normal grid.
You also avoid Moire patterns with that too.
I've had a number of issues over the years, but my current issue is that when I try to open CC the interface elements all freeze and are unclickable (even though the window is still scrollable – very strange behavior). So I went to uninstall it, but I can't because Photoshop is installed. So I went to uninstall Photoshop, but you guessed it, I can only uninstall PS through CC, which is unresponsive.
Smh.
https://helpx.adobe.com/creative-cloud/kb/cc-cleaner-tool-in...
Note that this isn't the same as uninstalling everything via the official process since this leaves behind stuff like Adobe's Genuine client which verifies you're not using pirated software.
I needed to update my CV recently and expected to spend 1h in InDesign. I spent 6h in the end.
– InDesign crashes while saving and destroys my document. 1h lost.
– InDesign crashes while exporting a PDF of my document (9 pages). I hadn't saved. 1h lost.
– InDesign crashes (reproducible) when adding/inserting a page (mind you, that's page 10). First time this happened I hadn't saved for half an hour. I was really considering changing the text because I couldn't solve this. Then I found [1]. Quote:
> [...] after speaking to Adobe chat help, they asked me to send my file to them. They sent it back to me and everything went back to normal. [...] "File was corrupted , we recovered it by using scripts and then saved as IDML."
– Because of the above I had the idea of exporting to IDML. Re-importing then allowed me to add the page but I had subtle formatting errors where the last character before a tab or a newline on lines that had the font changed via a character style had the wrong style. Fixing this: 1h.
– When I re-arranged parts of the CV via copy & paste entire sections I copied lost the small caps/italic styles they had assigned (acronyms/names). Going through the entire document to fix this: 1.5h.
I should have known better. Less than two years ago I helped a friend do a snail mail mass mailing where we used a CSV file with addresses to create hundreds of (two page) letters. All in InDesign. Everything worked until we tried to export as PDF, for printing. The solution was to export as 'interactive' PDF and only export about ~100 pages at a time.
I bought Affinity Publisher already when the thing with the letters happened. But I naively believed updating my CV would be quick in InDesign.
In retrospect typesetting the CV from scratch in Publisher would have been the better choice.
Last week I helped a friend with a commercial that was mostly 3D and some motion graphics done in After Effects (Ae). We couldn't get it to render in After Effects 2019. It would run out memory and then just not render the frame or crash. In the end we exported the project for an older version and went back to an Ae CC version from six years before. That worked without any issues.
All this is just shocking. I used InDesign from 1.0 and it was not that bad, a decade ago. Ae ... the same. See above.
As of a recent update, Acrobat Reader (free version) refuses to let me open any document w/o signing into CC first. Another wtf.
What a friend of mine replied when he heard about my InDesign adventure:
> I'm on CS6 for anything Adobe. Just junk now.
[1] https://community.adobe.com/t5/indesign/indesign-crashes-whe...
I don't understand why people are so obsessed with Adobe, since their software nowadays isn't that good. There are tons of alternatives out there that work better and do the same thing, if not more.
Is it just laziness/reluctance to learn something new?
If you are just doing graphic work yourself, there are other applications like Affinity that can work, as long as you are not collaborating.
Affinity Photo. Much more stable than Photoshop and I like it a lot, but there are things Photoshop does that Affinity doesn't and I can't think of anything that goes the other way.
Krita. Fairly sleek, especially for OSS. Becoming very competitive for digital illustration, but not (and not intended to be) a great photo editor.
paint.net. Good for fast edits but simplified, not a true competitor.
GIMP. Ancient, slow, ugly, clunky, severely lacking in features, and somehow even less stable than Photoshop.
Yes we’re not talking about accuracy here, just perceived resolution, no need to hammer on that.
I hear it’s just monstrously fast on the M1 too.
I haven't tried Gigapixel but I have used Topaz' Video Enhance AI, which is phenomenal. I've been using it to upscale old TV shows which never got an HD remaster, to UHD.
Right now it's running through the first episode of Firefly, converting from 540p to 2160p (540p as the bluray rip was basically upscaled to 1080p from its original production, so I converted to 540p first in Handbrake with zero noticeable loss in quality since I used a near-lossless compression factor.. this provides better upscaling):
https://i.imgur.com/hcRYM5n.jpg
When it's done I'll run it through Flowframes for framerate interpolation. Then maybe another pass in Handbrake to figure out an optimal size for the end file.
Then I'll run through the rest of the season using the same settings I tested with this first episode.
I noticed with this model that really fine lines will have a tendency to get smoothed out a little. There's a similar model which should pull a little more detail but typically this one seems to work best. It's less noticeable once the video is in motion, compared to a still image.
I also probably removed too much grain from this, hence the more 'silky' look. It's nice for skin but less so for textures.
This post about "Super Resolution" is interesting because it starts with RAW format (which contains information about camera sensor arrangements), hence, the machine-learned model should not only memorize artificial details (what hair should be look like, what a tree-leaf should be look like etc, and use that to "hallucinate" a higher-resolution details, I liked to call this "hallucination" for that reason), but also relationship between complex interference of different sensors in their corresponding arrangements.
You can read more about RAW format and why exposing RAW format for photography is exciting (on everyday's camera, i.e. your phone) from this post: https://blog.halide.cam/understanding-proraw-4eed556d4c54
I agree on the smoothness, it takes a little tweaking to get right.
Susan Sontag's "On Photography" is a great read on this topic for anyone marginally interested in not just photography, but art in general.
this is no different than any other sensor. It doesn't mean replacing data with guesses is better than the sensor representation of the world, which is the issue at hand today.
Previously there was a clear line between scene-referred image data, which was treated as objective record of a 2D slice from a 3D world by way of measuring light, and output-referred image data—one of the countless lossy adaptations of that data to fit the limitations of some particular medium (display, paper, etc.) in order to be actually viewed.
The scene- to output-referred data conversion is where objectivity inevitably went out the window, but not earlier—the original scene-referred data was mostly treated as immutable.
What these guys are doing actually happens at the demosaicing stage, and from what I understand the resulting “super resolution” image is still scene-referred—but it isn’t representing the actual captured light anymore! In other words, we’ll have raw images that are partially “guesses” and no longer an objective record.
This isn’t necessarily good or bad, but is somewhat of a paradigm shift I’d say.
As a side note, I wish Adobe released the mechanism so that it could be made one of the demosaicing methods available in open-source raw processors, but I take it this won’t be likely.
Or any other sort of analog filter/lens that changes the picture for that matter
I can’t parse this for some reason, could you rephrase?
> Or any other sort of analog filter/lens that changes the picture for that matter
The glass on the camera affects the shape of the 2D slice of the 3D world captured by the camera and can attenuate light of different frequencies; technically the better you know which equipment was used, the stronger the element of objectivity to raw data captured by camera sensor.
Addendum: I take back most of my original comment. Adobe’s super resolution tool operates at the demosaicing stage, so image data it produces is no longer strictly scene-referred (it can’t be both demosaiced and scene-referred). The actual sensor capture is still scene-referred.
(I think we’ll see tools that take scene-referred data and output scene-referred data eventually though.)
At one point do we go from picture to painting ?
My goal when editing photos is to make them more clearly express how I felt or how I saw. This is often quite divorced from what shows up on the back screen of my camera.
Ansel Adams was surely no stranger to post-processing.
https://photofocus.com/photography/a-look-inside-ansel-adams...
i.e. you had the original scene that was captured by a digital camera (a lossy operation) and then saved as an image file (often also a lossy operation), and then a tool like this makes an educated guess as to what information was lost in the 1st and 2nd steps.
On the other hand, there certainly is a difference between working with the information (pixels) you've captured, and inventing information by either drawing on the image or creating new data.
This methodology falls squarely into the gray area between those two.
When I volunteered for a small newspaper I did my usual, sometimes significant, editing (Lightroom-level, not PS) and didn't see anything wrong with it (neither did they; of course edits look better than plain JPGs). But I can appreciate how this becomes much more important with increasing range. Imagine if Pete Souza spoiled 8 years' worth of Obama presidency imagery just because he edited them in some obscure way.
EDIT: Still can't find it, but here's a list which includes a number of other wartime photos which are proven or suspected to have been staged in various ways: https://militaryhistorynow.com/2015/09/25/famous-fakes-10-ce...
All advancements have simply given us more control in how painterly we render our photographs, but have never _really_ brought us closer to the truth.
[1]: https://en.wikipedia.org/wiki/History_of_photography#/media/...
I imagine this will tend to reproduce things in the dataset, e.g. up-scaling blurry text may look like fonts that it has memorized more than the original. Or upscaling a feather will provide details like the feather of more common birds. Or upscaling blured out numbers will pick some numbers at random [1].
We need to make sure people don't rely on these details, e.g. in courts, HR reviews, when reddit sleuths try and investigate an incident, when someone looks for cheating partners etc.
[1] https://www.theregister.com/2013/08/06/xerox_copier_flaw_mea...
The bigger problem is informal settings. Propaganda, for one.
Forensic evidence has been and still is systematically abused:
> * a 2002 FBI re-examination of microscopic hair comparisons the agency’s scientists had performed in criminal cases, in which DNA testing revealed that 11 percent of hair samples found to match microscopically actually came from different individuals;
> * a 2004 National Research Council report, commissioned by the FBI, on bullet-lead evidence, which found that there was insufficient research and data to support drawing a definitive connection between two bullets based on compositional similarity of the lead they contain;
> * a 2005 report of an international committee established by the FBI to review the use of latent fingerprint evidence in the case of a terrorist bombing in Spain, in which the committee found that “confirmation bias”—the inclination to confirm a suspicion based on other grounds—contributed to a misidentification and improper detention; and
> * studies reported in 2009 and 2010 on bitemark evidence, which found that current procedures for comparing bitemarks are unable to reliably exclude or include a suspect as a potential biter.
> Beyond these kinds of shortfalls with respect to “reliable methods” in forensic feature-comparison disciplines, reviews have found that expert witnesses have often overstated the probative value of their evidence, going far beyond what the relevant science can justify.
(https://web.archive.org/web/20170120002449/https://www.white... page 16)
Even more:
* Tire and shoe prints: https://www.apmreports.org/story/2016/09/27/questionable-sci...
* Lie detector tests: https://en.wikipedia.org/wiki/Polygraph#Effectiveness
* Burn patterns: https://www.pbs.org/wgbh/frontline/article/forensic-tools-wh...
So yeah, real people have been harmed by bad matching algorithms.
On the first day of trials of deep-learning based facial recognition here, a random person was arrested because the algorithm confused that person with another one.
Even more stupid, is that the person with "outstanding warrant" was actually ALREADY in prison.
So yes, AI managed to arrest the same person, twice, one time the real person, one time a random look-alike.
<Insert other race> all look the same: https://onlinelibrary.wiley.com/doi/abs/10.1002/acp.898
This bus is an ostrich: https://arxiv.org/abs/1312.6199
Racist autofocus: https://sitn.hms.harvard.edu/flash/2020/racial-discriminatio...
Google thinks black people look like gorillas: https://www.wired.com/story/when-it-comes-to-gorillas-google...
And here is an argument to look forward to: should the training dataset be build to represent the general population, or the specific subgroup that the software would be most often encountering? Because that FBI UCR...
Except if it's used in security camera, it's going to be a disaster. Those things have super low resolution, and software will be cheaper than upgrading. And models have huge bias.
Tons of fun.
Personally, I think one reason we still do this has a lot to do with detective shows being so ridiculously popular that people think that it's some sort of scientific process, when it's not. As a thought experiment, we'd probably have flat-earther-ism be the dominant belief if was like 15/20 broadcast TV shows are dedicated to glorifying flat-earther-ism. This means the most dangerous thing about "Super Resolution" is the public has already been "primed" with cop shows having the "enhance" feature.
Basically you can imagine that the blue subpixel is always to the top-left of the pixel. If you shifted the blue down and right one half a pixel you would have a more "accurate" production. In this way you can add a new pixel with the blue value closer to the right spot, then interpolate the blue of the original.
Of course you can also do logic such as detecting lines of different lightness and applying those on top.
So yes, especially with their machine learning they are adding new detail, but that is also likely some detail that was already there, but could not be conveyed with with the lower resolution. I wonder how different this would be from the simple approach of realigning the subpixels on a higher-resolution image and interpolating the "missing" subpixels. This approach may look better but wouldn't add any data.
[1] https://en.wikipedia.org/wiki/STED_microscopy
[2] https://en.wikipedia.org/wiki/Super-resolution_microscopy
The 'resolution limit' (Abbe diffraction limit [1]) is related to a few things, but practically by the wavelength of the excitation light and the numerical aperture (NA) of the lens (d = wavelength/2NA). When we (physicists/biologists) say 'super resolution', we mean resolving things smaller than what was previously possible based on the Abbe diffraction limit. So rather than only being able to resolve two points separated by a minimum of 174nm with a 488nm laser and a 1.4NA objective, we can resolve particles separated by as little as 40-70nm with STED (but it varies in practice).
STED does not accomplish this by estimating PSFs and fitting Gaussians, it uses a doughnut shaped depleting laser to force surrounding fluorescence sources to a 'depleted' state, and an excitation laser to excite a much smaller point in the middle of the depletion (see the doughnut in the STED wikipedia page, Stephen Hell and Thomas Klar won the Nobel Prize in Chemistry for this in 1999 [2].
I know PALM/STORM uses statistics, blinking fluorescence point sources, and long imaging times to build up a super resolution image based on the point sources and computational reconstruction.
Not as familiar with that one or SIM, but I know the "Pure physics/optics" folks I work with regard STED as the most pure physics based one that doesn't rely on fitting, deconvolution, or tricks (not that any of that is bad or wrong!).
[1] https://en.wikipedia.org/wiki/Diffraction-limited_system#The... [2] https://en.wikipedia.org/wiki/STED_microscopy
Photography for record keeping/science and photography for aesthetics/artistry diverge long before you get to techniques like super resolution. Which is still a fuzzy boundary because anyone who has taken a picture including a sunny sky can tell you raw photos generally don't capture how it looks to your eyeballs.
[1] = https://pytorch.org/tutorials/advanced/cpp_export.html
https://deepai.org/publication/gimp-ml-python-plugins-for-us...
edit:
I think longer term stuff like neural rendering will make super resolution less relevant. If you can re-create a 3D scene from a single photo or otherwise reconstruct the photo in a less-resolution-dependent way, then playing the super resulting game is less interesting (for users and researchers alike).
If I want to up-rez strictly for printing purposes, the PS one looks like the winner. But, obviously, it's subjective.
Both of the options had winners.
Unless ofc you're genuinely curious as to how it's different to Gigapixel and not just knocking Adobe :)
* demosaicing: interpolates color from nearby pixels. Each pixel gets just one of the tree color components. The other two are interpolated.
* decompressing jpeg: tries to guess information the compressor lost.
* black field correction: adjusts the brightness at every pixel to compensate for the different sensitivity at each pixels.
* de-vignetting: compensate for the border of the image being darker than the center.
* auto white balance: compensates for the fact that your eye’s color consistency doesn’t work as it would in the natural setting. This is a complicated way to get you to see the color you would have you seen the full scene.
All of these try to recover some aspect of the signal that was irretrievably lost by a previous step. They do this by making plausible guesses.
Keeping one frame and discarding the rest is a bit wasteful in a sense. The other frames have useful information that could be extracted by a well-trained AI to provide super-resolution, increased DoF, additional blur or shake reduction, etc...
I've deliberately kept all of my RAW frames, even the not-so-sharp or slightly shaky ones, because I foresee that at some point in the future this will be an automatic thing that tools like Adobe Lightroom will do that maximise the available image quality.
Storage is cheap, but I can never go back in time and photograph my memorable occasions with a better camera from the future...
By projecting the discarded frames onto your keeper frame, you've potentially got multiple samples of the same pixels and their neigbours...
(Although I guess the "every image is a plane" 3D transform is a bit simple and doesn't account for lens distortion)
[1] https://en.wikipedia.org/wiki/Photosynth [2] https://en.wikipedia.org/wiki/Bundle_adjustment [3] https://en.wikipedia.org/wiki/Scale-invariant_feature_transf...
My wife routinely shoots photos on her Pixel 3 that get a better response on our family whatsapp group than the painstakingly post-processed DSLR shots I create and post.
This could be an indictment of my failures as a photographer. Or perhaps my family has no taste in photos. But it's also entirely possible that a Pixel 3 is all the camera you really need for family documentary work ... and I've wasted so much money on unnecessary hobby gear.
And then, the jury decides to use "Super resolution" to "enhance" the picture. The ML model decided that what it saw was a gun instead of a rose.