Super Resolution: Image-to-Image Translation Using Deep Learning in ArcGIS Pro
esri.com
esri.com
I would argue this is not true - rather, super resolution generates a plausible high resolution image that would look like a given low resolution image if it were downscaled (i.e., it's not going to recover real details, it's just going to sharpen lines and potentially show details that look real but might not be).
Edit: As an example, in the lower half of figure 5, the algorithm displays circular white dots on the roof, when in reality they are rectangles. An image analyst using this tool might, for example, incorrectly geolocate an image taken on the ground. This tool probably needs a warning label on it.
To make absolutely clear: any details revealed by this upscaling do not exist. They are guesses based upon other imagery
Super resolution as a name covers a number of techniques, but some absolutely can recover real details. Eg Superresolution from video will integrate over time, allowing sub-pixel accuracy. Imagine a grid of pixels, each of which covering a defined area of the source. Now, if that grid moves, the area covered by each pixel will be slightly different. The differences between frames can then be used to determine with accuracy a higher resolution output image.
There are superresolution methods that _can_ increase resolution, but they're in the form of combining multiple captures closely spaced with each other that are slightly offset. "Drizzle" was the original method in astrophysics, and while that method is long gone, it is common to do in remote sensing imagery for many instrument types. Similarly pansharpening (which interpolates multispectral information using a higher resolution black and white image captured at the same time) can actually improve resolution and is commonly used.
This is not improving resolution in any way. Two objects that blur together into one will still appear as one object.
no - if you relent like that perhaps
At this point everyone knows the memes about CSI Enhance, the Xerox compression image hallucinations, blah blah. Must we constantly revisit subjects at a grade-school level?
https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres...
8e2fcd496
I _constantly_ hear people claim that this method really does add information. My own company constantly makes similar claims, and internally, even for people who actually do work in the field, they honestly do think it adds resolution.
> "Figure 1: Recovering high resolution image from low resolution"
This is terribly misleading. There is no data being "recovered" here. A ML model is guessing at the result based on other training data. It may (and in fact is likely to) make stuff up entirely based on what it thinks should be there.
I'm generally pretty live and let live when it comes to ML-based upscaling, because if some drawing or personal photograph has some artifacting it's pretty harmless. But when you're doing it in a tool whose data will be relied upon for Real Stuff, one needs to be painfully accurate when it comes to what the system does and its limitations.
This kind of image transformation is cool when all you care about is esthetic quality of the image, but if you want to see details, then Mark I Eyeball is the best tool you can hope for, because if the thing is unrecognizable, you'll know it and won't make things up to pretend it is not.
This is more intuitive with height data. For example a 100m grid could have a 10m cell right beside a 20m cell. A point on the boundary between both cells is more likely to be at 15m than either 10m or 20m. And you can improve that estimate using other nuances. Picking a predicted value can be more truthful.
I agree that caution is important but that is true with any imagery analysis. A human analysing imagery will already be using a lot of intuition as it is.
Imagine you are looking for swimming pools to find water to fight a wild fire. All you have is a "blurry" image. That blue shape on the image could be a weirdly shaped tent or patio. And that would be very obvious on high resolution drone imagery. But 99% of the time it will be a pool and that is good enough.
All data has limitations and this is no different.
This was a big thing in the medical imaging community (where I did my stint as a CV researcher), folks were hallucinating microscope images and CT scans with no information theory justification as to why it worked.
Super resolution IS possible, but it must be done by synthesizing new pieces of information, not by inferring based on what other similar looking objects looked like. A cool technique by my former advisor does this with microscopes [1].
Deep learning has a place here, just not as a "lets create information" step, but as a way to learn how to synthesize additional information about images from more sources (i.e. more similar to how Google does Night Sight [2]).
Edit: if you want to see (an attempt) at using deep learning in this field you can checkout one of my papers [3].
[1]: https://en.wikipedia.org/wiki/Fourier_ptychography [2]: http://graphics.stanford.edu/papers/night-sight-sigasia19/ni... [3]: https://openaccess.thecvf.com/content/ICCV2021/html/Cooke_Ph...
Sometimes detail accuracy doesn't matter but the presence does.
Just about every image you ever view has had some manipulation applied. Sometimes that results in a "better" image.
Consider all astronomical images for human consumption, even smartphones adapt now to skin tone.
I'm playing hogwarts legacy, a recent AAA game which is very demanding, and where aesthetics are very important on a mediocre PC precisely because FSR from AMD (and if I had an Nvidea GPU DLSS and DLAA).
Deep learning image enhancement is totally appropriate in your smartphone, as there the goal is not accuracy but perceived quality. Doing this to satellite imagery where the primary consumer cares about accuracy is what I call "reckless"
I vaguely remember the Rittenhouse trial had an expert discuss if pinch-to-zoom could introduce false information.
What happens if there are multi-million dollar economic outcomes depending on the details of the remote sensing content, as in disaster response.
Firstly, this isn't just doing edge detection etc as happened 20 years ago. It's creating a deep learning model which fills in information based on extrapolating from looking at lots of other images and making an educated guess as to what is most likely to be present. It's a fairly new approach and works much better than previous methods, given enough training data.
This is, of course, imperfect, but claims that it's just to "look good" and is of no practical benefit are incorrect. For instance, we published a paper using DLSR in microscopy, to help experts identify synaptic vesicles. In the original images the experts had a 3x higher false negative rate compared to the DLSR images.
Finally, claims that this can't be called "super resolution" are ignoring years of peer reviewed published research in which this is exactly the name used for this approach. Yes, super resolution can also be achieved using other methods which take advantage of additional data (such as multiple images), but that does not mean DLSR has to be called something else.
(I'm not involved in ArcGIS Pro, but am the lead author of the fastai framework which underpins ArcGIS training behind the scenes, and have published papers and tutorials on super resolution using deep learning.)
ML-based "super resolution" is more trying to "guess" what the extra pixel data is based on images of other things. I thought that was (should be) called "upscaling".
Are we changing the meaning of this term? Or do I have it wrong? (Or do they have it wrong?)
Saying you can “recover” high resolution seems like a bit of a stretch.
Uh… I mean, it looks pretty but it’s basically just invented a bunch of random probabilistic crap on your map right?
Is this a thing? People actually want their maps to be pretty and wrong?
O_o
Strange times.
Particularly blurry blocks that are “maybe cars?” in a backyard being turned into high resolution cars. Or blue splat into “100% a swimming pool” seems… pretty dubious.
Of course if outliers / anomalies are very important for persons business or use-case they shouldn't use this feature.
A lot of things on maps repeat themselves a lot. Most roof HVACs look the same, road lines, cars, trees, etc. How likely is it that a car shaped blob that's the same as the other 10 million car shaped blobs in the training data isn't a car? Not very.
Satellite images have 0.4 m resolution or better these days, what matters most is good color quality (e.g. hyperspherical pansharpening), sensor dynamic range and good lighting conditions (sun elevation in particular).
I guess with this technique you could do ML image synthesis guided by SAR satellites - that way you could sort-of look through clouds from space, as long as you don't mind the image being largely fictional.
Depending on the application, this is a good model. It makes the maps look sharper. There are photo editing plugins that use similar technology.
I would hope courts of law would not admit any image edited in such a manner as evidence, though. The model is making it up.
Phone manufacturers haven't been making revolutionary advances in image sensors or the centuries long science of camera optics — they've learned how to take a few noisy signals and extrapolate them into plausible high fidelity data.
Computer, generate a list of applications where technology like this should absolutely not be used.
This seems terrifying — wouldn't this, for example, synthesize identifying details about motor vehicles not present in the sample, drawn from training data? About people? Etc etc
This sounds awesome for making a consumer mapping product feel higher quality. Hell, it could even serve an anonymizing function that is pro-privacy ("a roof" not "your roof"). But it feels incredibly reckless to direct this toward those listed industries.
People in surveillance and forensics etc. should be confronted with the limits of the quality of the data they are using, we should not try to synthesize extra confidence in their analysis by making the images seem higher quality than they are.
Is it that this is now available in a particular commercial mapping platform?
Please someone advice.