Learning to See in the Dark (2018)
github.com
github.com
This happens at the expense of detail in low-contrast areas, producing a plastic-like appearance of human skin and hair, and making low-contrast text unintelligible, which is why it's generally not done by default.
A lot of hand built filters (I see a lot of these in the audio space) have many hand tuned parameters, which work well in certain circumstances, and less well in other circumstances. One of the big advantages of NN systems is the ability to adapt to context more dynamically. The NN filters can generally emulate the hand designed system, and pick out weightings appropriate to the example.
I don't know what you mean by hand made filters and I don't know why that's a conclusion you jumped to.
[0]: https://raw.githubusercontent.com/cchen156/Learning-to-See-i...
Thank you, not only for setting me straight, but also for doing so as kindly as you did.
I'm sure you know exactly how much of which filter to apply for similar results. Laymen like ourselves will need a lot more trial and error. Their contribution here is to provide a push-button, automated mechanism.
I would have probably also tried something simple and given up due to the noise. So this is definitely interesting.
What you are describing is usually called automatic tone mapping. This is basically noise reduction and possibly color normalization from brightening a dark image. Them showing their black image as the starting point is silly, because jpg will make a mess of the remaining information. What they should show is the raw image brightened by a straight multiplier to show the noisy version that you would get from trying to increase brightness in a trivial way.
So presumably this neural net more or less does it for you.
Huh? At 1:40 in the video that's exactly what they do.
For those curious, our current approach differs in some very significant ways to the author's implementation, such as performing our denoising and enhancement on a raw bayer -> raw bayer basis with a separate pipeline for tone mapping, white-balance, and HDR enhancement. As well, we explored a fair amount of different architectures for the CNN and came to the conclusion that a heavily mixed multi-resolution layering solution produces superior results.
As other commentators have pointed out, the most interesting part of it is really coming to terms that, as war1025 pointed out, "The message has an entropy limit, but the message isn't the whole dataset." It is incredibly powerful what can be accomplished with even extraordinarily noisy information as long as one has a extremely "knowledge packed" prior.
If anyone has any questions about our research in this space, please feel free to ask.
We also have some more raw data[2] where there is the original bayer data available as .npy files with 40db analog gain applied, however I think the calibration targets show off what we are able to accomplish more dramatically. Finally, we have a short youtube video[3] that shows off how it works when applied to video.
[0] https://www.dropbox.com/s/0bm4dpxhn35vkhe/ALLIS_Investor_Int...
[1] https://www.dropbox.com/sh/k861saentyq1cs6/AADmO7X_L49nUkEI_...
[2] https://www.dropbox.com/sh/fv8omdf4fbx59m9/AABDnf6sdvv7rtIml...
Often flash is not the look people are going for, but would be okay with the flash firing in order to improve the non-flash photo.
As a proof of concept that this task can be tackled directly, a quick search brought up "DeepFlash: Turning a Flash Selfie into a Studio Portrait"[0]
Beyond denoising, we are already running experiments with very promising results on haze, lens flare, and reflection removal; super resolution; region adaptive white balancing; single exposure HDR; and a fair bit more.
One of the other cooler things we are doing is putting together a unified SDK where our algorithms and neural nets will be able to run pretty much anywhere, on any hardware, using transparent backend switching. (e.g. CPU, GPU, TPU, NPU, DSP, other accelerator ASICs, etc..)
What would happen if you
- begin capturing video (unsure of fps) on a phone-quality sensor in a near-dark environment
- pulse the phone's flash LED(s) like you're taking a photo
- do super-resolution on the resulting video to extract a photo...
- ...while factoring in the decay in brightness/saturation in consecutive video frames produced by the flash pulse?
I vaguely recall reading somewhere that oversaturated photos have more signal in them and are easier to fix than undersaturated. Hmm.
IIRC super-resolution worked with 30fps source video for better quality; I wonder if 60fps or 120fps source video would produce better brightness decay data, or whether super-resolution could actually help extract more signal out of the decay sequence too.
On the other hand, I'm not sure if super-resolution fundamentally requires largely consistent brightness in order to work as well as it does. :/
Perhaps individual networks could be trained/tuned to specific slices/windows of the brightness gradient. I also wonder if it would be useful to factor the superresolution process into each of the brightness-specific stages or just to do it at the end.
Nonetheless, it's kinda a neat idea, so I tried testing the feasibility of it. I set up a recent flagship phone that claims to have 960fps super-slow-motion video capture next to another phone with a strobe app at 12Hz with a short delay in between pulses.
https://www.dropbox.com/s/ha51ntucl3klkcb/cell_flash_960fps....
There are definitely a few frames where the LED is at an intermediate brightness, however teasing out the exact timings between the flash and the camera may prove to be difficult to correctly synchronize.
As for over-saturated images having more signal... although the PSNR calculation may give you a better number, in practice, a region that is over-saturated is just a blob of 1s on the image (assuming float64 pixel values of 0-1) and there is no information there to extract. With a black level near but not at 0, we've found there is often more information hidden in the 'dark noise' than can be discerned by the human eye alone.
Stepping back and forth throughout the frames (using mpv), the flash clearly enhances several spots of localized brightness where contrast pops out into clear relief.
The effect is clearest at the very bottom of the image which goes from "shadow blob" to "adequately discernible", but I think the area just above that (the 3rd vertical quarter of the image) is most interesting; the detail visible in frames 24-29 (immediately before 00:00:01 / 30.030fps) is excellent, and that's with the flash LED at peak brightness.
Flash synchronization would be effectively impossible to achieve (the camera would need to stream LED status information inside each frame), but achieving such synchronization may provide no net gain, even with "LED is on" information available, both because the exact point the hardware says "LED is off" will not necessarily correspond to the exact moment in time the light decays to zero (based on 1/960 = 1.0416 milliseconds per frame, the video suggests it takes apparently 2 frames or ~2.08 milliseconds for the light to decay), which will never be the same as the flash sends light outwards into arbitrarily different environments. I can't help but wonder if calibration references for everything from Vantablack to mirrors would be needed... for each camera sensor... and that there would then be the problem of figuring out which reference(s?) to select.
Staring at the video frames some more, two ideas come to mind: 1), analyzing all the frames to identify areas of significant difference in brightness, then 2), for each (perhaps nonrectangular) region of difference, figuring out the "best" source reference for that specific region. As an example reference, I'd generally use frame 13 for most of the image, and frame 44 or so (out of many, many possible candidates) for the bits that, as you say, become float64 1.00 :). Obviously a nontrivial amount of normalization would then be needed.
I'm not aware of how you'd do either of these neurally :) but the idea for (1) came from https://en.wikipedia.org/wiki/Seam_carving (although just basic edge detection may be more correct for this scenario), while the idea for (2) came from https://github.com/google/butteraugli which "estimates the psychovisual similarity of two images"; perhaps there's something out there that can identify "best contrast"? I'm not sure.
Trivial aside: I wondered why mpv kept saying "Inserting rotation filter." and also why the frame numbers appeared sideways. Then I realized the video has rotation metadata in it, presumably so the device doesn't need to do landscape-to-portrait frame buffering at 960fps (heh). I then realized the left-to-right rolling shutter effect I was seeing was actually a bottom-to-top rolling shutter. I... think that's unusual? I'm curious - after Googling then reading (or, more accurately, digging signal out of) https://www.androidauthority.com/real-960fps-super-slow-moti... - was the device an Xperia 1?
(And just to write it down for future reference: --vf 'drawtext=fontcolor=white:fontsize=100:text="%{n}"' adds frame numbers to mpv. Yay.)
[1] - https://github.com/cchen156/Learning-to-See-in-the-Dark/blob...
If you want to know what the next hot thing in software engineering will be, just pay attention to whatever Jeff Dean is doing.
There's also the issue of how hard it is to select the line of code (API call) that does the job (because the API surface is huge), which is nearly invisible attribute.
Mathematica is famous for being incredibly powerful with low custom coding, but also very hard to find the API call that does the precise thing you need.
IMHO, credit should always go to Alex Krizhevsky for the rapid spread of deep learning. He has shown us it was possible. Even without Tensorflow and PyTorch, we will be fine with Caffe, torch, mxnet or Julia.
But, many DNN concepts (and ML concepts themselves) can be described with a few lines of pseudocode. CNNs, RNNs, etc. can all be described in a few lines.
It's really quite amazing, most of the work goes into first creating the net work from theory, then training and tuning it until you get good results.
Generally, you just need to subtract the right black level and pack the data in the same way of Sony/Fuji data. If using rawpy, you need to read the black level instead of using 512 in the provided code. The data range may also differ if it is not 14 bits. You need to normalize it to [0,1] for the network input."
The Sony and Fuji training code looks mostly the same - they haven't bothered to pull out common code and re-use.
Machine learning is really machine-enhanced educated-guesswork, which has its place but also has its limits.
There is an entropy limit to the message, but the message isn't actually the only data.
One thing humans are great at is integrating existing knowledge into a messy situation and intuiting more than is available just from the raw message.
I.e. The message has an entropy limit, but the message isn't the whole dataset.
First of all, if you were to use the image as a communication channel, how much you could theoretically communicate is exactly the entropy of messages (by definition), and optimal communication means maximum entropy. Information theory already assumes shared existing knowledge in the form of codes; the codes essentially encode all this knowledge, which you could make an analogy in images to digits and shapes, etc. -- what makes them decodable is there is statistical redundancy, shapes do not occur arbitrarily (i.e. not every possible shape occurs, at least not with equal probability) and exhibit dependence between different pixels of a shape and even between shapes elsewhere in the image. Again the dependence is (generally) based on the statistics of the distribution of all possible images -- it essentially encodes all prior knowledge.
This redundancy allows reconstruction of losses in on part of the image from data elsewhere. It's the same principle used in error correcting codes, except codes are designed, while shapes are mostly natural (except things like alphabets, which are designed and indeed follow some principles of codes). But because they're not designed there's not guarantee of having a unique/reliable decoding (i.e. you can get a distribution).
I think that's a significant issue, because if your estimate doesn't match reality it could have important consequences for the use of the image: perhaps text goes from 'X is good' to 'X sucks'.
In this case a few things could be done:
1) Have some kind of watermark indicating the image was enhanced by a neural network, and possibly contains false information;
2) Have some kind of indication of reliability of the image: it should encode the multimodality/confidence of the decoding distribution -- how many different solutions does this have. If it is more or less unique, it would show as high confidence; otherwise it would show a low confidence indicator;
3) Instead of trying to convey uncertainty, the system could simply give up in cases where there is too much uncertainty, i.e. leaving the image dark. This could be done locally or globally, although locally it introduces a lightning consistency problem.
---
There's another important observation w.r.t. Information Theory/Statistics: it essentially assumes unbounded computational power (since this distribution could require analyzing arbitrarily large datasets). Of course this isn't true in reality. For example, the entropy of an encrypted of a redundant text is exactly the entropy of the plaintext string plus the entropy of the key (given an encryption ensamble or encryption prior) -- the process of (e.g. through brute force) finding the key doesn't concern statistics. However, with reasonable computational power, the (properly) encrypted stream is indistinguishable from a random string, hence it would have maximum entropy. So there are further computational limits beside statistical limits. In the case of encryption the function is again designed (to be not tractable), while in natural images the correlations are of simpler and hopefully more tractable nature (although I'm sure not always the case).
Eyewitness testimony is awful but given gold status.
Wait, did it? Isn't the middle photo being shown for comparison only, rather than as an input?
EDIT: Finally got the paper to load via the helpful wayback machine link provided in another thread. It looks like the goal was to simulate a long exposure with a short exposure. So whitewashing of "bright" areas in the original might be expected.
You do realize that, in a camera sensor, light is the signal, right? So the more light, the higher the signal-to-noise ratio, which means that yes it does have more information available to extract.
And yes, I quite realize that the input is (a). I'm guessing that in your display you are not seeing that there is a brighter spot in the middle of (a) corresponding to the whitewashed area in (c). Try maxing out your brightness if you're on a phone or laptop and you should see it. I can even make out letter shapes in (a) within this bright spot.
It's not trying to make things readable; it's trying to make things look like there was more light in the room when they were shot. In rooms with high lighting, some objects have glare. That's "correct"—it's what would appear in the training data.
For a long time, photographs were typically used as records. Even when they were art, they were typically records of something. Soon, we'll be typically doing so much more with them and will have to accept that photos can, but frequently won't serve as records.
Being able to read the title on the books in the example photo is great; you could rely on the title for evidentiary purposes, the smaller text probably not so much. So for a security camera it would do poorly at identifying the color of a car, but might well be sufficient to read the license plate.
You show in a courtroom a CNN-enhanced low light image of a car and it's there, literally 'clear as day' - the jury will find it pretty compelling. But maybe the data really wasn't there in the original image, and the CNN just filled in some blanks based on previous images of license plates, letters, and just random noise it had seen in the past.
The worry is when these kinds of algorithms get built in to basic image capture processes, so you never even see the raw data, only data that has already been filtered through the inbuilt prejudices of the CNN enhancement suite.
The camera never lies, but now it doesn't have to, because it can convince itself it saw something that wasn't really there...
You are right that the raw sensor data should always be preserved. But sticking with the license plate example, you could challenge a picture of a single car with a visible license plate far away in a wooded area, but it would be hard to refute a picture of the same car in a parking lot surrounded by other (non-suspect) vehicles whose presence there at the same time could be independently verified. In other words, if I can show that it accurately read the license plates of 9 other cars, the chances that it got yours wrong go way down.
That's assuming a single photo taken in the dark by an investigator. With a fixed security camera you would have an even larger basis of comparison, with a population of hundreds or thousands of license plates against which to rate it. I predict that before long we'll see preemptive certification for devices warranting the reliability of their image pipeline out to a certain distance at either the manufacturing or installation stage.
https://www.theregister.co.uk/2013/08/06/xerox_copier_flaw_m...
License plates are an ideal breeding ground for false enhancement owing to standardisation of appearance; an ML algo trained on lots of examples might, without due care, learn to replace as a well-known texture.
Also, y'all need to think more like prosecutors. Say you are dragged you into court on the basis of photos showing your car in the dark, and you object that the photo is from the ML 9000 security camera and it might be just imaging your license plate. The police/prosecutors will just 'borrow' your car and leave it there for a night and leave it up to the jury.
Forensic evidence can be and is regularly abused, but it can also be quite easily validated and it's massively persuasive to juries.
Let’s assume I’m innocent but some neural net has placed my car at the location of a crime.
You’re saying that if I challenge the evidence, the prosecutors will counter that by showing that if my car were there, the neural net would have produced a picture of my car? They don’t need to do that and it adds nothing to their argument. I’m not challenging that the neural net is capable of producing an image of my car.
No, the point is I am placed in the position of having to demonstrate that there exists some other car which under those lighting conditions the neural net would mistake for mine. That’s a far harder burden of proof for me to reach.
Honestly this is similar to the way fingerprint, DNA and hair sample matches are presented to courtrooms all the time so it isn’t a new problem. As you say, forensics are persuasive.
people bring this up all the time as hot-takes in these areas. it's conditional inference. it's no more disingenuous than linear regression.
This seems to replicate the post-processing we do in our brain (which is also a giant neural network). I wonder if the process is similar?
That’s not to say that the same “thing” is happening at the granular level at all.
But this is distinctly different from standard filtering functions, which can only work with entropy already present in the source image. So there’s a neat distinction.
The output from the CNN is essentially an “artist interpretation” of the source image. As such there could be “clarifying details” in the output that were in fact totally invented and not actually present in the source.
”The human eye can detect a luminance range of 10¹⁴, or one hundred trillion (100,000,000,000,000) (about 46.5 f-stops), from 10−6 cd/m2, or one millionth (0.000001) of a candela per square meter to 10⁸ cd/m2 or one hundred million (100,000,000) candelas per square meter. This range does not include looking at the midday sun (10⁹ cd/m2)[21] or lightning discharge.”
“ The pretrained model probably not work for data from another camera sensor. We do not have support for other camera data. It also does not work for images after camera ISP, i.e., the JPG or PNG data.”
Would be cool to see how they come up with better models that would allow them to overcome the above limitations
https://github.com/cchen156/Learning-to-See-in-the-Dark/blob...
If you take the dark image (a) from that and balance its color, the information that is present in it simply cannot contain the text from the book covers and so on. In fact, it's full of JPEG artifacts despite the image being a PNG. It would be useful if they presented a histogram equalized image of (a).
So if you're planning to do crime, make choices where the evidence relies on spectra rather than geometry. Steal Rothkos rather than Mondrians; baggy coveralls are in, form-fitting ninja wear is out.
http://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_...?
- Did they create a special network topology for this problem?
- Does the network need to see the entire image, or only an NxN subblock at a time?
- How did they obtain the training data? Is it possible to take daylight images and automatically turn them into nighttime images somehow?
I didn't even think this was possible. Have people ever done this manually before? Like without AI?
X27 is also using some kind of neural algorithm to denoise and get maximum out of the CIS
>Still images: ISO 100-102400 (expandable to ISO 50-409600),
[1]:https://www.sony.co.uk/electronics/interchangeable-lens-came...
So in this instance they're processing lossily on top of an image already processed lossily in-camera.
The other option is spelling it out.
Most people will read CNN as the news channel. Even those familiar with neural networks.
Agree that title should be changed.
Said differently: The percentage of people who are not from the US - but are aware of CNN as the Cable News Network, is higher than the percentage of people who are not machine learning experts - but are aware of CNN as a Convolutional Neural Network
2. CNN exists outside of the US
See also: https://en.m.wikipedia.org/wiki/Thiotimoline
The major peculiarity of the chemical is its "endochronicity": it starts dissolving before it makes contact with water.
The heart of each Predictor is a circuit with a negative time delay — it sends a signal back in time.
But still, could we make an effort not to devolve into what has happened on Reddit, i.e. comment sections which mainly consist of puns and other low effort jokes?
Would make sense to add Tensorflow to make it more specific.
And it is even more "HN" to comment on details of the title or the article because you don't really know what to say about the article.
Look at my comment.
It's a fight worth fighting. "C'mon it's just a joke lol" or "it's only a comment about the title" only assist in that transformation.
https://www.thedrive.com/the-war-zone/25803/this-is-what-col...