> Subtract one frame from another, that leaves us with samples of raw thermal noise values.
> Calculate the mean of a noise frame. If it is outside of ±0.1 range, we assume the camera has moved between frames, and reject this noise frame.
> Delete improbably long sequences of zeroes produced by oversaturated areas. For our 1920x1080=2Mb samples and a natural zero probability of 8.69%, any sequence longer than 7 zeros will be removed from the raw data.
> Quantize raw values from ±40 range into 1,2 or 3 bits: raw_value % 2^bits.
> Group quantized values into batches sampled from different R,G,B channels, at big pixel distances from each other and in different frames to minimize the impact of space and time correlations in that batch.
> Process a batch of 6–8 values with the Von Neumann algorithm to generate a few uniform bits of output.
> Collect the uniform bits into a new entropy block.
> Check the new block with a chi-square test. Reject blocks that score too high and therefore are too improbable to come from a uniform entropy source.
This reads like a highly ad-hoc process with nothing resembling a formal justification for any of its steps; nor its general outline, nor any of the magic numbers used in it. It's unclear what properties are being achieved and how exactly the steps guarantee those. There is no analysis of the predictability of the data by an adversary, either.
What does this get you that SHA512'ing the entire raw image bitmap doesn't? Using statistical tests makes sense to verify that the camera data isn't pathologically anomalous (say, all zeroes or all 255), but I don't understand why this sort of procedure is preferable to using a strong hash function to extract randomness from an image sensor's output.