The article says that the highest light level in a scene is 1,000,000 times brighter than the lowest. It does not say, or even imply, that you can see the difference between those levels of light to a one in a million, which is what 20 bits needs.
Can you see what I mean? The factor between the lowest and the highest is entirely separate from the question of how many levels between you can see, which is what directs how many bits of resolution you need.
One thing to add: Wikipedia sayeth "the eye senses brightness approximately logarithmically over a moderate range". If we go with that, then you presumably want to encode brightness logarithmically, and the number of bits you have available will determine the ratio between your adjacent quantized levels. In that case I believe the ratio between adjacent levels would be exp(ln(max_range_ratio)/(2^bits)).
I explained it in another comment but basically the "20 bit" value is based on an idealized digital image sensor rather than a game rendering pipeline (which operates in floating point). It has admittedly proven somewhat confusing.
All we know is that the range is N to 1000000N, but we do not know in and of itself that N is the smallest delta perceivable or possible.
I don't know anything about games but at least in cameras the relation you want between the bit depth and dynamic range is determined by what you want to measure rather than a formula. An eye tracker I have is a 10-bit IR camera because a normal 8 bits are insufficient to both note that IR LED reflections are way brighter than everything else, while also having sufficient detail in the lower values that the edges of the pupil are easy to detect. At least, without having the scale be logarithmic or discontinuous or otherwise compressed.
The real benefit of HDR is in the "more values" part since, as the author notes, our displays have a very limited dynamic range anyway.
Hell, I like to think I'm somewhat knowledgeable about these things and I read straight past that line thinking "that's a million, that's about twenty bits, okay".
But you are absolutely right, now that you pointed it out it's obvious.
And indeed you can get the same range using only one bit (per channel, that is) and if you had high enough (very high) resolution and proper dithering, you'd totally get away with it, too. In that case, the tonemapping goes just before the dithering.
I didn't interpret the articles premise that if you had 20 bits, you could 100% reproduce the light in a scene.
You only need 1 bit if you declare an encoding scheme where 1 == 1 million. You need 20 bits to define 1 million DIFFERENT luminances. Their ratio is subject to an arbitrary scaling.
> To express a ratio of 1000000:1, you need at least 20 bits.
You mean to express all the integer factors in the range 1000000:1, you need at least 20 bits. With 20 bits you can represent 1 times brighter, 2 times brighter ... 1000000 times brighter.
But there's nothing special about those coefficients. 20 bits does not allow you to represent 1.5 times brighter. That's still in the range 1000000:1 but 20 bits isn't enough to represent it.
If we're happy to skip 1.5 times brighter, why can't we skip all the even integer times brighter values, and use 19 bits?
Things get wonky if you don't have a linear scale with a true zero; in such a scale the low end of your N:1 contrast ratio (in the smallest representation) has a value of 1, and the high end has a value of N.
Also, if that type of "wonky" throws off your rendering pipeline, you're bound to get something else wrong.
Such as ever having a linear scale with a small number of bits in your pipeline. The linear scaled brightness stays afloat all the way through (cause floats have this handy feature of being transparently sorta-logarithmic in the way they use their bits, even 16-bit floats beat 20-bit ints for that purpose), only at the very end you apply the tonemap+gamma function(s), then dither, then truncate to fixed (8) bit integer.
edit: And aren’t output color spaces already highly nonlinear due to gamma correction?