It's a bit mindblowing, but I'm pretty sure that every single person who has responded to you doesn't understand what you're asking and doesn't understand how to answer it.
The short answer is: You're right that it's wrong to say that you need 20 bits to cover the range. Because clearly you can cover the range with 10 bits or 100 bits by subdividing differently. More bits is better, but there's no rule of required multiplication of bits for fabricated metrics like integral brightness multiples. 20 bits is just required to cover the range with integers without skipping any, which is an entirely arbitrary decision.
The long answer is:
The reason you want more bits to represent the larger range is exactly as you say because one doesn't actually care about _range_ but rather about the utilization of increments within the range. Now I'm going to describe something that you clearly already intuit for completeness...
Say, for example, that you have only 3 steps (let's call them 1, 2, and 3) in your range instead of a million, but you want to represent the full range of darkest dark to brightest bright with those 3 steps (for simplicity you could also be thinking in black and white instead of color). Now, you take a picture of a normal indoor environment with some gentle indirect light, and maybe all of external values would be, on the million range, between 350k and 650k. If you linearly divide your range onto 3 steps, when you look at your resulting image you will have a flat solid block of 2s with no 1s or 3s at all. And yet your original setting had an absolutely massive 300k spread. But it didn't have significant brights or significant darks, because you were indoors with some indirect light and there was no fireball sun or deep cavern of despair in the image, so you didn't measure any differentiation in values.
So you decide to switch it up and go inside a dark cathedral with some wonderful stained glass windows (everyone loves cathedrals with stained glass when looking at HDR). This time you do the same thing, but instead of measuring all 2s, you measure all 1s with a chunk of all 3s for the sunlit windows, but no 2s at all.
So then you say, "What I should do is take some of those 3s and some of those 1s and pretend that they're actually 2s when differentiation is too high, and take some of those 2s and split them into 1s and 3s when differentiation is too low! That way I'll see more details, because there will be more efficient utilization of the segments of my available output spectrum!"
But _then_ you say, "Oh crap! I can't do that! I didn't actually capture the differences! I captured what I captured, and what I captured was the aforementioned terrible undifferentiated data!"
So the reason you want to be able to represent values on a _densely_ _segmented_ range is just so that you can differentiate things that are only slightly different in absolute overall term (think photon counts, I guess) because aesthetically those slight differences matter. It means you can do things like know when some of your 2s should become 1s and 3s, because the 2s aren't actually 2s anymore; they're a bunch of tiny fractional variations between 1.5 and 2.5, and that's data you can work with.
You can perform tone mapping with any number of bits. If record 10 increments instead of only 3, you can obviously convert those 10 down to 3 in a way that gives a better image than just starting with 3. And the more bits you have, the more linearly you can record the environment without losing the differentiability between things that are in absolute terms on a large scale very similar. And then you get to do fun things like add a light bloom effect around something that is _very_ bright compared to its surroundings without losing details within the very bright thing and without losing details within the very dark surroundings.
Having more representational bits just means you can tone map scenes that have really bright parts, really dark parts, and really medium parts without losing any of the details inside those parts.