Imagine back to when digital photography just started, when we were stuck with 10-bit RAW images (before processing them to 8 bit JPGs). That's 1024 different values per channel, not a whole lot of room to begin with. Because we're converting a linear signal to a logarithmic one, that means that highest stop in the image takes up 512 of those values.
EDIT: As rightfully pointed out in a reply, 8-bit JPGs are not in linear colour space any more - so this logic does not apply there (however, JPG compression is more harsh on shadow details, since it's lossy, and it tries to move that loss of information to the parts of the image where you're less likely to notice, like the shadows, but we're digressing)
By that logic, the low light data should have a lot more banding than the highly exposed data. Try it with level tools on a RAW file and see for yourself.
So you can imagine that when the dynamic range of (part of) the scene that you want to capture is less than the dynamic range of your sensor, you want to be on the right edge of the histogram (all else being equal - if this causes motion blur due to long exposure it's pretty useless)
Caveat: this is potentially outdated knowledge, perhaps the analog/digital converter in modern cameras actually converts the voltage of a CCD to a (pseudo)logarithmic scale for the digital output, which might partially circumvent this issue (although the sensor is still linear). And nowadays many of the better cameras have 14-bit A/D converters - the usefulness of which also largely depends on the quality of the sensor - so there's a lot more detail in the shadows than before.