At any given number of bits used for representation, using floating-point numbers instead of fixed-point numbers (integers are a special case of the latter) increases the so-called dynamic range, i.e. the ratio between the greatest and the smallest representable numbers.
This advantage is paid by increased distances between neighbor numbers inside the subranges, because the number of representable numbers is the same for floating-point and fixed-point, but the floating-point numbers are spread over their wider dynamic range.
Depending on the application, either the disadvantages or the advantages of a greater dynamic range are more important, which determines the choice of floating-point or integers (actually fixed-point), and when floating-point numbers are chosen, one can allocate more or less bits for the exponent depending on whether the dynamic range or the rounding errors are more important.
For ML/AI applications, it appears that the dynamic range is much more important than the rounding errors, which has caused the use of the Google BF16 format, which has great dynamic range and big rounding errors, instead of the IEEE FP16, which has a smaller dynamic range and smaller rounding errors, and which is preferable for other applications, like graphics (mainly for color component encoding), where the rounding errors of BF16 would be unacceptable.
In the parent article, there is a figure that is confusing, because in it the dynamic range appears to be the difference between the positive number and the negative number with the greatest absolute values.
This is very wrong. The dynamic range is the ratio between the (strictly) positive numbers with the greatest and the smallest absolute values. The dynamic range can be computed by subtraction only on a logarithmic scale, which is why in practice it is frequently expressed in decibels.
For instance, for INT8, the dynamic range is not (+127)-(-127)=254 as it appears in that figure, but it is 127 divided by 1, i.e. 127. Similarly, for FP16, the dynamic range is not (+65504)-(-65504)=131008 as it appears in that figure, but it is 65504 divided by 2^(-14), i.e. 1073217536, a much larger value, which demonstrates the advantage in dynamic range of FP16 over INT16 (the dynamic range of the latter is 32767).
With a dynamic range defined like in that figure, there would be no advantages for floating-point or for BF16, because with an implicit scale factor taken into account, one could make that "dynamic range" as great as desired, for any integer numbers, including for INT8. Nothing would prevent the use of an implicit scale factor of one billion, making the "dynamic range" of INT8 as 254 billion, or of an implicit scale factor of 10^100, resulting in a "dynamic range" of INT8 much larger than that of FP32.