Float Toy
evanw.github.io
evanw.github.io
I remember this when somebody says f128 will never be a thing, because if you'd asked me in 1990 whether f16 will be a thing I'd have laughed. What use is a 16-bit floating point number? It's not precise enough to be much use in the narrow range it can represent at all. In hindsight the application is obvious of course, but that's hindsight for you.
https://en.wikipedia.org/wiki/Half-precision_floating-point_...
Yeah, floating point is nothing more than your standard scientific notation of numbers, e.g.
digit.xyz... * 10 ^^ +/- some exponent
The exponent is simply shifting where the decimal point is. The only different for floating point is that everything is base 2 because computers :DInterestingly you're right that a bunch of fp functions can be faster than integers equivalents (although I'm still not convinced that this isn't simply due to the reduced number of bits involved), and more fun the relative performance of operations can actually change vs what they would be in integers. Also this is in the context of doing it in software vs hardware, where again the perf costs of things change.
1. Since the normalized significand will always be 1.bbbb, the '1' bit is stripped from the significand representation, except:
2. To extend the range, the lowest 'zero' value of the exponent drops the leading '1'. This is referred to as the subnormal range
3. The highest exponent value, when the significand is zero, is used to represent positive and negative infinity
4. The highest exponent value with a non-zero significand is used to represent NaN
5. There are many different values usable for NaN by software, including a differentiation between 'quiet' NaNs and a (I believe implementation optional) 'signaling' variant, which will raise an interrupt when used. The idea is that these can be used to convey additional information, and that the signaling variant as well as the right interrupt handlers can be used to add additional functionality such as variable substitution.
6. Zero is signed
The technical details of how it handles _every_ case weren’t particularly relevant.
However just to address 1,2 with some hilariousness (autocorrect wants this to be “hilarious mess” which may be more correct).
Ieee754’s 80 bit format was the first widely deployed format, and was largely used by intel to get the other manufacturers to stop trying to reduce the functionality of ieee floating point because “it couldn’t be implemented, implemented efficiently, etc”. However because of that it has a quirk that was fixed for fp32,64,etc.
FP80 uses an explicit bit for the leading 1. That means it can do 1.0 * 2 ^^ N, or 0.1 * 2 ^^ N it should hopefully be immediately obvious why this could be a problem :)
Not only do the multiple representations for a single value result in sadness, it also gives us a variety of concepts like pseudo-denormals, pseudo normals, pseudo infinities, pseudo nans, etc all of which cause their own problems.
Mercifully by default the only hardware fp80 implementation now (since 286 maybe?) defaults to just treating them as invalid and converts to Nan. But you can set a flag to make it treat them as it did originally.
There's some notes about the advantages of half-float pixel in the openexr documentation: https://openexr.readthedocs.io/en/latest/TechnicalIntroducti...
I don't think "obvious" was the best adjective, but "small memory/file size footprint" is probably the quality that's easiest to understand.
Presumably they do benefit from the dynamic range as otherwise you'd think int16 would be sufficient, and not suffer the conversion costs.
Right now the following bits
0 10001 1000000000
show the calculation as
1 x 2^2 x 1.5 = 6
when it's more clear to actually say
1 x 2^2 x (1 + (512 / 1024)) = 6
Like the famous Quake3 Fast Inverse Square
Honestly I think having fp80 would be super interesting as it was the first end-user available hardware implementation of ieee definition of floating point, and also as a by product shows some design "mistakes" that were ironed out by the time fp32/64 were formalized.
In financial apps, it is perhaps more common to use specific numerical types that don't lose precision.
One simplistic option is to simply represent the smallest unit (say, tenths of a penny) with a fixed point number rather than a unit like USD dollars in floating point.