That is no more true than my current body temperature being precisely 36 4/10 degrees Celsius because the thermometer reads 36.1.
All numbers in that range alias to the same double, yet differ wildly in their decimal digits beyond the fifteenth. That "long tail" of digits is a meaningless residue arising from the arbitrary difference between the chosen center-of-range point and the actual number it approximates.
> Printing only 15 digits is not doing programmers any favors
Note that the ISO C function printf uses 6 digits of precision under the %g conversion specifier by default. So a team of experts can find it justifiable to severly truncate precision on printing. I find that justifiable also, like this: printing is not only for constants, and trivial expressions like 0.1 + 0.2, but for the results of complex calculations, which accumulate significant rounding errors. Six digits of precision is prudent: for many complex calculations, it will avoid misleading the user with too much precision. Fifteen decimal digits will be wrong after just a few operations; there is hardly any "headroom" to absorb error.
> just makes the lie even worse
If merely neglecting to reveal some aspect of the stark truth is tantamount to lying, then you're also lying when you pull a number with 54 significant decimal figures out of a 64 bit double. Revealing some truth without an explanation can also be construed as "lying", as in "[N]one of woman born shall harm Macbeth".
I personally find it convenient that when I bang the token 0.3 into my REPL, it comes back with a tidy 0.3. I know that there are several values of double which will print as 0.3, and don't compare with exact equality except in special circumstances (like when deliberately using integral values not far from zero).
[1] https://www.cs.berkeley.edu/~wkahan/JAVAhurt.pdf
[2] https://math.mit.edu/~stevenj/18.335/fp-myths.pdf
[3] http://lyoncalcul.univ-lyon1.fr/IMG/pdf/FloatingPoint.pdf
See Misconception #2 of http://lipforge.ens-lyon.fr/www/crlibm/documents/cern.pdf.
Floating point numbers are _not_ intervals. If you read carefully any formal definition of floating point numbers (The IEEE standards, TAoCP Chapter 4, or Higham's _Accuracy and Stability of Numerical Algorithms_, to name just three possible references), you will see that floating point numbers by definition form an exact rational subset of the extended real line.
> All numbers in that range alias to the same double, yet differ wildly in their decimal digits beyond the fifteenth. That "long tail" of digits is a meaningless residue arising from the arbitrary difference between the chosen center-of-range point and the actual number it approximates.
See Misconceptions #1 and #2 on Kahan's list (https://www.cs.berkeley.edu/~wkahan/JAVAhurt.pdf).
You are not being clear about the set of real numbers a user may input and their ultimate representation as a floating point number. While it's true that many real numbers round to the same floating point number _f_, that's not the same as then saying that _f_ carries that interval around with it. The latter is false, since the intervals do _not_ propagate in floating point arithmetic. It's also false that the floating point number _f_ is the midpoint of the set of numbers that round to _f_; the precise set depends on the rounding mode and the granularity of the set of floating point numbers around _f_.
> that floating point numbers by definition form an exact rational subset of the extended real line.
Sure, that's what they form; it's not what they (usually) denote.
> It's also false that the floating point number _f_ is the midpoint of the set of numbers that round to _f_;
It is close enough to being the case for my purpose in the grandparent article. Whether the cluster of numbers is lopsided one way or the other doesn't really detract from my point.
So 1.0 + 2.0 == 3.0.
And of course 1x10^7 + 2x10^7 != 3x10^7, right?
Let's say you want to represent "two-million" in 32bit float:
Your answer is: 0x1e8480 * 2^0.
However, I was expecting something like: 0x1.e848 * 2^X. But then "125-thousand" is 0x1.e848 * 2^Y, so how do you standardize what is the correct exponent for integers, X or Y? And how do you know there is not a pair (M,Z) != (0x1.e848, X) such that M * 2^Z is also two-million?
let's take a simple case where you have one bit of fraction.
exponent 0, 2^0 = 1, fraction (1)0 = 1.0b = 1
exponent 0, 2^0 = 1, fraction (1)1 = 1.1b = 1.5
exponent 1, 2^1 = 2, fraction (1)0 = 10b = 2
exponent 1, 2^1 = 2, fraction (1)1 = 11b = 3
if we have more fraction bits it looks like this: exponent 0, 2^0 = 1, fraction (1)00 = 1.00b = 1
exponent 0, 2^0 = 1, fraction (1)01 = 1.01b = 1.25
exponent 0, 2^0 = 1, fraction (1)10 = 1.10b = 1.5
exponent 0, 2^0 = 1, fraction (1)11 = 1.11b = 1.75
exponent 1, 2^1 = 2, fraction (1)00 = 10.0b = 2
exponent 1, 2^1 = 2, fraction (1)01 = 10.1b = 2.5
exponent 1, 2^1 = 2, fraction (1)10 = 11.0b = 3
exponent 1, 2^1 = 2, fraction (1)11 = 11.1b = 3.5
So for any given exponent multiplication with a fraction with an invisible bit creates products that are hemmed between 1.0x2^n and 2.0*2^n so there's no collisions between any pair (e,f) in your representation.If the exponent is 1, then the mantissa is multiplied by 2 to get the actual value. Put another way, the mantissa is the value divided by 2 and rounded, i.e. shifted right one bit.
> 1x10^7 + 2x10^7 != 3x10^7, right?
This particular case is interesting.
If we're talking about 64-bit doubles, then you have 53 bits of precision for integers. All those values are well within that range, so this arithmetic and comparison are done with precise integer values, and the expression will compare equal.
If we're talking about 32-bit floats, the range of precise integer values is -16,777,216 through 16,777,216, or -2^24 through 2^24.
1x10^7 is 10,000,000, within that range, but the other two values are outside the range.
So you might expect that != would be the answer here.
But if you test it, that isn't the case: the expression compares equal!
The reason: although 20,000,000 and 30,000,000 are outside the range where every integer has a precise representation, they are within the range where every even integer is precise: -33,554,432 to 33,554,432. Values within this range but outside the range of precise integers are rounded to a multiple of 2.
Similarly in the range -67,108,864 to 67,108,864, all integers which are a multiple of 4 are represented precisely.
Basically, as you go outside the range of precise integers, values get rounded to a multiple of 2, 4, 8, etc. as required.
When the mantissa overflows the available precision, it is shifted to the right enough so that it fits without losing the most significant bits. Instead, the least significant bits are discarded, and the exponent is incremented by the number of discarded bits.
Of course you could choose other values where the rounding doesn't work out in your favor, and then you'd get the unequal comparison you expect. A simple example:
33333333.0f + 1.0f != 33333334.0f
Here, the value 33,333,333 is rounded down to 33,333,332. Add 1 to that and you get 33,333,333, which is again rounded down to 33,333,332.
33,333,334 is represented precisely, so the comparison fails.
https://en.wikipedia.org/wiki/Single-precision_floating-poin...