9999999999999999.0 – 9999999999999998.0
geocar.sdf1.org
geocar.sdf1.org
The fact is that IEEE 754 is an exceptionally good way to approximate the reals in computers with a minimum number of problems or surprises. People who don't appreciate this should try to do math in fixed point to gain some insight into how little you have to think about doing math in floating point.
This isn't to say there aren't issues with IEEE 754 - of course there are. Catastrophic cancellation and friends are not fun, and there are some criticisms to be made with how FP exceptions are usually exposed, but these are pretty small problems considering the problem is to fit the reals into 64/32/16 bits and have fast math.
Floating-point numbers (and IEEE-754 in particular) are a good solution to this problem, but is it the right problem?
I think the "minimum of surprises" part isn't true. Many programmers develop incorrect mental models when starting to program, and get no feedback to correct them until much later (when they get surprised).
It is true that for the problem you mentioned, IEEE 754 is a good tradeoff (though Gustafson has some interesting ideas with “unums”: https://web.stanford.edu/class/ee380/Abstracts/170201-slides... / http://johngustafson.net/unums.html / https://en.wikipedia.org/w/index.php?title=Unum_(number_form... ). But many programmers do not realize how they are approximating, and the "fixed number of bits" may not be a strict requirement in many cases. (For example, languages that have arbitrary precision integers by default don't seem to suffer for it overall, relative to those that have 32-bit or 64-bit integers.)
Even without moving away from the IEEE-754 standard, there are ways languages could be designed to minimize surprises. A couple of crazy ideas: Imagine if typing the literal 0.1 into a program gave an error or warning saying it cannot be represented exactly and has been approximated to 0.100000000000000005551, and one had to type "~0.1" or "nearest(0.1)" or add something at the top of the program to suppress such errors/warnings. At a very slight cost, one gives more feedback to the user to either fix their mental model or switch to a more appropriate type for their application. Similarly if the default print/to-string on a float showed ranges (e.g. printing the single-precision float corresponding to 0.1, namely 0.100000001490116119385, would show "between 0.09999999776482582 and 0.10000000521540642" or whatever) and one had to do an extra step or add something to the top of the program to get the shortest approximation ("0.1").
I'm just a sample size of one, but isn't this kind of class a requirement for CS majors?
I don't know what's taught to CS majors, and as others have pointed out, programmers don't necessarily study CS.
I believe the issue is just that the pitfalls of floating point are not apparent without a certain level of math education.
But there may be one more pitfall, which is that those of us using FP regularly, also happen to be "scientific" or "exploratory" programmers who haven't learned a lot of formal software engineering discipline (including me). So we understand the math but might be more prone to making mistakes with it.
I do kind of like the idea of flagging any number that is potentially exposed to a FP issue. We all make mistakes. Displaying all floats in exponential notation by default would be a good enough warning to the wary. We only display them as decimals for readability.
I'm a little amazed in retrospect that the Math department was where one had to go to get a class in the mechanical details of computing when as a subject it's usually considered (and often in practice is) notably up the ladder of abstraction. And this was a weird outlier as the single most practical upper division class offered by the Math department at the time...
I think most CS departments dropped numerical analysis from their requirements by the end of 1980s. Nowadays you are more likely to find such a course in some dusty corner of math or engineering departments.
The problem is more that we don't have the tools to track and understand how the errors are altered as we do the math (e.g. how would you even begin to try representing catastrophic cancellation at compile-time) and doing the numerical error analysis on the abstract math itself is hard once the math gets complex let alone trying to figure it out after you've optimized the code for performance & tweaked the algorithms for real-world data/discrete space.
Now perhaps it could be possible to do it at runtime in some way but I suspect the performance of that is prohibitive to the point where arbitrary precision math or decimal numbers is going to be a better solution.
> My reasoning was based on the requirements of a mass market: A lot of code involving a little floating-point will be written by many people who have never attended my (nor anyone else's) numerical analysis classes. We had to enhance the likelihood that their programs would get correct results. At the same time we had to ensure that people who really are expert in floating-point could write portable software and prove that it worked, since so many of us would have to rely upon it. There were a lot of almost conflicting requirements on the way to a balanced design.
I imagine that the number of people writing code without having taken a numerical methods class has only increased since the late 1970s being talked about, or even in the two decades since that interview.
https://en.wikipedia.org/wiki/Unum_(number_format)
http://www.johngustafson.net/pdfs/BeatingFloatingPoint.pdf
(They are definitely not any easier to understand than the floating point system, though.)
Floating point in base 10 is already in the standard since 2008:
https://en.m.wikipedia.org/wiki/Decimal_floating_point
The standard is not to blame, the lack of demand for that feature is.
Most of the potential users don’t know that they can demand from their software and hardware suppliers that feature. Using it there would be less “surprises.”
Nevertheless, you have a good point and what I take away from this is that showing a range for the end result of a computation (instead for a given number directly entered by the user/programmer) can be misleading if the result of the exact computation wouldn't actually have been in that range.
It's time to look at other ways to depict fractional parts of numbers in a computer. I know that one can express any rational number as a integer fraction. And our computers are incapable of expressing a irrational number exactly - it does so to a certain precision... In other words, every number a computer expresses is a rational number.
The exception is if the computer could express irrational numbers as symbolics, then we could work with the symbolic instead. And then as a last pass, the symbolic could convert to a imprecise rational depiction, or express as its native type.
(0) https://itsfoss.com/a-floating-point-error-that-caused-a-dam...
I would say the fact that floats can have data larger than INT16_MAX is hardly a surprise. That was just a bug, not some great and surprising drawback of floating point.
We recently had an exponential memory growth bug because rationals can not always be simplified, for example if you start with (2/3) and repeatedly square it. Fortunately, this was not in user-facing code, so there was no chance of a denial-of-service attack, but that's definitely something to watch out for with rationals.
Floating point numbers are the optimal minimum message length method of representing reals with an improper Jeffery's prior distribution. A Jeffery's prior is a prior that is invariant under reparameterization, which is a mandatory property for approximating the reals.
In this case, it is where Prob(log(|x|)) is proportional to a constant.
Thus, we aren't going to ever do better than floats if we are programming on physical computers that exist in this universe. There is a reason why all numerical code uses them. Best to learn their limitations if you are going to use them, otherwise use arbitrary precision.
Most logic is multiplicative. For example, apply a 30% tax on a dollar quantity and display both subtotal and grand total. With floats, there are inequalities. With decimal there usually aren't unless you're dividing, but we already have to deal with divide errors in base ten, and it is much more likely to need to represent 0.30 than 1/3 and because decimal shares a base with binary (since it's factors are 5 and 2) binary doesn't really get us anything but headaches anyway. It's true that there are still gotchas, but they happen less often and usually don't end up looking stupid and weird for no reason. That 0.1 + 0.2 = 0.300000000000001 is dumb and we all know it.
Is “academic world” now a shorthand for “all numerical computing”?
Decimals basically never make sense, except possibly in some calculations related to money. Those make up a minuscule part of modern computer use.
Maybe decimals are also better for homework assignments for schoolchildren?
The type of applications where decimals are useful are by and large insensitive to compute speed and need no special hardware support. You can easily write your code for decimal arithmetic on top of integer arithmetic hardware.
Those of us who need binary floating point for graphics, audio, games, engineering, science, .... won’t stop you.
Lots of applications don't need that, but a surprising amount do, so it's not a global solution.
In some situations, there may even be industry or legal standards as to what can be considered rounding noise.
Citation needed, because this isn't really true.
Even if one concedes your (unspoken) idea that only financial transactions aren't "academic" (which also isn't true), in the real world financial transactions will typically include currency conversions, and those will have all sorts of weird non-decimal factors.
Floating point makes them all much easier to do well.
We chose to _not_ do _any_ calculations client side in Javascript...
How did you use ints to compute compounded interest on loans? I asked that above, and you avoided it. I ask again.
For example, suppose you have a mortgage where you lent $100K at 5% annual, compounded 12 times a year, for 30 years, and you need basic values regarding this loan.
Often in such calculations you need to compute 100K*(1+0.05/12)^360. How do you do that with integers? Naively you need (1+(1/20)/12)^360, which as a reduced fraction each of the numerator and denominator have over 2800 binary digits. Do you really do this with integers?
Now put that in a mortgage trading or pricing system where it needs to do millions/billions of those per second.
Doing this as double gives enough precision to make the difference between computed and infinitely precise negligible (approx. 10^-17 error).
It's easy to make examples where doing incremental calculations, rounding to pennies and storing, results in long term error. In these cases I don't see how do to it with integers without massive overhead.
And this is a trivial, common example. Doing stuff like hedge fund stuff, or anything using numerical integration to make models for pricing, would be astounding hardly to do with integer only math.
What finance software did you write? A simple ledger works fine as integers. Anything more complex will hit performance and scaling issues soon after the basics.
Currency is stored as a count of cents (millicents if being fancy). Therefore the two main features of floats are not useful:
- Support for very small numbers is not needed. Floats dedicate approx half their range to numbers between -1 and +1, this is wasted when counting whole cents.
- Support for very large numbers at the expense of precision is actively bad, as the precision must always be down to individual cents.
So the useful range of floats is much reduced when using floats for counting, approx 54 bits out of a 64 bit float are used. Instead ints (“counting numbers”) are much better for counting cents than floats (which approximate the continuous real numbers in a finite number of bits).
Floats don't have the precision required.
IBM's new and rising supercomputer architecture, POWER9, supports hardware IEEE binary128 floats (quad precision). Their press claims the current fastest supercomputer in the world uses POWER9.
The ppc64 architecture (still produced by IBM) supports "double-double" precision for the long-double type, which is a bit hacky and software-defined, but has 106 bit mantissa.
And ARM's aarch64 architecture supports IEEE binary128 long-doubles as well, though it is implemented in software now (by compiler). Maybe they plan a hardware implementation in the future?
The x87 instructions are basically for legacy compatibility, or if you manually use long doubles on some platforms.
The idea behind extended precision registers was good in theory, but ultimately caused too much hassle in practice.
(Sorry, this is more for the folks who aren't familiar with this, since it seems like you are familiar, but I didn't want it to seem like this isn't widely supported when they read "legacy" or "some platforms")
Here is a good toy examples that runs into the same numbers shown in the parent, showing the two different instruction types, and that long double can give you the correct answer, while still being run in hardware, vs. going all the way to float128s which are currently emulated in software!
Code w/ Assembly: https://godbolt.org/z/W3ZmqJ Output: https://onlinegdb.com/Sy_I3Q1ME
The current hardware trends seem to be providing instructions for compensated arithmetic, like FMA and "2sum" operations. I think this is ultimately a better solution, and will make it possible to give finer control of accuracy (though there will still be challenges on the software/language side of how to make use of them).
https://en.cppreference.com/w/cpp/language/types
https://software.intel.com/en-us/articles/size-of-long-integ...
They are definitely still supported in modern Intel processors. That said, there can be some confusion because they end up being padded to 16 bytes for alignment reasons, so take 128 bits of memory, but they are still only 10 byte types.
They are a distinct type from the "quad" precision float128 type, which is software emulated as you mentioned.
All that being said, you are right that most of the time float math ends up in SSE style instructions, but as soon as you add long doubles to the mix, the compiler will emit x87 style float instructions to gain the extra precision.
Example: https://godbolt.org/z/PMZVdb
Rightly so, because programmers want their optimizing compiler to decide when to put a variable on the stack and when to elide a store/load cycle by keeping it in a register. With 80 bit precision, this makes a semantic difference and you end up in volatile hell.
EDIT: If I had some application where I needed the extended range, like maybe I was going to run into the exact numbers above, I'd appreciate the ability to opt-in to this. Totally agree I wouldn't want the compiler to surprise me with it, but also not terrible, or useless.
Code w/ Assembly: https://godbolt.org/z/W3ZmqJ Output: https://onlinegdb.com/Sy_I3Q1ME
In that case, the extra precision is present and "carried across" operation when registers or the dedicated floating point stack is used, but is discarded when values are stored to a narrower 64-bit location. So the problem is one really of mismatch between the language/ABI size and the supported hardware size. Of course, 80 bits isn't a popular floating point size any more in modern languages, so this happens a lot.
Well x86-64 is not binary compatible with x86 so that's not the reason. It is mostly for software relying on either the rounding quirks or the extended 80 bit precision I guess.
> However the actual x87 registers are shadowed by the vector registers so you can only use one
You are confusing with the legacy MMX registers which are deader than the x87 for stack. XMM registers do not shadow the for stack.
(It’s inevitably floats or doubles)
Do you have a link to a proof or discussion of this? I haven't heard this before and I would love to have this statement unpacked a little more.
Maybe, but that doesn't mean the particular implementation of floats being used is the best one. See also: Unums and Posits
Not quite. The difference in computation time is the current the lack of hardware support, not something inherent to the underlying encoding method. So in practice you are right, but in, for example, embedded contexts without floating point hardware, the performance advantages of IEEE floats should disappear (especially if using a 16 or 8 bit posit suffices).
Posits are simpler to implement than IEEE floats (less edge cases) and use more bits for actual numbers whereas IEEE floats waste about half on NaNs. The use of tapered precision is also nice.
In my defense, the comment he replied to got downvoted and I thought it was nestorD, so I was "primed" to misinterpret his comment as criticizing unums in general.
> I can’t understand how they think that’s going to be possible in a finite 64 bits.
You apparently stole 32 of them to make your bat.
If you put them back your tests balloon to half a century each.
"You can rent a Skylake chip on Google Cloud that'll perform 1.6 trillion 64 bit operations per second for $0.96/hr preemptively. That's enough to run one instruction over a 64 bit address space exhaustively over 120 days, or for ~$2800"
It might not make economic sence to actually make this happen for any realistic test, but it's interesting that it might actually be feasible to do it on any kind of human timescale...
I think most unit tests are best served by testing key values- e.g. values before and after any intended behavior change, values that represent min/max possible values, values indicative of typical use.
The unit test can serve as documentation of what the code is intended to do, and meaninglessly invoking every unit test over the range of floats obscures that.
There are certainly cases where all values should be tested, but I don't think that's all cases.
I presume they were saying 'and for 32-bit floats you also get this property that you can...'
You of course need an unbounded but finite amount of space to store these numbers, which is perfectly fine.
I don't think that's really quite true. The point of FP is that you don't get any wierd statefulness in your compute complexity as values accumulate, every operation basically has O(1) compute time where N is the number of previous operations you've done. For rationals and algebraics that isn't the case.
Very far from a floating point expert here, but what I do is to scale-down by a few odd prime-power factors as appropriate:
Scaling down by powers of 5 is obviously appropriate for decimals, currency etc.
Scaling down by powers of 3 is good for angles measured in the degrees, minutes, seconds system.
If one scales down a lot there is an increased risk of overflow, so one can compensate by scaling up some powers of 2.
The way I think of this is as using my own manual exponent bias [0].
>the exponent is stored in the range 1 .. 254 (0 and 255 have special meanings), and is interpreted by subtracting the bias for an 8-bit exponent (127) to get an exponent value in the range −126 .. +127.
So, for example, even single-precision number are always exact multiples of 1/(2^126), and I'm just changing the denominator to contain powers of 3, 5, 7, ... etc.
In financial calculations I've seen, figures are given in standard magnitudes (per cent, per mille, basis points, integer cents, etc.) which, if you're lucky with your language, can be encoded as types which can be promoted to higher precision (somewhat) transparently.
> I appreciate that it may be surprising that 0.1 + 0.2 != 0.3 at first, or that many people are not educated about floating point, but I don't understand the people who "understand" floating point and continue to criticize it for the 0.1 + 0.2 "problem."
That's not a calculation that should require a high level of precision.
Decimal fp is still 'wrong' for, say, 1/3 + 1/3 = 2/3.
Maybe we really should move back to base-60 like the Babylonians used, then you could at least divide by 3.
The argument "we want to look at base-10 in the end so it should be the internal representation" is really weak and ignores basically every other practical aspect.
The only example off the top of my head that is floating point is C# "decimal", which actually originates from the Decimal data type in OLE Automation object model (which could be seen in VB6, and can still be seen in VBA):
https://msdn.microsoft.com/en-us/library/cc237603.aspx
Note this bit:
"scale: MUST be the power of 10 by which to divide the 96-bit integer represented by Hi32 * 2^64 + Lo64. The value MUST be in the range of 0 to 28, inclusive."
The reason why it's limited to 28 is because the 96-bit mantissa can represent up to 28 decimal digits exactly. The way it's enforced, any operation that produces a result outside of this range is an overflow error (exception in .NET).
They don't - all 11 bits of the exponent (for float64) are in use, so you can have something like 1e300, and then you can't e.g. add 1 to it and get a different number.
>>> x = 1e100
>>> x
1e+100
>>> y = x + 1
>>> y
1e+100
>>> x - y
0.0 if x != y, x - y != 0
which holds for all x and y, but is different to if dx != 0, x + dx != x
which fails for some (many!) x and dx.9999999999999995.0 – 9999999999999994.0 == 10.0000
Filtered by languages I care about, I guess I have no choice but to learn perl 6 if I want correct (but presumably slow) floating point with elegant syntax (my taste might not match yours).
I’d be curious to know what the random GPU languages and new vector instruction sets do with this computation. I don’t think they’re all 754 compliant.
No surprises.
You answered your question. 99% of the time being exact is a requirement and calculation speed is utterly unimportant, thus using IEEE 754 results in programs that are fundamentally broken.
For example, entering 9999999999999999.0 into "double" gives https://float.exposed/0x4341c37937e08000 and entering 9999999999999998.0 gives https://float.exposed/0x4341c37937e07fff
My wishlist for such a page would contain two additional features:
1. Allow entering expressions like "a OP b == c", so that one can enter "0.1 + 0.2 == 0.3" or "9999999999999999.0 - 9999999999999998.0 == 1.0" and see the terms on the left-hand side and right-hand side.
2. Show for each float the explicit actual range of real numbers that will be represented by that float. For example, show that every real number in the range [9999999999999999, 10000000000000001] is represented by 10000000000000000, and that every real number in the range (9999999999999997, 9999999999999999) is represented by 9999999999999998.
The author of this one has a blog post about it: https://ciechanow.ski/exposing-floating-point/ and I also like a shorter (unrelated) page that nicely explains the tradeoffs involved in floating-point representations and the IEEE 754 standard, by usefully starting with an 8-bit format: http://www.toves.org/books/float/
Nice that it reformats the input to "10000000000000000.0", gets the point across that a 64 bit double float just doesn't have enough bits to exactly represent 9999999999999999.0, but that it does happen to be able to represent 9999999999999998.0.
IEEE 754 64-bit floats have 53 significant bits ("mantissa").
Try playing around with half precision, it makes things a lot easier to understand.
9999999999999998.0 in IEEE754 is 0x4341C37937E07FFF
"9999999999999999.0" in IEEE754 is 0x4341C37937E08000 - the significand is exactly one higher.
With an exponent of 53, the ULP is 2 - so parsing "9999999999999999.0" returns 1.0E16 because it's the next representable number.
Using one of these workarounds requires a certain prescience of the
data domain, so they were not generally considered for the table above.
Doing arithmetic reliably with fixed-precision arithmetic always requires understanding of the data domain. If you need arbitrary precision, you'll need to pay the overhead costs of arbitrary-precision: either by opting-in by using the right library, or by default in languages like Perl6 and Wolfram.If you want arbitrary precision, use an arbitrary precision datatype. If you use fixed precision, you'll need to know how those floats work.
Pointless article, imho.
No, I don't think so. Where does that come from? The page doesn't mention FP standards at all.
> If you want arbitrary precision, use an arbitrary precision datatype.
That's the point. Half of them don't offer this feature. The other half make it very awkward, and not the default.
We went through this exercise years ago with integers. These days, there are basically two types of languages. Languages which aim for usability first (like Python and Ruby), which use bigints by default, and languages which aim for performance first (like C++ and Swift), which use fixints by default. It's even somewhat similar with strings: the Rubys and Pythons of the world use Unicode everywhere, even though it's slower. No static limits.
With real numbers, we're in a weird middle ground where every language still uses fixnums by default, even those which aim for usability over performance, and which don't have any other static limits encoded in the language. It's a strange inconsistency.
I predict that in 10 years, we'll look back on this inconsistency the same way we now look back on early versions of today's languages where bigints needed special syntax.
I'm sorry you thought so. It pops up pretty often and always seems to spark a lot of conversation, so I think most programmers that give it any thought can find it a very interesting area of study.
There's an incredible amount of creep: We have what starts with nice notation (like x-y) and have to trade a (massively increased) load in either our minds or in the heat our computer generates. I don't think that's right, and I think the language we use can help us do better.
> What is the "right answer"?
What do you think it is?
Everyone wants the punchline, but this isn't a riddle, and if this problem had a simple answer I suspect everyone would do it. Languages are trying different things here: Keeping access to that specialised subtraction hardware is valuable, but our brains are expensive too. We see source-code-characters, lexicographically similar but with wildly differing internals. We want the simplest possible notation and we want access to the fastest possible results. It doesn't seem like we can have it all, does it?
If you subtract two numbers close to each other with fixed precision you don’t know what the revealed digits are. (1000 +/- .5) - (999 +/- .5) = 1 +/- 1.
Thus 0, 1, and 2 are all within the correct range.
Floating point numbers have X digits of accuracy based on the format. (Using base 10 for simplicity) Let’s say .100 to .999 times 10^x.
But what happens when you have .123x10^3 - .100x10^3. It’s .23? x 10^2 but what is that ? we might prefer to pick 0 but it really could be anything. We can’t even be sure about the 3. If the numbers where .1226 x 10^3 and .1004 x 10^3 that just got rounded the correct number would be .222 x 10^2
Y However, you aren't going to do any better without using vastly more expensive arbitrary precision.
For example, CPU integers aren't like mathematical integers. CPU integers wrap around. So CPU integers aren't "really" the integers—CPU integers are actually the ring of integers modulo 2n , with their names changed!
I'm not sure what the name of the ring(?) that contains all the IEEE754 floating-point numbers and their relations is called, but it certainly exists.
And, rather than thinking of yourself as imprecisely computing on the reals, you can think of what you're doing as exact computation on members of the IEEE754 field-object—a field-object where 9999999999999999.0 - 9999999999999998.0 being anything other than 2.0 would be incorrect. Even though the answer, in the reals, is 1.0.
By all means embrace the surprise and educate today's 10,000, by why not actually explain why these are reasonable answers and the mechanics here behind the scenes?
Thank you for the relevant xkcd!
If we're going to try to find a "right" answer from a language view without knowing the exact program and use cases then the most reasonable compromise is likely "error" because types weren't specified on the constants or parsing functions.
The appropriate default is, I would argue, the one which preserves the mathematically correct answer (as close as possible) in the majority of cases and enables coders to override the default behavior if they want to specify the exact underlying numerical representation they desire (instead of it being automatic). That goes along with the "principle of least surprise" which is always a good de facto starting point for any human/computer interaction.
Consider a similar example, pointers. Some languages (like C and C++) use pointers heavily and it's expected that devs using those languages will be experienced with them. However, pointers are very "sharp" tools and have to be used exceedingly carefully to avoid creating programs with major defects (crashes, memory leaks, vulnerabilities, etc.) They are so hard to get right that even software written by the best coders in the world commonly has major defects in it related to pointer use. This problem is so troubling to some that there are many languages (java, javascript, python, C#, rust, etc.) which have been designed to avoid a lot of the most difficult to use aspects of languages like C and C++, they use garbage collection for memory management, they discourage you from using pointers directly, and so on. However, even those languages do very little to protect the user from blundering into a mindfield of floating point math.
Consider, for example, simply this statement:
x = 9999999999999999.0
Seems rather straightforward, right? But it's not, it's a lie. Because in many languages the value of x won't be as above, it'll be (to one decimal digit precision) 10000000000000000.0 instead. Whereas the value of ....98.0 is the same as the double precision float representation to one decimal digit precision (thus the difference between the two comes out as 2.0 instead of 1.0). Now, maybe in a "the handle is also a knife" language like C this is fine, but we have so many languages which go to such extremes everywhere else to protect the user from hurting themselves except when it comes to floating point math. And here's a perfect case where the compiler, runtime, or IDE could toss an error or a warning. Here you have a perfect example of trying to tell the language something you want which it can't do for you in the way you've written, that sounds like an error to me. The string representation of this number implies that you want a precision of at least the 1's place in the decimal representation, and possibly down to tenths. If that's not possible, then it would be helpful for the toolchain you're using for development to tell you that's impossible as close to you doing it as possible, so that you know what's actually going on under the hood and the limitations involved.
Something which would also drive developers towards actually learning the limitations of floating point numbers closer to when they start using them in potentially dangerous ways than instead of having to learn by fumbling around and finding all the sharp edges in the dark. The sharp edges are known already, tools should help you find and avoid them not help new developers run into them again and again.
I'm a little concerned if merely knowing the existence of floating point arithmetic constitutes "prescience."
Mathematica. But it's not particularely fast.
> Unless you're using numbers that scale from very small to very large, like 3d games or scientific calculations, you don't actually want to use floating point.
Unfortunately, we can sometimes only use floats in 3D graphics and floats aren't even good for semi-large to large 3D scenes. Unity is a particular bad offender. It's not even necessary for meshes but having double precision transformation matrices would make life so much easier. Could simply use double precision world and view matrices, then multiply them together and the large terms would cancel out in the resulting worldView matrix, which can then by cast back to single precision floats.
I'm not aware of any language (other than Wolfram) that defaults to storing something like 0.1 as 1/10 - i.e. uses the decimal constant notation for rationals, rather than having some secondary syntax or library.
In[1]:= Precision[0.1]
Out[1]= MachinePrecision
In[2]:= Precision[1/10]
Out[2]= \[Infinity]I don't currently have a license, so out of curiosity is 0.3 == 3/10 in wolfram?
According to the documentation of Equal†,
> Approximate numbers with machine precision or higher are considered equal if they differ in at most their last seven binary digits (roughly their last two decimal digits).
Which is why in Mathematica, 0.1+0.2==0.3 is also True.
If you need a kind of equality comparison that returns False for 0.3 and 3/10, use SameQ. Funnily, SameQ[0.1+0.2,0.3] is also True, because SameQ allows two machine precision numbers to differ in their last binary digit.
Casting to bigint doesn't work because the problem occurs when converting the decimal constant in the source to floating point. You would have to convince the parser to parse the constant as something besides a float.
Perhaps it is time for the developers of new programming languages to consider using a different approach to representing approximations to real numbers, for example something like the General Decimal Arithmetic Specification [2], and to relegate fixed-size binary floating-point numbers to a library for use by experts.
There is an analogy with integers: historically, languages like C provided fixed-size binary integers with wrap-around or undefined behaviour on overflow, but with experience we recognise that these are a poor compromise, responsible for many bugs, and suitable only for careful use by experts. Modern languages with arbitrary-precision integers are much easier to write reliable programs in.
[1] https://en.wikipedia.org/wiki/Kahan_summation_algorithm [2] http://speleotrove.com/decimal/decarith.html
most popular previous discussion: https://news.ycombinator.com/item?id=10558871
That said, even in Common Lisp I think its only CLISP (among the free implementations) that gives the correct answer for long floats.
CLISP:
[1]> (- 9999999999999999.0L0 9999999999999998.0L0)
1.0L0
SBCL, CMUCL and Clozure CL: * (- 9999999999999999.0L0 9999999999999998.0L0)
2.0d0
The standard only mandates a minimum precision of 50 bits for both double and long floats, so there's no guarantee that using long floats will give the correct answer, as we can see.http://www.lispworks.com/documentation/HyperSpec/Body/t_shor...
https://stackoverflow.com/questions/588004/is-floating-point...
No, it's just that a lot of people don't understand its limitations.
https://cacm.acm.org/magazines/2017/8/219594-small-data-comp...
9999999999999999 – 9999999999999990
9999999999999999 - 9999999999999971 == 0
9999999999999999 - 9999999999999970 == 30
9999999999999999 - 9999999999999969 == 32
9999999999999999 - 9999999999999966 == 34 #include <boost/multiprecision/cpp_dec_float.hpp>
#include <boost/lexical_cast.hpp>
#include <iostream>
using fl50 = boost::multiprecision::cpp_dec_float_50;
int main() {
auto a = boost::lexical_cast<fl50>("9999999999999999.7");
auto b = boost::lexical_cast<fl50>("9999999999999998.5");
std::cout << (a - b) << "\n";
}
works int main() {
fl50 a = 9999999999999999.7;
fl50 b = 9999999999999998.5;
std::cout << (a - b) << "\n";
}
doesn't, even if you change fl50 out for a quad precision binary float type.Note that in your code sample you're not actually using user-defined literals (https://en.cppreference.com/w/cpp/language/user_literal). This works (based on on your earlier code sample and adding user-defined literals):
#include <boost/multiprecision/cpp_dec_float.hpp>
#include <boost/lexical_cast.hpp>
#include <iostream>
using fl50 = boost::multiprecision::cpp_dec_float_50;
fl50 operator"" _w(const char* s) { return boost::lexical_cast<fl50>(s); }
int main() {
fl50 a = 9999999999999999.7_w;
fl50 b = 9999999999999998.5_w;
std::cout << (a - b) << "\n";
}printf("%Lf\n", 9999999999999999.0L - 9999999999999998.0L);
In my x86_64 computer it breaks when you add enough digits. At this point it started outputting 0.0 as the difference:
printf("%Lf\n", 99999999999999999999.0L - 99999999999999999998.0L);
With 63 bits for the fraction part you more or less get around 19 decimal digits of precision, and the expression above uses 20 significant digits.
[1] https://en.wikipedia.org/wiki/Extended_precision#x86_extende...
I guess this comes down to most of them having implementations of arbitrary precision decimals.
If you stored your values in table rows as DOUBLE PRECISION, you would of course get the wrong answer.
Python 2.7.3 (default, Oct 26 2016, 21:01:49)
[GCC 4.6.3] on linux2
Type "help", "copyright", "credits" or "license" for more
information.
>>> from decimal import *
>>> getcontext().prec
28
>>> a=Decimal(9999999999999999.0)
>>> b=Decimal(9999999999999998.0)
>>> a-b
Decimal('2')
That is unexpected. >>> a = 9999999999999999.0
>>> a
1e+16 >>> Decimal('9999999999999999.0')-Decimal('9999999999999998.0')
Decimal('1.0')A specialist number representation is made for exact representation of values in geometric calculations (think CAD). Numbers are represented as sums of rational multiples of cos(iπ/2n).
Exact summation, multiplication and division (not shown) of these quantities are possible, and certain edge-cases (eg. sqrt) have special-case handling.
The system was integrated into and tested on an existing codebase.
The speaker was also one of the authors of Herbie, if other people remember that.
PHP 7.2.10 (cli) (built: Oct 9 2018 14:56:43) ( NTS )
$ php -r "echo 9999999999999999.0 - 9999999999999998.0;"
2
$ php -r "echo bcsub('9999999999999999.0', '9999999999999998.0', 1);"
1.0
5.55 * 1.5 = 8.3249999999999....
26.93 * 3 = 80.7899999999999....
I raised it with the supplier some time ago, they said it's just the calculator app and the main program isn't affected. Quite shocking that they are happy to leave it like this.
Why?
My thoughts: 1) more inefficient programs because encountering an arbitrary-precision expression requires arbitrarily large memory and computation, 2) more complicated language implementation.
The dangerous bit is, that just extracting a variable from constant expressions might change the result slightly. That should not be a problem, unless you are depending on exact values.
Hint 1: Can you imagine ever moving a magic number from an expression into a constant?
import Data.Scientific
(9999999999999999.0 :: Scientific) - (9999999999999998.0 :: Scientific)
Of course for others it should "fixable" where needed:
~ python3
Python 3.6.7 (default, Oct 22 2018, 11:32:17)
[GCC 8.2.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> from decimal import Decimal
>>> Decimal('9999999999999999.0') - Decimal('9999999999999998.0')
Decimal('1.0')But I don't think such a warning would be all that helpful. How often do we use literals with 16 significant digits, expecting exact representation?
The bigger gotcha here is catastrophic cancelation. This is the issue of an insignificant rounding error becoming much more significant due to subtracting of very nearly equal numbers. You can't generally detect this at compile time if you don't know all your numbers in advance (e.g. you're not working with only literals).
Well, the perl6 result surprised me, since that means it's using something more precise than double precision floating point :)
$ perl6 --version
This is Rakudo version 2018.03 built on MoarVM version 2018.03
implementing Perl 6.c.
$ perl6 -e 'print 9999999999999999.0-9999999999999998.0;print "\n";'
2
$
(Incidentally, I would have used "say" rather than "print" with an explicit newline.) https://github.com/rakudo/rakudo/commit/fec1bd74f97e257d4c88673cd62fdcae39f587a3 › perl6 -e '.say for $*PERL.compiler, 9999999999999999.0-9999999999999998.0'
rakudo (2018.11)
1Iv'e seen article after article of how "horrible it is". So, are there default libs to use Binary Coded Decimal (BCD) or something like that?
We use it because it's fantastic.
64 bit integers are actually good for a lot of things that floating point gets used for; coordinates already should never have been floating point (the accuracy of measurement is independent of the magnitude for coordinates), but you could represent the entire solar system in millimeters without overflowing 64 bit integers (compared to not even the entire earth in mm for 32 bit integers).
Also, it's popular to blame javascript for all today's problems, but the double-precision floating point is the only number type it has, which has some effect on its use.
bc is doing its job!
1> let foo = 9999999999999999.0 - 9999999999999998.0
foo: Double = 2In this case,
9999999999999999.0`17-9999999999999998.0`17 does indeed return 1.
perl -Mbignum -e'print 9999999999999999.0 - 9999999999999998.0'
1
The BigFloat solution on the website is suboptimal. perl6 -e'print 9999999999999999.0 - 9999999999999998.0'
1I guess the first link is converting the float constants to ints at compile-time?
(edit: oh, it's actually mentioned in the article. I should read more carefully)
Plenty of languages are going to get upset if you add 2 billion to 2 billion.
Whenever I see these examples I do get annoyed st the Haskell one because we never are told what type it gets defaulted to, which only happens silently in the ghci repl, but will trigger a warning if it’s in a source file that’s being compiled.
MariaDB [(none)]> SELECT 9999999999999998.0 - 9999999999999999.0; -1.0
writeln(FloatToStr(9999999999999999.0 - 9999999999999998.0));
1
println(9999999999999999.0-9999999999999998.0) 1.0
But oh well.
Perhaps inflation will fix it. Or the first quadrillionaire AI.
If you want more than 53 significant bits, you need something wider than a 64-bit IEEE double. And anyone who cares about such precision knows this and will expect to use a higher precision library. (Or is an incredibly specialized math language, like Wolfram.)
If you think having only 53 significant bits is bad, you should see how many bits common trigonometry operations lose in libm implementations.
Another fun example: video games (historically) have thrown away even 53 significant bits for the higher performing 32-bit floats.