Java’s floating-point hurts everyone everywhere (1998) [pdf]
people.eecs.berkeley.edu
people.eecs.berkeley.edu
The main author of this report, W. Kahan, was also the original author of IEEE-754. He was (and is) strongly unsatisfied on the state of floating-point in practical systems, and it was not his first time to complain about these problems. One can find his most recent critique from a few years ago.
My understanding is that, Kahan's IEEE-754 is meant to:
1. The use of double and extended precision should be encouraged to safeguard non-experts from floating-point errors. Everything should be at least double precision by default, and extended precision serves as an additional safeguard. In fact, when IEEE-754 was being drafted, Kahan believed 128-bit floating point should also be supported as computers become more powerful in the future. The computational cost was too high at that time, so he settled on 80-bit as seen on the 8087. He criticized Java for not supporting it.
2. Floating-point exceptions should be used and turned on everywhere to safeguard programmers from making mistakes, and possibly to allow programs to handle them as special cases in the logic at runtime. He criticized Java for not supporting it.
3. Unsafe optimizations should not be done, such as automatically using FMA, or using algebraic identity in compiler optimization. He criticized Java for allowing it in some cases.
Unfortunately, as far as I can see, these ideals of IEEE-754 have all but largely disappeared in real-world applications since then, for various practical reasons.
The industry did not move to 128-bit floating point because its performance overhead is too large, the original assessment of IEEE-754 in the 1980s was too optimistic. Similarly, the industry did not accept 80-bit extended precision as a standard but saw it as an oddity of the Intel 8087, even Intel has abandoned it - x86_64's SSE or AVX has removed all support of that, making it an exclusive feature of obsolete i386/i686 machines. My impression is that anything above double precision is no longer used in the industry (IBM POWER does have native quad-precision support, to their credit).
A secondary overhead of higher-precision floats is the memory wall, arguably a more serious problem today. Memory bandwidth has become the most serious overhead for many numerical programs. Modern computers have a machine balance of 100:1, it means you need to do as many as 100 floating-point operations after loading a single value from RAM to reach the machine's peak performance. But this is not compatible with many algorithms with an inherently low arithmetic intensity, including important physics simulations. The use of of FP80 or FP128 will make them unacceptably slower. As a result, today's trend is moving from FP64 to FP32, and even to FP16 or a custom 16-bit format if possible, not vice versa.
The use of floating-point exceptions similarly became unpopular because of performance problems. I'm not an expert on this, but it was my impression that signaling a floating-point exception was so expensive on both hardware and operation-system level that it was never seriously used in most practical programs. So in contrary to IEEE-754, these exceptions never became an integral part of the programming environment. In addition to an expensive Unix signal, potential exceptions also inhibit the efficient pipelining of modern out-of-order, superscalar CPUs.
Finally, the use of unsafe optimizations is prevalent in many applications when rigor is sacrificed for speed.
So overall, the original spirit of IEEE-754 was long gone - for better or worse, unfortunately.
[1] https://people.eecs.berkeley.edu/~wkahan/ieee754status/754st...
Honestly I don't understand how this would constitute a "solution" or even a "safeguard", really. Using any kind of FP arithmetics without being acutely aware of its quirks is going to cause headache no matter the precision. Conversely, when you do know how to deal with FP, then FP32 can suit many scenarios just fine.
Floating point exceptions worked well enough in Delphi. The most common exception in practice is divide by zero and it's usually a programming error.
I agree with most of what you said, but automatic use of FMA is not an unsafe optimization. If FMA breaks your code, your code was always broken in the first place.
Automatic FMA is absolutely an unsafe optimization. Don't do it unless software requests it.
Crazy. Gosper and I spent several dinners trying to come up with a use for the range or precision of 128. We figured there had to be some classified application up at LLL because even at quantum & cosmological scales it didn't make sense.
It also seems like a huge overkill for avoiding accidental over/underflow problems by naive programmers.
As for processors, according to Wikipedia [4], "IBM POWER6 and newer POWER processors include DFP in hardware, as does the IBM System z9 (and later zSeries machines). ... Fujitsu also has 64-bit Sparc processors with DFP in hardware."
[1] https://en.cppreference.com/w/c/23
[2] https://gcc.gnu.org/onlinedocs/gcc/Decimal-Float.html
[3] https://learn.microsoft.com/en-us/dotnet/api/system.decimal?...
A flag is a type of global variable raised as a side-effect of exceptional floating-point operations. Also it can be sensed, saved, restored and lowered by a program. When raised it may, in some systems, serve an extra-linguistic diagnostic function by pointing to the first or last operation that raised it.
Any modern programmer knows that side effects are, while inevitable, hard to tame and some discipline is needed. Pure functional programming, mutable XOR shared, software transactional memory, you name it. This part of talk completely handwaves a difficulty of side effects and forces every language to be handcuffed with those global states and side effects. No good.
[1] Joe Darcy, one of the students of William Kahan, later went to Sun to improve fp support in Java; see https://web.archive.org/web/20090402234711/http://blogs.sun.... and https://web.archive.org/web/20100224062224/http://blogs.sun.... for example. But his work was more about library supports AFAIK.
As a numerical analyst, Kahan is pretty obsessed with use-as-much-precision-as-you-can. But there's a useful rule of thumb: you need about twice the amount of working precision as your final result. Since double precision has a 53-bit mantissa (~16 decimal digits), that means if you need only 8 or fewer decimal digits, you're completely fine with double precision. And furthermore, the experiences I've had with many programmers suggest that getting bit-equivalent results from different machines is a higher priority than squeezing the best possible numerics out of your hardware. HPC does tend to care about the latter a lot more, but that's also an area where the solution is almost always to just use your system's advanced math libraries (e.g., MKL for dense linear algebra).
Ironically, using your system’s math libraries will probably make replicating floating point results harder. Especially on macOS. Accelerate is a curse.
I disagree with your gloss of Kahan’s philosophy. His approach is more along the lines of “do not waste precision”. But this philosophy is not the complete truth; as close as I can state it briefly, my modification would be “do not waste precision that may be needed later”.
(Comic villain sitting in his fast-math lair) Foiled yet again!
[1] But really, if -ffast-math does turn -funsafe-math-optimizations on, it should have been named similarly. There is a possibility of much safer -ffast-math with almost zero breakage (by assuming a subset of IEEE 754, like the fixed rounding mode). The current -ffast-math is so reckless [2].
[2] https://simonbyrne.github.io/notes/fastmath/#flushing_subnor...
Plus, comparing against strict math as I go tends to highlight where I might have been about to do something dodgy anyway.
“The impetus for changing the default floating-point semantics of the platform in the late 1990's stemmed from a bad interaction between the original Java language and JVM semantics and some unfortunate peculiarities of the x87 floating-point co-processor instruction set of the popular x86 architecture.”
The nature of the beast is that as soon as you change the order of arithmetic you're going to get a different result. Optimized code is going to give you different results on different hardware due to the fact that you need to optimize things differently. Threading, memory alignment and/or different versions of the library software are likely to lead to different results even on the same machine unless the authors of the library go out of the way to promise repeatability.
(If you want to get the same answer, run on a single thread, page align everything you feed in, and never upgrade your system; alternatively write a scalar loop in C, compile with -O0 and pray the compiler doesn't change the order of things on its next upgrade).
We did it for Wasm, which follows IEEE-754 semantics exactly for 32-bit and 64-bit floats. (The only nondeterminism is the exact bit pattern you get for NaNs in some circumstances.) Rounding is 100% well-specified. And CPUs have done that for decades. Even vector ISAs have learned that non-IEEE results are not what software wants; all vector ISAs are converging on IEEE-754.
> Optimized code is going to give you different results on different hardware due to the fact that you need to optimize things differently.
This is due to C/C++ (and to some extent Fortran) semantics. It is not hardware.
What do threads have to do with floating point precision?
There’s also the weirdness that in C++ the floating point environment is thread-local, which can cause all sorts of chaos.
Different microarchitectures (e.g. how many vector instructions of what size need to be in flight for full occupancy), different numbers of cores (see threading discussion below) and often even differently aligned memory (does it need repacked or not for best performance?) will all require different order of operations to obtain maximum throughput, which means different (but equally valid) results.
For threading in particular if you want to get the same bit-exact answer, you end up constraining yourself to a particular ordering on reduction operations. This in turn either outright prevents techniques such as work-stealing or fires a very prescriptive reduction tree that itself constrains parallelism.
This is entirely driven by hardware and its impacts on performance of algorithms, and applies regardless of the language you're writing in if you want to obtain the best possible performance from a given chip.
As far as I know, two sectors claim they need it: finance and climate.
"Do you want a better answer?"
"No, I want the same wrong answer that I got last Tuesday."
Science/Mathematics can't fix this.
I can confidently say that this is not the only good reason. Other reasons include:
- You want to compare different runs by hashing outputs (e.g. to find the first computation step where they diverged). Very useful for debugging, and also useful to determine whether you accurately reproduced a result (e.g. a customer problem).
- If your program has a single floating point comparison, there is no such thing as "enough significant digits" - with reasonable assumptions about the distribution of "unreproducability", your logic is now divergent (and your output will jump between different values) with a certain probability. At that point we're no longer talking numerical analysis, it's straight up "divergent results".
This doesn't seem right, or at least it's not very general. The more operations you do, the more rounding errors you have, and each operation has the potential to magnify earlier errors. In solving ill-conditioned problems (which are not uncommon) the errors can easily be magnified so much that they're bigger than your signal, even with relatively small and simple situations.
phave you ever played Kerbal Soace Program? have you met the Kraken?
the game simulates a spacecraft you build piece by piece. So you are flying around in space and a thruster fires and maybe some part of the ship is poorly attached and it wobbling. Everything is fine, but then you clip upper atmosphere of mars, and everything goes to shit - tleven though it shoupd pose no threat, the spacecraft spontaneously shakes itself to pieces
That game is plagued by massove problems with floating point errors , they kill you vrew, they ruin your missions, your speed of rotation becomea a NAN
Simulating human-scale physics in a solar-system-scale world needs far more than 8 decimal digits. Neptune's orbit is 12 orders of magnitude larger than your spacecraft, no wonder your floats are acting up. You're going to need at least 50-bit precision to make your simulations accurate to the millimeter.
If anything, the fact that KSP is pretty much the only game with serious float issues shows that doubles are just fine for most applications.
Instead, this is a rambling, mostly unstructured document where it's nearly impossible to follow the many scattered threads of thought, or even catch which side of some of the arguments he's actually on.
For business and accounting systems, this seems like an obvious choice.
How Java’s Floating-Point Hurts Everyone Everywhere (1998) [pdf] - https://news.ycombinator.com/item?id=6585828 - Oct 2013 (72 comments)
> Conclusions
> ...
> To win, Java has to surpass Microsoft's J++ in in attractiveness to software developers. This means better design better thought through, less prone to error, easier to debug, ... and many other things.
This document is a bit dated.
Because that’s the alternative to your “sad state.”
The premature optimization is beyond uncalled for. Learning to use floating point type is something that most developers should do. No need for weasel words, either (annoying, counterintuitive).
I can't think on any widely-used language that does it right nowadays (except Rust, if you change some flags, and not by default).
For sure, 20 years ago there were some dying languages that behaved differently, and yes, the Java position there did hurt everybody, but it's the same one everybody else took.
[1] https://doc.rust-lang.org/rustc/codegen-options/index.html#o...
There are many non-professional programmers who get baffled by various "computerisms" such as
0.1 + 0.2 != 0.3
Excel tries to hide this but ends up doing even stranger things if you push it hard enough. I think a lot of people who could use computers to put their skills on wheels just give up because of this "lack of empathy" that manifests here and in other places. If you do 0.1 + 0.2 - 0.3
on a pocket calculator you get the right answer and you should get the same right answer in a Jupyter notebook. The only person who should be exposed to the base 2 arithmetic of the computer is a professional programmer who knows assembly language.There are numerous social consequences of this that are harmful such as the perception that computer programmers are "grinds" and "nerds" and the idea that "idea people" are more worthy than the people that execute, etc.
there is a significant performance penalty for bcd arithmetic, so bcd floating point has never been attractive to the customers for floating point: cray buyers, fortran programmers, gamers, analog circuit designers, climatologists
those people don't really care about the beginner issues you mention
if you're willing to accept a performance penalty in exchange for easier learning, you can implement decimal arithmetic in software; many people have, and it's in the cobol and sql standards and the python standard library. the version i implemented in bubbleos is 83 lines of c, https://gitlab.com/kragen/bubbleos/blob/master/yeso/decimal.... https://gitlab.com/kragen/bubbleos/blob/master/yeso/decimal.... and compiles to about 1.3k of code
but a lot of us are using cpython in jupyter to prototype algorithms we want to run as fast as possible, so we want it to behave like the floating-point hardware does
But if you are programming, there are tons of quote-unquote computerisms that we have for a good reason but are not really intuitive for newcomers. The whole concept of variables, pointers and general indirections, zero-based indexing, (pseudo)randomness, Unicode, time complexity, concurrency and parallelism and so on. Many (but not all, I admit) professional programmers take them as granted but they are just as arcane as the concept of base-2 floating points.
I do think it would be a good idea for programming languages to expose more exact number representations (for example, rational numbers, which are supported by some programming languages). But numbers are complicated, and you will never be able to imbue a computer with a number type that always behaves exactly the way that someone who knows nothing of computers would naively expect. Even mathematica does not quite manage it, and mathematica takes some somewhat extreme measures which would not be considered acceptable in many other general-purpose languages.
Then what is the job of the language? Programming languages should nudge their users in the right direction; they should be safe and maintainable by default. Sometimes sharp-edged parts are necessary (e.g. for performance) and they should be made available to those who need them, but they shouldn't be front and center. E.g. floating point literals should be a high/arbitrary precision decimal type by default, similar to what Python does with integer literals.
> Decimal floating-point will be prone to very similar malfeasances, if that is what you are proposing; somebody unaware of the issues of binary floating point will not be helped by decimal floating point.
Disagree. Decimal calculations getting rounded off is a problem that is a lot easier to see and understand.
If you carefully, tediously work by hand some calculations such as addition, multiplication, or division (but not, say, square roots, exponentials, or sines, since there are no generally known methods for computing those by hand), you may be able to replicate some specific results of the computer without needing to learn anything new. I don't really see what the point of doing that would be. Oh, and you must use banker's rounding, which is also not generally done. Otherwise, the results will appear to manifest as just some inaccuracy incurred after every operation, just the same as with binary floating point. The full complexity of error analysis remains.
(Note also in particular that addition, multiplication, and division are closed over the rationals, which can be readily represented exactly by computer—I did mention I think rationals are a good thing.)
Adding 0.1 and 0.2 to get 0.3 is hardly advanced mathematics.
> Otherwise, the results will appear to manifest as just some inaccuracy incurred after every operation, just the same as with binary floating point.
Perhaps, but it will be a comprehensible inaccuracy. Running out of decimal places is a well-known and understood phenomenon, and something already experienced with regular calculators.
"float" is usually the second thing taught in the math tutorial, right after "int", in many languages.
JavaScript and many other languages hide many things away and just call everything "a number" which is a float by disguise.
"Decimal" types are usually hidden in some standard library, and not as a first class primitive type of the language.
Stop victim blaming, that's not going to get the discussion anywhere. You HAVE to acknowledge that the first option is non-obvious of its flaws.
It's not real analog noise, but reasoning about rounding is hard if you're not good at math, it's something you can explain easily, and it prevents you from doing clever stuff that will confuse the next person.
It's the opposite of being an idea person. I just accept them as a practical approximation, not a tool for doing real math, and move on. If you come from electronics it makes perfect sense. Voltages are pretty much always noisy even if it's picovolts.
If I need precision I use ints or dedicated libraries.
>should get the same right answer in a Jupyter notebook. The only person who should be exposed to the base 2
"Why does my notebook take hours to do a simple task"
>There are numerous social consequences of this that are harmful such as the perception that computer programmers are "grinds" and "nerds" and the idea that "idea people" are more worthy than the people that execute, etc.
The benfits of floating point arithmetic easily outweigh people having to learn it.
“Because the chip it is running on doesn’t have native base-10 math”?
See also “why my notebook [is fast but] yields incorrect results”.
The benfits of floating point arithmetic easily outweigh people having to learn it.
GP talks about saving non-low-level programmers from base 2 FP, not about removing it. CPUs could use an additional block (or mode) of base 10 exponent FP.
This and other geeky issues make programmers programmers instead of making everyone a programmer. The consequence of this is much heavier than any benefits of base 2 FP.
Base 10 fixes exactly zero of the problems with floating point.
Decimal floating point exists as a standard already. It is even part of the upcomming C standard.
But again decimal arithmetic is just as weird as floating point arithmetic. The change of base s irrelevant, except for a few niche applications.
>The consequence of this is much heavier than any benefits of base 2 FP.
Totally false. Decimals don't fix floats. They are just as weird. Changing the base is irrelevant to the inherent properties of floats and using base 10 instead of base 2 just means that with a very high chance you get something even worse.
If you do not understand the basics of floating point arithmetic you should not be programming software. Tough world out there for people who refuse to learn, I know.
The problem with binary floating point is that it is input and output as decimal floating point.
Every floating point number is really a fraction of the form
M
-------
b^E
When you write x = 0.1
you are really asking for 1/10. If b=2, however, you can only get a denominator of 1/2, 1/4, 1/8 or something like that. 1/10 just doesn't exist in that number system, but there is a number A that you get when you ask for 0.1 that round-trips back to 0.1, and the same is true for 0.2 (B) and 0.3 (C). The trouble is those substitute numbers aren't the real numbers and A != B + C
This has nothing to do with numeric precision, it's always going to be off a bit even if you are using 1024-bit floats or 1048576-bit floats.The problem punches above its weight because it is targeted right at two fault lines of the mind: (1) people flinch at inconsistencies, 0.1 + 0.2 = 0.3 is an identity and when basic identities are wrong people feel uncomfortable and don't want to proceed; if you are working with an accountant, for instance, and they see something that is not consistent in a way they've been trained has to be consistent, they will just stop until it is consistent. (2) A certain kind of laziness leads people to not get to the root of a problem like this and instead waste a huge amount of time and energy into non-solutions (rounding!) that are just like pushing a bubble around under a rug.
Note that decimal floating point does not require that you use BCD. That is, the mantissa and the exponent of a floating point number are just integers, and unlike the floats, base 7 and base 192 integers are the same numbers.
An obvious idea is to use base 2 for the mantissa and exponent, just have the exponent be base 10. One difficulty is the cost of sliding two numbers so they have the same exponent before you add them, for instance to add
43
723
---
766
you have to express the 723 and 43 in the same base to add them (multiply/add by ten) and multiplication by powers of ten is a lot hard with base 2 math than with base 10 math. You also can represent decimal floating point numbers with a decimal mantissa and face the tradeoff of BCD numbers being wasteful of bits but the factors of 10 being easy to deal with. There is an efficiency gap between hardware binary floating point and hardware decimal floating point but it's not as bad as you might think.>1/10 just doesn't exist in that number system
And 1/3 doesn't exist for b=10.
I really have no idea what you are on about. The number system you want does not exist, it is a mathematical theorem, floats are the best approximation of real numbers and choosing b=10 is dumb outside of specific applications.
Decimal floating point fixes nothing it just moves around the issues where they less affect numbers in base 10. It is exactly as broken as base 2. You still violate basic identities in b=10, just different ones.
Again, the number system you want doesn't exist and it can't. You can never approximate the real numbers with constant bits, division and arithmetic consistency.
431*2^-9
where 2^-9 is 1/512 there would be no semantic gap here.The fact that you find it so hard to get what I am talking about isn't reflective of your intelligence, it is that this is something that strikes people where they are absolutely weakest, where the gap between the map and the territory can leave you absolutely lost. That's what is so bad about it.
Of course that is why computer professionals suffer with the "grind" and "nerd" slurs, because we tolerate things like the language that puts the C in Cthulhu.
Personally I think Cantor's phony numbers suck and it is an insult that we call them "real" numbers. It's practically forgotten that Turing discovered the theory of computation not by investigating computation but by confronting the problem that there are two kinds of "real" numbers: the ones that are really real because they have names (e.g. 4, sqrt(2), π) and the vastly larger set of numbers that can never be picked out individually (e.g. any description of how to compute a number has to be finite in length but the phony numbers are uncountable.)
I wish Steve Wolfram would grow some balls and reject the axiom of choice.
Now that is a really dumb ask. Surely designing computers around the numbers of finger we have for no reason is insane. Again, b=10 DOES NOT FIX FLOATS. It has the exact same problems.
>Personally I think Cantor's phony numbers suck
Personally I think they are the greatest description of the continuum.
>I wish Steve Wolfram would grow some balls and reject the axiom of choice.
Yes. Really looking forwards to such hits as "you can't split the continuum" and "points do not exist", contradicting 100% of human intuition.
The point is that when the literals don't match your numbers you get particularly strange problems that I think scare people away from computers. We just lose them.
The other problems with floating point math create much less cognitive dissonance than that does.
I actually hope that "we" loose people who do not put in that tiny amount of effort to learn something so simple. They have absolutely no buisiness developing software.
If you are unwilling to learn such absolutely basic concepts as what base 2 is, you really should be excluded.
Like with random number generators - computers are so powerful now that it makes sense to make the default PRNG cryptographically secure, to avoid misuse, and leave it up to the experts to swap in a fast (not cryptographically secure) PRNG in the rare cases where the performance is needed (and unpredictability isn't), for example Monte Carlo simulations.
One could argue that for many use cases (most spreadsheets handling money amounts, for example) computers are powerful enough now that "non-integer numbers" should default to some base-10 floating point, so that internal and user representation coincide. Experts that handle applications with "real" numbers can then explicitly switch to "weird float". It is worth a thought.
[1] https://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.31....
https://www.cs.tufts.edu/~nr/cs257/archive/florian-loitsch/p...
https://en.wikipedia.org/wiki/Floating-point_arithmetic#Bina...
The only difference is that different numbers are representible. You still have:
- They do not obey mathematical laws for real numbers
- Equality comparison is meaningless
- Multiple operation lead to unboundedly large errors
>Experts that handle applications with "real" numbers can then explicitly switch to "weird float". It is worth a thought.
The weirdness does not go away in decimal floats. It is exactly as weird.
>One could argue that for many use cases (most spreadsheets handling money amounts, for example) computers are powerful enough now that "non-integer numbers" should default to some base-10 floating point
Base 10 floats do not fix money amounts. (1/3)3 is not* equal to 1 in base 10 floats. You can not correctly handle money with floats, as money is not arbitrarily dividable. Changing to base 10 does not fix that.
The core problem with money is that the divison operation is not well defined. As for example $3.333... is not an amount of money that can exist. Even the mathematically correct operations are wrong, you can not fix that with imperfect approximations.
And what you see is what you get.
1/3 = 0.333333, rounded, makes sense.
0.333333 x 3 = 0.999999, makes sense.
>But you fix some, for example 0.1 + 0.2 == 0.3 in base 10 floats.
And you get the exact same amount of NEW problems. You are just shifting around where they are.
0.1 + 0.2 != 0.3 also makes a lot of sense, if you have tiniest knowledge about floats.
There are corners we don’t need to cut anymore, because a half century of Moore’s Law has already paid for better tools if only we would claim them.
Floats are generally better near zero, but better (although less useful) approximations exist.
>Rational numbers don’t need to be approximated
Nonsense. "rational" datatypes are far less useful than floats. Approximating rationals by floats is usually better.
FYI rational datatypes can not support operations which result in irrational numbers, such as exp, sqrt.
>where the hardware itself limits precision
Hardware does not limit precision. It implements IEEE 754.
> letting an end user dictate how much precision he wants to see at the moment.
Those are arbirtrary precision numbers.
>There are corners we don’t need to cut anymore, because a half century of Moore’s Law has already paid for better tools if only we would claim them.
Total nonsense. Fast floats have never been more important than they are now.
This is an unsolvable problem. If your floats are base-10 "0.1 + 0.2 = 0.3" might work out, but it isn't going to fix "(1/3)3 = 1". And it gets even worse if you do anything involving π.
It is mathematically impossible for a computer to handle all* reals correctly, so something has to give. Binary floats are a reasonable approximation in practice, and anyone wanting something else is free to use arbitrary-precision libraries.
(1/3.0)*3.0-1.0==0.0
and got back the answer True
no problem there.Pythons is unusable for numerics without numpy, which makes it a tiny bit less unusable.
That isn't an "elitist" view. I don't get what you are on about, seems ridicolous. People need to learn to use the systems they are working with any highschooler can understand what floating point numbers are and why they have flaws. It is really simple.
Every problem is the mathematical consequence of its features.
While many people complain about it, IEEE 754 is a very thoughtful and thorough standard, and the reason it is still around is that nothing unambiguously better has been proposed yet.
Save floats for performance optimizations. Start out correct, optimize when you can show both need for speed and non-need for correctness.
Consider an object with a very small velocity traveling away from the origin with no forces acting upon it. With floating-point arithmetic, it will eventually stop at some distance away from the origin. In general objects with small velocity will only move if they are close to the origin.
A simulation that computes bulk properties by simulating particles directly is very rare. At least for me. Observable like temperature, electric/magnetic field strength, stress, velocity, etc are much more common outputs. These often all vary over many order of magnitude across a domain.
Working with continuum mechanics, performing large inverted solves on matrices in fixed precision, or heck, even an FFT sounds terrifying to guard against overflow.
And almost any code that would work with 54 bit fixed point works even better with 64 bit floating point, and the floating point version is much easier to code.
So sure, floating point isn't optional when you start off assuming the same number of bits. But if you treat it as a small overhead to make code more robust to large numbers, easier to code, and often faster to run, then it looks a lot better.
For many situations though, I find the graceful degradation of floats to cause more subtle bugs, which can be a problem.
If integer overflows trap (curse you Intel, C89), then repeated additions (important in many simulations) will either work as expected or crash.
Floating point operations are (for practical purposes) highly privileged because of extremely mature hardware implementations. Hardware implementations that make other forms of calculations more tractable are possible (and have existed in the past) and should be considered when evaluating FPs fitness for purpose, otherwise we will be stuck in the IEEE local maximum forever.
How so? Is that because of inherent problems with the outcomes of fixed-point arithmetic, or are they just clumsy to use and no-one's written a decent library that makes dealing with fixed-point numbers straightforward?
I think GP is alluding to issues you get when you have a mix of scales. Like floating-point, fixed-point will have a limited precision range around the origin. In contrast to floating-point, the range for fixed-point is smaller but evenly spaced.
Say you have 32:32 fixed-point. You can then represent numbers in multiple of ~0.2e-9. So if you need to calculate distances in nanometers, perhaps due to a short timestep in a simulation, you have hardly any precision to work with.
The obvious way around this is to pull out a scale factor, say 1e-9 since you know you'll be in the nano-range. Now the numbers you're working with are back on the order of 1 with lots of precision, however you need to apply the scale factor when you actually want to use the number, and now you have to be careful not to overflow when multiplying two or more scaled numbers. This is part of the programmer overhead GP alluded to.
Do that rescaling automatically and have a system that just keeps track of the scale and you've just implemented floats. (:
Because you need to be extremely careful about overflows/underflows. All operations you perform suddenly become difficult and require careful analysis. With fixed point numbers you need to ensure that every intermediate result of your operation gives an in range result.
You can not remove that tedium by a library, since any potential application has different requirements and now you need to start out your program by defining those requirements for your library. All operations you perform need to be analysed based on your initial requirements.
Also, fixed point arithmetic does not fix floats. You have an uncountable number of real numbers and you try to fit that into 32 or so bytes. It seems simple to understand that any such projectiom whatever you try has enormous drawbacks.
That floats do not behave like real numbers is the consequence of their design requirements. Fixed point just means instead of being able to pretend that floating point operations are sometimes inaccurate real number operations, you have to deal with constan, domain specific renormalization and have to be extremely careful aboit choosing scaling.
How is that any different from ordinary integer operations/arithmetic with ints/longs/etc...?
> Also, fixed point arithmetic does not fix floats.
I wasn't expecting it to?
It is different, because you never want to take the square root of an integer.
Floats allow you to pretend you are working with sometimes inaccurate real numbers. That is the magic behind it and why fixed point arithmetic was abandoned almost immediately when floating point became fast and easily available.
One of the key advantages of floating-point is that it is scale-independent--it doesn't matter if you're doing your calculations in meters, feet, miles, kilometers, parsecs, AUs, nanometers, Planck lengths--you'll get the same relative accuracy. If you're using fixed point (or integers), you instead have to take care to make sure that you scale things such that your units are not too large or too small.
Floating-point is essentially binary scientific notation, and it should be no surprise that it's a good format if you're in a field that already uses scientific notation all the time.
Yes, but if I'm working with lengths where I know micrometers are good enough, but I also want micrometer (or, at least, better-than-millimeter) precision, but I want to actually deal in meters and I know I'm not going to be dealing with lengths longer than 100km, then fixed point would be ideal for that.
Simlarly, if I want to deal in dollars, but need cent precision (or, tenth of a cent precision), and need it to be precise, then fixed point would be ideal for that too.
I think the disconnect might be that I wasn't considering fixed-point arithmetic as a general replacement for floating-point arithmetic, I'm considering it as a complement to floating-point, to be used in the cases where it makes sense to do so.
The impression I got from the commenter I was originally replying to was "fixed point is absolutely terrible and there is never a reason to prefer it over floats". If that wasn't an intended implication, I guess I'm kind of arguing into the void.
Huge fallacy. No, you can not use fixed point numbers like that, it can not work. It is irrelevant what the actual maximum/minimum scale you are dealing with is. The thing which matters is the largest/smallest intermediate value you need. You need to consider every step in all your algorithms.
Imagine calculating the distance of two objects being 50km in x and 50km in y direction apart. Even though the input and output values fit within your range, the result is nonsensical if you use naive fix point arithmetic. Floating point allows you to write down the mathematical formula exactly, using fixed point arithmetic, you can not do that.
Looking at the maximum and minimum resolution you need is a huge fallacy when working with fixed point arithmetic and one big reason why everyone avoids using it. You need to carefully analyze your entire algorithm to use it.
>The impression I got from the commenter I was originally replying to was "fixed point is absolutely terrible and there is never a reason to prefer it over floats".
My position is that sometimes there might be a situation where fixed point arithmetic could be useful. If you are willing to put in a significant amount of time and effort into analyzing your system and dataflow it can be implememted successfully.
It is a far more complex and errorprone system and if you aren't careful it will bite you. In all cases floating point should be the default option and only deviated from if there are very good reasons.
Forget everything else, I need to pay my bills, this involves money.
I didn't say anything else because it is hard not to make sarcastic jokes and devolve into a rant.
You break down a problem to it's smallest parts then you solve those small problems. It seems very basic.
The smallest problem is c = a + b NOT a + b
One can do c = 2.03 , there is no problem storing the number.
One can also do 103 + 100 like c = ((a100) + (b100))/100
This looks even more hideous than c = asfdsADD(a,b) but at least you don't need to hurl around a big num lib to do 1+1.
If things get ever so slightly more complicated than 1+1, what should be a bunch of nice looking formulas looks horrible. Heaven forbid one wants a percentage of something. I would have to think how to accomplish that without the lib.
I have a very slow brain with very little memory and many threads, I'd much rather spend cpu cycles.
This is UNSOLVABLE. There is no solution to this if you want to approximate real numbers in constant size.
This is also a very stupid argument, since it misses the point of floating point numbers.
It's normal and expected for ML these days (not completely custom, but smaller than the IEEE floats). And I think for gaming graphics too.
> If you need simd, or vectorization, including gpu acceleration, you use the format the hardware and compiler/optimizer support to meet your throughput needs. And on x86_64 aarch64 or spirv, thats a ieee floating point for better or worse.
If you need to do it on the CPU, sure. (Even then, you probably don't want to use full IEEE754, you will likely disable subnormals for performance). If you're using the GPU then you'll likely be using something else. The niche that IEEE754 fits is pretty narrow - it works for non-HPC physics simulations because it was designed for that, but that's about it.
[1] I believe the biggest use of quad precision is evaluating the accuracy of double precision arithmetic.
They are usually useless. Almost all programs need fast floats.
I never needed fast anything, things are plenty fast for my needs, insanely fast.
If your software isn't manipulating Gigabytes of data or solving a hard math problem everything should be instantaneous.
Just because you might not understand basic CS knowledge is no reason to deprive the world of the best approximation to real numbers we have.
By the way. Big num libraries absolutely do not solve the problems of float. They, just as floats, can suffer from unbounded large errors. And 0.3 isn't representable regardless how long you make mantissa and exponent.
I appreciate the sentiment but in this case I do not want to hurl around a bignum lib to add up 2 tiny numbers.
> If your software isn't manipulating Gigabytes of data or solving a hard math problem everything should be instantaneous.
I agree, NO amount of bandwidth should be spend on this, build in functionality should do the small number arithmetic.
> Just because you might not understand basic CS knowledge is no reason to deprive the world of the best approximation to real numbers we have.
Meanwhile, back on the construction site I'm hammering screws.
No, I don't want to take your nails away from you.
That is why integers and floats exist.
>build in functionality should do the small number arithmetic.
Every CPU has that.
Floating point numbers CAN NOT be fixed. It is mathematically impossible. The inconsistencies always will be there.
Many people approach them as "oh, they're decimals" when they absolutely are not... but there are no convenient alternatives in most languages. So people use the convenient one that they see 99% of other code using, rather than something that is more likely to fit their expectations.
That's the problem. People are surrounded by screws and hammers and they understandably choose to hammer in the screws rather than 1) knowing a screwdriver exists, and 2) hunting down the few three-handed screwdrivers in existence and trying to use them.
Floats are the best in almost every circumstance. Which is why they are the right default.
It's theoretically impossible to compute with the general reals. Full stop.
It's theoretically possible to compute with the general rationals (Common Lisp for example lets you do so) but it's often impractical because you can easily end up running out of memory and/or time for the computation to complete.
It's practical to compute with floats because you can guarantee hard limits on memory usage and time. You just have to live with the fact that you can't represent the vast majority of real numbers and you lose associative arithmetic.
There are various other, weirder, computational number systems out there but the tradeoffs don't go away; they just move around.
You could limit yourself to the computable numbers. Although you do have an issue determining equality on general computable numbers...
But yeah, I would agree that floating-point is generally the best compromise out there for a general-purpose computational approximation to real numbers.
While that is true. It is possible to make mathematical true statememts about results of real number arithmetic on a computer.
>you lose associative arithmetic.
People like to complain about this, but it is a mathematical consequence of the reuirements. You can not fix that problem.
Floats can be fixed just as much free energy can be invented.