0.30000000000000004
google.com
google.com
Financial arithmetic? Convert to smallest unit and use integers or the currency data type du jour in your language, and don't act surprised when operations on 32-bit floating point don't yield the intuitively correct values.
If you understood the representation format, you'd understand why.
http://www.schemers.org/Documents/Standards/R5RS/HTML/r5rs-Z...
The suggestion in the R5RS spec is that you shouldn't need to worry about the segregation of integer and real types (unless you need to). And indeed, why should you?
It's a conceptual shift away from char, int, long, float and double, which exist only because they're easier for the computer to deal with. Scheme was the first language I used that attempted to treat numbers like numbers.
Besides, Exactness is required in R5RS and the full numerical stack is no longer optional in R6RS.
Even competent coders who usually work in higher-level languages can sometimes forget that the decimal representations they work with are actually approximations of binary numbers, which is responsible for other seemingly weird behavior (why does ~2 = -3?), and I doubt even most skilled computer scientists are used to converting decimals to arbitrary-precision binary in their head.
It is unintuitive, and it's not like it's hard to explain why. "You'd understand if you were smarter" is a cop out.
Whether you can convert a decimal number into its binary IEEE-754 equivalent, or not, is besides the point.
You just need to know that it is lossy for certain classes of decimal numbers, and how to avoid or mitigate those effects.
But prime factorization seems to be taught for almost superstitious reasons, as some sort of math trick rather than a fundamental aspect of understanding numbers at even the most basic of levels, with neither students nor teachers nor the curriculum writers really understanding why this is in the lesson plan, so, yeah, I suppose you have to be some sort of super expert genius to swiftly realize that 1/3 can't be represented in a base-2 number. But it shouldn't be that way.
It works in Excel...
Clojure here, FWIW:
(+ 0.1M 0.2M) => 0.3M
That 'M' suffix denotes a BigDecimal, which provides for arbitrary-precision decimal math (which Clojure's arithmetic ops dispatch to as necessary).
A similar 'N' suffix is coming in the next release that denotes (contagious) BigInteger math (though as fast as longs when values are < 2^31), so one doesn't have to worry about overflow issues (which rate up there with misunderstandings of internal floating point representations in terms of frequency).
Other reader syntax is provided for other common notations (e.g. 10e6, 16rFF, 0xFF, 0220, 2r100010101).
Though, I assume non-suffixed literals are still regular floats?
It's also slow as molasses, which is why very few languages default to decimal floats and use IEEE floats (or doubles) instead, via their hardware implementation. The behavior of IEEE floats and doubles is very well defined, and though they are unfit in some sectors (you do not want to count money using them), have to be massaged a bit when displayed and don't deal well with great differences in powers-of-10 (e.g. 1e21 + 1 == 1e21) they work well enough in practice. And they're implemented in hardware.
Ah yes. Well not all languages require a bunch of boilerplate either. In Python for instance, there is a type decimal.Decimal which you can just alias to `d` and write:
val = d('0.1')Seems like other delta's (e.g. how many NaN's are implemented) don't much matter.
http://java.sun.com/docs/books/jvms/second_edition/html/Conc...
That said, the above notation makes advice to those who are having trouble a lot easier.
1> 0.1 + 0.3.
0.4
1> 1.0e21 + 1 == 1.0e21.
trueIf you used a decimal fractional representation instead, then one transaction of higher precision would contaminate all accounts it touches, making printing exact balances a pain.
If you want to do this correctly, then there should never be a balance that has an amount under the minimal unit. Then it doesn't matter if you use a BigDecimal or an int.
0.33333… is not an approximation.
Oh, this applies to me btw.
What does worry me is how many times this question is asked on The Web. This is a sad indictment on the quality of education in general. Given this resource, where information is far easier to find than in any library, there are still so many people who can't be bothered to look things up for themselves.
OK, I'm probably giving away too much about my age :-)
And some had better things to do anyway...
That said, yes, this does get misprogrammed in financial applications very frequently. FWIW, I wrote a little tuturial about this very issue a while back: http://roboprogs.com/devel/2010.02.html
How do we know it's not freshmen, high school students, and self taught developers asking the question?
From what is currently the top hit for that search:
"I am writing a program for class project..."
Two books I'd recommend: _Write Great Code vol. 1: Understanding the Machine_ by Randall Hyde, and _The Elements of Computing Systems_ by Nisan & Schocken (http://www.idc.ac.il/tecs/). WGC1 is about things typically learned about the machine while learning assembly (but leaves assembly itself for later). TECS guides you through building a (virtual) computer, starting with simulated NAND gates and building up. (Admittedly, I just got that one and I'm only a few chapters in, but excellent so far.)
I think every developer would benefit from knowing at least one low-level language (such as C), and one very high level language (such as Prolog).
I'm posting my project solutions to my github account, fwiw: http://github.com/happy4crazy/elements_of_computing_systems
Yes, though you may want more precision then that. For accuracy you could use, say, exact rational numbers, and for speed a fixed precision below e.g. Cents might be useful. (Though you probably meant something like nano-Cent when you said, convert to the smallest unit.)
I'm surprised someone hasn't translated it into Klingon or Sanskrit yet.
It's pure snobbery that makes people leave responses like this. This document falls into the further reading for enthusiasts category. Nobody should be linking to it as an introductory answer to the question. At least not if they want people to learn and stop asking the question eventually.
Many people have no interest in being computer scientists. They want to learn enough software engineering to get something done using a computer. Outside of CS, programming is only a means to an end.
Higher level languages ought not to expose the user/programmer to this sort of thing, even at the expense of performance; for every 'why is it that...' question, there are probably several more buggy programs being used in production environments. Also, it's a confusing distraction for kids trying to learn programming, who may not be able to wrap their heads around the theory of FP implementation, although they may be quite competent to handle the task they're trying to implement in software.
I used to work with a Motorola 56k series DSP, so named because it offers up to 56 bits of integer precision. Dealing with data that's usually represented as gloating point involves a normalization conversion as you describe for the currency unit which soon becomes second nature. But for newcomers to the platform, that step at first seems like an inefficient limitation, adding complexity to an otherwise simple algorithm. Not everybody wants or needs to know how the underlying hardware works, often they just want to implement some fairly trivial math or logic without worrying about accuracy.
I can't bring myself to correct that spelling error :)
There's a ton of stuff on similar level that does matter and it adds up.
Phenomena like knowing O(n) of basic algorithms and data structures. Understanding basic file formats and how they are implemented. Basic knowledge of network protocols and their building blocks. Basic knowledge of assembler. Understanding how IO happens and how your use case is bound (IO or CPU). Knowledge of character code pages, etc.
Nothing of the above is necessary to program in high level language. But if you want to do it well they can't possibly harm you. Or you will get involved with a vendor product that was written by someone who didn't know the above and you will be forced to reverse engineer the goddamn thing just to get your job done. :)
There are lot of people that might need to do a bit more programming - more elaborate than what a spreadsheet easily allows, but which doesn't justify becoming/hiring a pro. I was surprised to find last year that there are programming languages designed specifically for knitting, for example.
If it was at least a homework question on stackoverflow there would at least be a good description, and possibly a reasonable explanation.
14479.14
-
152.36
=
(result is 14326.78)
1143
/
78
=
(result is 14.6538461538461)
14479.14
-
152.36
=
(result is 14326.7799999999884)
Note that the first and third calculations are the same, yet they resulted in different displayed results!I never understood this bug. I understand floating point, so understand that some numbers are not exactly representable. However, it should at least still be consistent! The same calculation should give the same results every time.
* FP means the rules of algebra do not hold, so if the calculation is done in a different equivalent form, you can get a different result.
* With Intel's old FPU unit, if a value is written to memory rather than staying in registers you can get a different result.
http://webcast.berkeley.edu/course_details_new.php?seriesid=...
also available via iTunes U. I'm currently listening to them on my commute. Note, if you do actually go through the whole course -- you'll need to listen to a different year for lecture 24 or so -- that one is skipped. One of the highlights of my day is actually coming home and re-looking up what he's talking about.
(Oh, and of course you can just listen to the two Floating Point lectures. It has to do with the non-uniform -- or at least non-linearly uniform mapping of numbers, represented with a significand/mantissa and an exponent, onto the set of real numbers + the fact that the exponent used is in base 2, in the hardware, so the floating point numbers are spread about in a particular way. The difference between real numbers, as you tick up the odometer with each bit, varies depending on where you are in the number line (with big numbers, it's actually much more, with smaller numbers it's pretty minimal, but not unnoticeable as seen with this example. Does that make sense? Maybe I'm off about this... Anyway, still obviously recommend the lectures. And now, I'm going to read up more on ALUs and MUXs..)
Because why not? I've populated it with some languages that I can convenient access to an interpreter for. If you post/send me .1 + .2 in any other languages, I'll try and put them up.
$ ghci
GHCi, version 6.12.1: http://www.haskell.org/ghc/ :? for help
0.1 Loading package ghc-prim ... linking ... done.
Loading package integer-gmp ... linking ... done.
Loading package base ... linking ... done.
Prelude> 0.1 + 0.2
0.30000000000000004
And just for fun with GHC's rational numbers: Prelude> :m + Data.Ratio
Prelude Data.Ratio> (1 % 10) + (2 % 10)
3 % 10
Hugs (Haskell): $ hugs
Hugs> 0.1 + 0.2
0.3
bc: $ bc
0.1 + 0.2
.3
Gforth: $ gforth
0.1e 0.2e f+ f. 0.3 ok
dc: $ dc
0.1 0.2 + p
.3bc and dc are probably installed on your Linux or Unix box.
By the way, you should fix "Below are some examples of sending .1 + .2 to standard output in a variety of common languages." to "[...] in a variety of languages."
It also seems like the names of bc and dc are non-capitalized.
For the Haskell entry, please just shorten it to "0.1 + 0.2". The "Prelude>" thing is just a prompt for the REPL. ":m + Data.Ratio" loads the rational number module, please take the entry about Haskell's rational numbers out since yours is a page about floating point. (You might want to replace it with a comment, that Haskell supports rational numbers. But so do lots of languages in their libraries.)
C would be a good addition. (Plus Fortran, Cobol, Ada and J.)
$ racket
Welcome to Racket v5.0.1.
> (+ .1 .2)
0.30000000000000004
> (+ 1/10 2/10)
3/10
Steel Bank Common Lisp: $ sbcl
This is SBCL 1.0.43 [...elided]
* (+ .1 .2)
0.3
* (+ 1/10 2/10)
3/10 package main
import "fmt"
func main() {
fmt.Println(.1+.2)
}
0.3 PS C:\> 0.1 + 0.2
0.3the CS: "Hey, it's round off error. Get used to it"
the Mathematician: "Fix IT!!"
Well, it's fixable. You lose your deal of performance in order to get some absolutely valid rational numbers representation.
It's only that most tasks don't require this kind of precision.
x = 2
y = 2
1.upto 100 do
x = x ** (1 / 2)
end
1.upto 100 do
x = x ** 2
end
x == y # => true
because it handles the math by only evaluating the operations at the very end."It's a base 2 fractional number with no exact decimal expansion with a finite number of digits. Display it in base 2 fractional form if you don't want to see an approximation."
Binary numbers are a countably infinite set. Decimal numbers are a countably infinite set. You can therefore map binary representation to decimal representation 1-1 with a simple mapping function.
OTOH, floating point numbers are an uncountably infinite set. The only way to map binary numbers to floating point numbers is to map to a strict subset.
The subset happens to be different when mapping binary to floating point and decimal to floating point when using IEEE754.
Anyway, mathematicians don't use numbers at all - debasing the equations by performing caculations with them is so.. gauche. Leave that to the physicists, chemists and engineers.
We'd rather calculate with knows than numbers. (See http://en.wikipedia.org/wiki/Knot_theory)
Then I'd just just use PI as a symbol and do symbolic algebra for as long as possible. Don't evaluate it to anything until the user actually asks for a numeric representation. Once a numeric representation is required I'd pick a value from the above structure based on what has been asked for.
Every programmer forum gets a steady stream of novice questions about numbers not 'adding up.'...
Out of curiosity, I'm wondering how much trade you would have to be doing for floating-point imprecision to cause an actual problem.
Taking 0.2+0.1 as an example and figuring an imprecision of $0.00000000000000004 per $0.30, figuring a loss of one cent as being significant, I have 0.01/(0.00000000000000004 / 0.3) = 7.5e13, or... seventy-five trillion dollars?
Never mind that you're as likely to get 0.6+0.1 = 0.69999999999999996, which should roughly cancel out the error over time.
This is basically just an aesthetic problem in finance, yes?
Most of the time the only result is an imperceptible rise in noise, but it's not uncommon to have threshold-dependent routing of signal flow or for audio signals to used as input to modulation processors and vice versa. Every audio synthesis tool I've ever used had at least one bug resulting from error accumulation. They don't stand out very well in testing, but then people start reporting things like I was playing a tune with this patch and stopped to eat dinner, but when I came back my keyboard would only make a horrible noise, is it broken?
A normal double (IEEE 64) has only ~15 digits of precision, so when the amounts grow large, you loose the precision for the cents.
Example (in Java, which uses IEEE): double d = 1e9; System.err.println("d: " + d); d += 0.01; System.err.println("d: " + d); d -= 1e9; System.err.println("d: " + d); d -= 0.01; System.err.println("d: " + d);
What is 'd' at the end? Not zero. This adds a gazilion weird cases in the code that you have to handle.
Some countries have laws are very strict about how rounding should be performed and that all amounts must be an integral number of 'cents', so using doubles are completely out of the question in those cases.
To be pedantic, if the errors of size e are as likely to go one way as the other for each of N steps, the expected magnitude of the total error will be sqrt(N)e. It's a random walk.
0.1 + 0.2 = 0.30000000000000004
it's not a problem.
However, my problem is that modern language are hiding these things for you.
If so, are there any IEEE-754 alternatives?
There is alternatives but don't forget that it is hardware issue not software. It looks like http://en.wikipedia.org/wiki/IEEE_754-2008 has decimal format thus it should address that.
Python has a Decimal type that can represent decimal values exactly. It must be imported first though. It is probably much slower than IEEE-754 floating point, but for many uses that is not an issue. http://docs.python.org/library/decimal.html
Decimal arithmetics, which incurs a severe performance loss as IEE754 floats and doubles are implemented in hardware.
Edit: or it was standardized but still not gaining acceptance:
In the special case of 0.1 + 0.2 != 0.3, this is specific, and using decimal will "fix" it. But decimal, or for that matter any representation will have the exact same issue for other cases. Arbitrary precision means exactly that: arbitrary, as in not infinite precision. How will you represent irrational numbers with decimal ? It is theoretically impossible.
That's why most advices about using decimals instead of IEEE 754 are not appropriate. In the special case of accounting, it is appropriate, but I am somewhat doubtful most people don't need to do occasional computation with their numbers involving transcendental functions (e.g. log, exp, etc...). As soon as you start using this, you will see the issue cropping up, whatever representation you may want to use. And IEEE 754 representation has been conceived by people who really knew what they were doing (e.g. W. Kahan)