How to Print Floating-Point Numbers Accurately (1990) [pdf]
lists.nongnu.org
lists.nongnu.org
After asking around, I came across the Steele & White paper, and the implementation known as `dtoa()` written by David M. Gay. At the time, this seemed to be understood to be the "correct" answer.
But then I looked at the code: http://www.netlib.org/fp/dtoa.c
Uh.
There's a lot to dislike about that code, but arguably the worst thing is that it isn't thread-agnostic. It mutates global variables and protects them with a global mutex. A global mutex lock, just to print a number! Whyyyyyyyy?
So then I tried something different: I wrote some code that would do sprintf() with 15 digits precision first, then parse it with strtod() to see if it came back exact. If not... then I did sprintf() with 17 digits precision... and called it a day.
In benchmarks, this turned out to be just as fast as calling dtoa().
And so that's how the Protocol Buffers library deals with numbers when writing TextFormat:
https://github.com/google/protobuf/blob/ed4321d1cb3319998411...
But the bigger lesson for me was: Transmitting floating-point numbers in text is awful. People have no idea how ridiculously complex this is, because it seems like it ought to be simple. When you send JSON with numbers in it, you are probably invoking code that looks something like dtoa(), over and over and over again. And that's just to write them out; I have no idea how complex the parsing side is.
Please, folks, think of the CPU cycles. When sending numeric data, use a binary format.
Binary representations of floating point number are a bit of a nest of vipers too, though. You can probably just pipe IEEE floats across the network as bytes, in practice, but it's risky.
It's probably safer in many cases to just transmit things in fixed point.
- not all hardware supports them
- subnormal numbers are annoying
- NaNs can be tricky to handle, such as signalling and (if payloads are used) security risks through NaN payloads.
Which are you thinking of? Or something else?
> - not all hardware supports them
let's be reasonable : how many people who run architectures which don't have any float would actually send floats to them over the network ?
0x3.0p-2 = 0.75
You were making so much sense up until that point. :(
There is very exciting news coming up with printing floating point. Ulf Adams from Google will be presenting a new algorithm called Ryu that appears to be super fast, simple, and perfectly accurate. Assuming the claims are correct, Ryu ought displace all of the current algorithms.
I look forward to reading the Ryu paper.
printf("%a\n", 0x1p-10);
It's a shame that JSON doesn't accept them.
[0]: http://www.cs.tufts.edu/comp/150FP/archive/florian-loitsch/p...
William D. Clinger, "How to Read Floating Point Numbers Accurately," Proc. ACM SIGPLAN '90, pp. 92-101.
pdf: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.45....
And David Gay's implementation: http://www.netlib.org/fp/
FoundationDB tuples have type codes that segregate values of different types, so that the strings "1" and "2" sort before the integers 1 and 2, which sort before the single-precision floats 1.0f and 2.0f, which sort before the double-precision floats 1.0 and 2.0.
The database that I test supports double-precision IEEE floats, and a proprietary decimal float with a signed 64-bit significand and signed 8-bit exponent. When converted to string for use as keys, these collate as expected. The price of this is that you don't get shortest representations of the sort sought by this paper and others. Otherwise, a binary and decimal float that compare unequal could convert to the same string.
I guess it's kind of unusual to use floats as keys, and more unusual still to use both binary and decimal floats, but I wonder if there is another strategy for collating them.