Many people think floating point numbers are magically precise. They're not. They're far more accurate at lower magnitudes and precision fades as you work with larger values.
Many people think floating point numbers are magically precise. They're not. They're far more accurate at lower magnitudes and precision fades as you work with larger values.
Even worse, it kicks in earlier than one might naively think. People are often surprised to hear a value as simple as 0.1 has no matching floating point representation.
(For those interested, http://www.exploringbinary.com/why-0-point-1-does-not-exist-... has a nicely illustrated explanation.)
It does a variable number of fraction bits, so that the values near 1 are even more densely represented than under the usual IEEE floats, while also providing greater dynamic range (but at reduced precision).
I made a little diagram showing a visual comparison of every represented value in a toy 8-bit version vs. IEEE-style floats: https://groups.google.com/group/unum-computing/attach/e80274...
Designing a good arithmetic system takes a lot of attention to detail, and he doesn't make a good impression by being that loose with details in his introduction.
As to "no big deal".... show me your code.
Don't worry, I'll be able to understand it. After 20 years as a CPU designer I've learned how to understand bit-bashing.
And as a CPU designer, are you disagreeing that most FP units today take 4 cycles (fully pipelined) (add and multiply, of course)? And that main memory is a lot farther away than that?
Given 20 years of experience, you missed out on the Cray PVP machines that were the start of this sub-thread. But the cycle counts I'm giving are modern ones. Cray did eventually implement IEEE in these machines without any significant problems, but that was more than 20 years ago.
Maybe I'm misunderstanding your terminology, but it seems like you are saying that operations involving denormals have the same latency as normal floating point multiplications and additions. At least for multiplication on Intel chips through Haswell, I think the current case is that subnormals between 0 and FLT_MIN still have abysmal performance --- 100+ cycles of penalty.
Here's Bruce Dawson from a few years ago:
Performance implications on SSE
Intel handles NaNs and infinities much better on their SSE FPUs than on their x87 FPUs. NaNs and infinities have long been handled at full speed on this floating-point unit. However denormals are still a problem.
On Core 2 processors the worst-case I have measured is a 175 times slowdown, on SSE addition and multiplication.
On SandyBridge Intel has fixed this for addition – I was unable to produce any slowdown on ‘addps’ instructions. However SSE multiplication (‘mulps’) on Sandybridge has about a 140 cycle penalty if one of the inputs or results is a denormal.
https://randomascii.wordpress.com/2012/05/20/thats-not-norma...
And here's the overview from a slightly more recent paper:
C. Subnormal Performance Variability
Due to the complex nature of the floating point numbers, processors struggle to handle certain inputs efficiently. In particular, it is well understood that operating on subnormal values can cause extreme performance issues, including slowdowns of up to 100× [19]. As an example, on a Core i7 processor using SSE instructions, performing standard mul- tiply between two normal numbers takes 4 clock cycles, whereas the same multiply given a subnormal input takes over 200 clock cycles.
https://cseweb.ucsd.edu/~hovav/dist/subnormal.pdf
Are you saying this has been fixed in recent (or non-Intel) chips? Or maybe you were considering only NaN and Inf when you said denormals?
There are problems with repeatability of float operations such as loss of precision when moving value from registers to memory and data alignment issues, but it seems that these can be avoided with proper compiler options, at the expense of speed.
https://gcc.gnu.org/bugzilla/show_bug.cgi?id=323
https://software.intel.com/en-us/articles/run-to-run-reprodu...
There are reasonable criticisms of IEEE 754; that's not the issue.
How do you know that?
One of the half-truths in the presentation (in the intro) is that IEEE 754 is a mixture of requirements and recommendations. IEEE 754 does have both requirements and recommendations, however what the presentation doesn't say is that, within a given format like double precision (aka binary64), the basic operations like add, subtract, multiply, divide, squareRoot, etc.) have exactly one possible result for any given input (except that NaNs may have some implementation-defined bits, though this is usually unimportant).
To put it another way...it's about significant digits. A float has about 7 and a double about 15. It does not matter whether you use 0 to 360 or 0 to 3.6 million...the number of digits that are meaningful remain the same.
Using IEEE floating point to represent angle measures is horribly wasteful of bits, regardless of how you do it.
As an alternative, consider not storing angle measure at all, but instead using a stereographic projection (“half-angle tangent”) of the circle onto the whole line, and then using floating point numbers in the range (–∞, ∞] for your representation. https://en.wikipedia.org/wiki/Stereographic_projection
What you really want as a working format for rotations, ray directions, or points on a circle is to use a 2-dimensional square-grid (Cartesian coordinate) format like unit-magnitude “complex numbers”. The angle measure or stereographic projection is primarily useful as a compression format when you need to save bits transferring or storing data. Angle measures work okay for this, but can be annoying for requiring transcendental function evaluations to convert them to/from Cartesian coordinates, whereas to take the stereographic projection or its inverse only requires a single multiplicative inverse (i.e. division operation) plus some addition and multiplication per point.
It's more of an issue when you're dealing with large sums composed of very tiny ones. If you add them together incorrectly the number becomes so large the tiny values stop mattering even if in aggregate they're important.
Like adding 1e-6 to itself a hundred million times gives you a result different than 1e-6 * 1e8. In the first case I get 999.999998191639 when the multiplied version is 1000.
These errors can accumulate to a dangerous degree if you don't do your operations in the right order.
Really, 52 bits of mantissa is more than enough for representing anything. You have to be wary of error building up on calculations, not of original representation errors (unless you are living on the edge; and I'd tell you: don't).
Useless answer: No, the precision is fixed for any magnitude. When you multiply your values, you will also multiply the interval, and there are as many useful numbers inside that new, larger, interval than were inside the old, smaller one.