How to defeat the purpose of IEEE floating point (2008)
yosefk.com
yosefk.com
This usually means you need regularization, or smoothing, or low pass filter, or averaging in some guise or the other (in some abstract way they are the same thing). There are some frequently used operations in ML that are very unstable under inaccuracies, numeric differentiation is one, just dont do it without smoothening. Inverting matrices is another, but you better have a real solid reason to ever invert a matrix, yes there are reasons when you do really want the inverse.
In any case the final point is that I have had to baby talk the compilers into, "yes, please please dont worry about it, just do these commutative (for real numbers) operations in any order that you want". This exposes more parallelism and other compiler optimizations. Often I would gladly take the hit over some precision in higher order decimal places to gain speed. But this should not be done blindly, and a smattering of numerical analysis helps gauging when it would be safe to do so. I doubt whether numerical analysis figures prominently in the trajectory of a graduating machine learner, although I think it ought to.
* (a,b,c) can take advantage of those tricks, whereas the programmer can still be explicit and go * (a, * (b, c)), for example.
http://gafferongames.com/networking-for-game-programmers/flo...
If your code changes behavior based on the timing of the various sections, you have a bug, a race condition. The solution isn't to lock down your code and build with only one configuration. The solution is to fix the bug.
In the same way, if the behavior of your code is dependent on the order of operations in floating point, it is very likely that you are doing something wrong. The solution is not careful ordering of floating point operations, but using algorithms appropriate to floating-point (and avoiding floating point when it isn't applicable.)
Code that has been ported to many platforms tends to be more stable, because the heterogeneous environments have exposed bugs and flawed assumptions in the code's construction.
Would you say that anyone who tries to make a game that has the same property is "doing something wrong"?
They don't! The system can tolerate answers which are slightly wrong... as long as all machines calculate the SAME slightly-wrong answer.
The same logic applies to synchronized PRNGs. It doesn't (directly) matter what the "random" dice-roll result is, as long as all clients get the same answer.
- Reproducible results across platforms
- Reproducible results with different build settings (such as debug builds or optimization levels)
- Useful test cases for regression testing
- Test cases that help him find misbehavior in the optimizer
- An improved chance of finding more subtle bugs, absent 'noise' from small changes in floating point operations
In particular, he gives examples of how changes to nearby code can change the output of floating point operations, due to hairiness in how registers are allocated. You can see how this would make testing quite painful.
These all seem like worthy goals that deserve a better response than "you're doing it wrong", along with some non-sequitur about race conditions.
> It's not just floating point that's inconsistent across modes – it's code snippets with behavior undefined by the language, buggy dependence on timing, optimizer bugs, conditional compilation, etc.
I cannot take a floating point rant seriously without knowing Kahan's arguments. And a crazy requirement to compare it to slow and false debugging code does not really help.
He could also use fdlibm ( http://netlib.org/fdlibm/index.html ), at least for anything other than primitive operations. But, again, it trades performance for consistency.
Sure, 0.1 can be represented exactly. But that's about the only issue with floating-point that that solves.