Floating-Point Determinism (2013)
randomascii.wordpress.com
randomascii.wordpress.com
Though I don't actually do any professional work with floats more than whatever distributed SQL correctness issues I get with sum() of doubles in the right/wrong order.
The error I try to correct are for visual/game code I wrote for fun & push through SSE.
More recently have started looking at bfloat16 which doesn't seem to have a lot of work done for this area.
Anyway floating point numerical stability is a different issue: some expressions produce results with less precision than other, different expressions, even though they should be the same thing algebraically (if we were dealing with real numbers or at least rational numbers). The issue here is that FP has limited precision, and it makes it lose some algebraic properties like associativity; and this limited precision sometimes lead to catastrophic cancellation and other problems that sharply reduce the precision (this doesn't happen if you use something with arbitrary precision like rational numbers).
Floating point determinism on the other hand is about always giving you the same result when running the same code (even if you happen to run in a different machine; and that's sometimes challenging because some architectures may not support features like subnormal numbers). It's okay if different code gives different results as far as determinism is concerned.
You can absolutely have determinism across different architectures, if you stick with architectures that support IEEE 754-2008. Also if a compiler break determinism this is a serious bug (as long as you don't enable -ffast-math of course - otherwise you're asking for breakage).
Unfortunately this means no SIMD, also you need to bring your own transcendental functions (sin, cos, etc). Also you need to be wary of other forms of nondeterminism like threads.
There's a physics engine meant for games and robotics called Rapier [0] that has a feature enhanced-determinism that will enable cross-platform determinism (by disabling threads and simd). If you spot any form of nondeterminism, it's absolutely a bug (either on the compiler or on the library itself), just like it would be a bug if it used only integers (which, by the way, was how nphysics, the antecessor of Rapier, implemented determinism: it used fixed-point math with integers [1])
[0] https://rapier.rs/docs/user_guides/rust/determinism/ - also it has amazing docs
[1] https://rustsim.org/blog/2020/06/01/this-month-in-rustsim/#m...
Care must also be taken not to risk using the x87 80-bit registers differently, but these days this is much easier - just don’t use x87 at all and use SSE instead.
I write software that does finite element analysis and similar things and results are reproducible across machines and across versions of the software.
My statement was about non-SSE scalar floating point operations. Using different optimizations ('Debug'/-O0 versus 'Release'/-O3 mode) will possibly produce different results, unless you are very careful. Using a different compiler (gcc versus clang versus visual studio) will likely produce different results.
If you want results to be reproducible across different CPU architectures (x86, arm64 etc) one option is using fixed-point arithmetic.
> Using different optimizations ('Debug'/-O0 versus 'Release'/-O3 mode) will possibly produce different results, unless you are very careful.
Yes. But .NET I think is much more predictable in that case as I don’t observe differences from optimization either. Having a spec and a decent memory model and no undefined behavior makes the compiler worse at optimizing things but better at consistency I guess. If the language spec strictly defines what math operators do and prevents reordering and similar, then there isn’t much outside transcendental functions that can go wrong. On 32bit .NET this did go wrong because whether or not something was in an 80bit x87 register or spilled to a 64 bit memory value seemed to depend on the moon phase. Those were bad times.
> If you want results to be reproducible across different CPU architectures (x86, arm64 etc) one option is using fixed-point arithmetic.
Luckily I never had that need. .NET (C#), x86-64, Windows. In that target, things look extremely stable now.
Step outside into 32bit or non-windows or C/C++ or even non x86 then all bets are off obviously.
I’m surprised you say it’s common as I haven’t yet bumped into the problem and I have a large suite of complex FP tests (engineering calculations) that have reproduced to the last digit over 15 years and dozens of different test machines of all flavors.
Results calculated 2004 on P4’s still work out to the same results after millions of calculations on every cpu since then and dozens of versions of the software.
Edit: SSE actually do transcendentals at all so it’s possible I’m saved by stable software defined transcendentals in .net (effectively then the Windows x64 implementations)