Some instructions don't guarantee the maximum possible precision, and implementations can differ in the least significant bits of their result. Addition and multiplication give you the correctly rounded result. However, there is typically some leeway in fma fusion. That is, if you multiply then add, it is in some cases implementation defined whether you get a result that was rounded after the multiplication and before the addition, or you get the correctly rounded result of doing the multiply and add as one operation without intermediate rounding.
All this is with the most conservative compiler settings. Be very, very careful about your optimization settings.
And let's not get into nans.
All this adds up to a situation where writing floating point code that is guaranteed to produce the same bit exact result on all implementations is basically a superhuman task. Better to just use fixed point and be done with it.
I wrote a multiplayer strategy game back in mid 90's issuing floating point and it was deterministic. No problems. Maybe chip optimizations have affected this?
We looked at fixed point, but with careful scheduling we'd get zero fpu (x87) stalls for float operations, so it wasn't a real win to go fixed. And it gave us the benefit of having more registers to use without needing to use the main stack as much, which also made the asm easier to read.
Edit: typos