We did it for Wasm, which follows IEEE-754 semantics exactly for 32-bit and 64-bit floats. (The only nondeterminism is the exact bit pattern you get for NaNs in some circumstances.) Rounding is 100% well-specified. And CPUs have done that for decades. Even vector ISAs have learned that non-IEEE results are not what software wants; all vector ISAs are converging on IEEE-754.
> Optimized code is going to give you different results on different hardware due to the fact that you need to optimize things differently.
This is due to C/C++ (and to some extent Fortran) semantics. It is not hardware.
What do threads have to do with floating point precision?