Understanding Fast-Math
pspdfkit.com
pspdfkit.com
> It turns out that, like with almost anything else relating to IEEE floating-point math, it’s a rabbit hole full of surprising behaviors.
immediately before describing that they disabled IEEE floating-point math is a bit funny. The standard isn't surprising, floating point numbers (arguably) are. The whole point of standardizing floating point path was to reduce this surprisingness. Can't complain about IEEE floating point numbers if you tell the compiler not to use them.
But almost nobody spends this much effort to familiarize themselves with the floating point system. So it keeps surprising people.
If it were a harmless flag you'd think it would be enabled by default. That's a clue that you should look before you leap.
We have a few files in our code base that compile with -ffast-math but the rest don't. Those files were written with that flag in mind.
But that perspective aside, I agree: generally people correctly assume that their tools are OK and their code is at fault.
I wouldn't consider -ffast-math a bug: it's a sharp tool that should only be used by experienced users. If there's a bug at all it's that, in retrospect, the flag should have had a different name.
I feel like a lot of programmers just type stuff and assume if it compiles it should ship.
This should be the main takeaway. Don't enable ffast-math if floating point calculations aren't a bottleneck, and especially not if you don't understand all the other optimizations it enables.
As the article demonstrates, there is: with ffast-math, floating point subtly behaves in ways that don't match the ways you've been taught it behaves.
As far as I remember fast math also breaks things like NaN and infinity, which makes filtering out invalid values before they hit something important fun since they will obviously still exist and mess up your results but you can no longer check for them.
Also, there are other uses of FP besides games..
Just like cranking the warning levels as high as I can at the beginning of a project, I also like to build and test with -ffast-math (really -Ofast) from the very beginning. Keeping it warning-free and working under -ffast-math as I go is a lot easier than trying to do it all at once later!
And much like the warnings, I find that any new code that fails under -ffast-math tends to be a bit suspect. I've found stuff that tends to break under -ffast-math will also frequently break with a different compiler or on a different hardware architecture. So -ffast-math is a nice canary for that.
I got a four times speedup on <cmath> functions with no loss in accuracy.
See also "Why Standard C++ Math Functions Are Slow":
https://medium.com/@ryan.burn/why-standard-c-math-functions-...
Compiling a big program can be a bit of a pain, though, so it is probably only worthwhile if you have a program that you use very frequently. Also compilers aren't magic, the bottleneck in the program you want to run could be various things: CPU stuff, memory bandwidth, weird memory access patterns, disk access, network access, etc. The compiler mostly just helps with the first one.
Also, note that some libraries, like Intel's MKL, are able to check what processor you are using and just dispatch the appropriate code (your mileage may vary, they sometimes don't keep up with changes in AMD processors, causing great annoyance).
My currently used optimization flags are: -O3 -fno-math-errno -fno-signed-zeros -march=native -flto
Only use -march=native if the program is only intended to run on your own machine. It carries out architecture-specific optimizations that make the program non-portable.
Also look into profile-guided optimization, where you compile and run your program, automatically generate a statistical report, then recompile using that generated information. It can result in some dramatic speedups.
https://ddmler.github.io/compiler/2018/06/29/profile-guided-...
I mostly use Java, and my impression is JIT inside JVM introduces hardware specific optimization without any user intervention. But for C/C++ if dependency is included as source - I can use compiler flag to enable platform specific optimizations. But if the dependency was included in the form of a pre-compiled binary, such as a .dll or .so, I am probably not using the the most optimally compiled version of the dependency. Am I right so far?
My employer spends many millions of dollars annually running numerical simulations using double precision floating point numbers. Some years ago when we retired the last machines that didn't support SSE2, adding a flag to allow the compiler to generate SSE2 instructions had a big time and cost savings for our simulations.
That's kind of a special case, though. Without SSE2, you're using x87 for floating-point numbers, and even using scalar floating point on x87 is going to be a fair bit slower than using scalar floating point SSE instructions. Of course, enabling SSE also allows you to vectorize floating point at all, but you'll still be seeing improvements just from scalar SSE instead of x87.
1. Using SIMD can be a big win, so yes.
2. SIMD (vectorization) is not the only optimization your compiler can do, the compiler has a model of the processor so it can pick the right instructions and lay them out properly with as many tricks as they can describe generically.
3. Compilers have PGO. Use it (if you can). Compilers without PGO are a bit like an engine management unit with no sensors - all the gear, no idea. The compiler has to assume a hazy middle-of-the-road estimate of what branches will be exercised, whereas with PGO enabled your compiler can make the cold code smaller, and be more aggressive with hot code etc. etc.
I like this because it only makes sense in some accents. For example it wouldn't work in Boston where the r would only be pronounced on one of the words (idea).
You're already in mortal peril if you're working with "coefficients near zero" because of denormals, another bad idea that should have been disabled by default and turned on only in the vanishingly-few applications that benefit from them.
There is even a good reason for that, math-errno is a posix requirement and completely optional in both C and C++ standards. If your code is intended to be portable it should avoid relying on this anti-feature anyway.
The handling of isnan is certainly a big question. I can see wanting to respect that, but I can also see littering your code with assertions that isnan is false and compiling with normal optimizations and then hoping that later recompiling with fast-math and all its attendant "fun, safe optimizations" will let you avoid any performance penalty for all those asserts.