SLEEF Vectorized Math Library
sleef.org
sleef.org
Personally I do a lot of approximations of math functions, which usually give me about a 10x speed increase relative to <math.h>. From the looks of it, this isn't quite that fast.
LLVM doesn't have a vector math library, so SLEEF could help you there.
If you're using single precision and AVX512, a >10x speed increase is likely. Otherwise, you'll probably get less than that. These functions are very accurate, most to withing 1 ULP of least precision. That is, if the answer they provide isn't the correctly rounded floating point answer, then it'll be either the next or previous representable floating point number.
There is a lot of room for giving up accuracy in the name of speed (eg, using less terms in the polynomials).
Sorry I wasn't clear, by "10x" I meant 10x faster than the standard library with the fastest compiler options (-O3 -ffast-math -avx512), but I've only tested with clang.
It is much faster (when vectorized) than what you get in base Julia, but lags behind gcc (glibc) and the Intel compiler's vectorized math libraries in performance.
For special functions, LoopVectorization relies on SLEEFPirates.jl, which is a fork of SLEEF.jl, a Julia port of version 2 of SLEEF. Most of the changes in SLEEFPirates are so that it works when you use llvm-vectors as arguments, but I also switched to using Estrin's rather than Horner's method of evaluating polynomials for a few functions (which more recent SLEEF versions did as well).
The code is pure Julia (or Julia + LLVM call; either way it does not need any external dependencies aside from Julia itself). It does need performance work, but I have many higher priorities at the moment.
You'll need at least GCC 8 to use them automatically, as well as the -ffast-math flag: https://godbolt.org/z/PL26up