Yes, I have an AVX512 double precision exp implementation that does this thanks to iperm2pd.
This approach was also recommended by the Intel optimization manual -- a great resource.
I just went with straight math for single-precision, though.
I just went with straight math for single-precision, though.
No comments yet.