Thanks for clarifying!
Without benchmarking, I would expect atan2f to be around 20-30 cycles per element or less with either Intel's or Apple's scalar math library, and proportionally faster for their vector libs.
Without benchmarking, I would expect atan2f to be around 20-30 cycles per element or less with either Intel's or Apple's scalar math library, and proportionally faster for their vector libs.
By the way, your writing on floating point arithmetic is very informative -- I even cite a message of yours on FMA in the post itself!