It's really fascinating to think about how all of this would work.
It's really fascinating to think about how all of this would work.
I think that’s one of those things Jai and Zig agree on - compile time functions have a place in preventing magic numbers that cannot be debugged.
Square roots are implemented in hardware: https://www.felixcloutier.com/x86/sqrtsd
> In software, you'd normally iterate Newton's Method
Software normally computes trigonometric functions (and other complicated ones like exponents and std::erf) with a high-degree polynomial approximation.
But how does that hardware implementation work internally?
The point I'm trying to make is that it is probably an (in-hardware) loop that uses Newton's Method.
ETA: The point being that, although in the source code it looks like all looks have been eliminated, they really haven't been if you dig deeper.
https://en.wikipedia.org/wiki/Methods_of_computing_square_ro...
Something like Heron's method is a special case of Newton's method.
For example: Pipelined RISC DSP chips have fat (many parallel streams) "one FFT result per clock cycle" pipelines that are rock solid (no cache hits or jitter).
The setup takes a few cyces but once primed it's
aquired data -> ( pipeline ) -> processed data
every clock cycle (with a pipeline delay, of course).
In that domain hardware implementations are chosen to work well with vector calculations and with consistent capped timings.
( To be clear, I haven't looked into that specific linked algo, I'm just pointing out it's not a N.R. only world )
I don’t know, but based on performance difference between FP32 and FP64 square root instructions, the implementation probably produces 4-5 bits of mantissa per cycle.