Am I overlooking it, or is this method missing from Wikipedia's "Division Algorithm" entry?
https://en.wikipedia.org/wiki/Division_algorithm?useskin=vec...
How it works is easy enough to figure from that description. You rescale your divisor to fit in the range (0.5, 1.0], then expand it as the series 1/(1 - x) = 1 + x + x^2..., which converges, and does so linearly at >1 bit per term (since x < 1/2). What's useful about this, is you can calculate the finite sums with a circuit that's logarithmic depth, or maybe better. You could do the repeated squarings to get: x, x^2, x^4, x^8...; then you have a bunch of parallel hardware multipliers filling out all the intermediate terms, which are products of zero- or one- of each of these basic terms; then you just find the sum with a logarithmic-sized tree of adders.
Somehow, I can't find the name of this algorithm (?)