On my computer it is indeed four times faster, but the marvellous method is on its own again four times slower than SSE, using the rsqrtps instruction, with the advantage of a way lower maximum error.
To take the square root of an array of 1024 floats:
Naive sqrtf() CPU cycles used: 48160, error: 0.000000
Vectorized SSE CPU cycles used: 2970, error: 0.000392
Marvellous Carmack method CPU cycles used: 11330, error: 0.002186