This isn't quite true, and the reality is actually way cooler than this.
Quake used a hybrid approach. It did both fixed point math on the ALUs and floating point math on the x87 FPU. The fixed point math on the ALU was used for game simulation, moving around, physics, stuff like that. The floating point math was used for graphics and drawing to the screen. Doing both sorts of math at the same time was non-trivial, but it had a huge upside: the ALU and FPU are independent of each other, so code executes in parallel. By doing FPU+ALU math simultaneously, you are effectively doubling the computational power accessible to the program, and as a result Quake looked head and shoulders more advanced than its contemporaries.
Not all CPUs had independent ALUs and FPUs. The Cyrix 6x86 ALU ops would delay its FPU ops and vice versa. So even though the 6x86 was faster than the Pentium at ALU ops and in the same ballpark for FPU ops, the 6x86 had absolutely terrible performance running Quake. This led to an undeserved reputation of the 6x86 of being slow and its benchmarks being fake marketing material amongst a certain segment of the population, which arguably led to Cyrix's decline and ultimate downfall.
--------------------------
Regarding square roots- the underlying architecture has an integer square root function, and fixed point numbers are represented as a left of decimal integer and a right of decimal integer. Why not just do integer square root on the left of decimal integer, and then do one (or two) iterations of Newton's method? This would be much faster and IMHO less complicated to code.