> This is a bit ironic because on x64 all math is SIMD math, it is just inefficient SIMD math.
I don't understand this statement at the end of the article? Can anyone explain? TIA.
I don't understand this statement at the end of the article? Can anyone explain? TIA.
float my_func(float lhs, float rhs) {
return 2.0f * lhs - 3.0f * rhs;
}
Becomes: my_func(float, float):
addss xmm0, xmm0
mulss xmm1, DWORD PTR .LC0[rip]
subss xmm0, xmm1
ret
(addss, mulss and subss are SSE2 instructions.)The instructions are still there even in 64-bit long mode, they use their own registers, and there are enough idiosyncrasies (80-bit extended-double precision, stack-based operations, etc.) that I would expect it to be easier to just include a dedicated scalar x87 FPU than try to shoehorn x87 compatibility into the SIMD units.