> especially since floating point operations are generally compiled to use the x87 fpu (so the floats will be moved back into gprs, and end up being slower than a regular swap)
gcc, at least, will use SSE rather than x87 by default on 64-bit x86. Looking at godbolt there's still a bunch of conversions that come out to about the same number of instructions, although I suspect doing it with SSE is probably faster than doing it with x87. Doing it with a temporary variable instead rather than coercing through a series of casts is probably even faster, though.