say more?
Better to just use more variables and let the register allocator in the compiler decide what to do. If it’s a loop then unrolling it once could remove the need for any swapping at all, for example.
This is only true for AMD cpus, on Intel xchg is 3 uops. Still better than the xor trick, though.
In theory, maybe.
But if that happens in your application and is performance critical, you probably should change it such that you're swapping pointers to them instead...
a = load(ap)
b = load(bp)
store(ap, b)
store(bp, a)
ap += step; bp += step;
any instructions to "do the swap" are a waste because generally load-store are separate instructions in SIMD instruction sets (and even if that wasn't the case, that's how they would get executed anyway)if you want to avoid polluting the cache there are SSE instructions for loading without caching, which might be worthwhile
edit: this might be useful in a SIMD context where you need to swap two registers, where the cost of using another register is higher than the cost of the 3 arithmetic instructions. i could totally imagine that happening, but it's nothing to do with caches or memory