x86 CPUs have a dedicated instruction to swap two registers, or a register with memory.
A note of caution: using that trick on modern hardware and/or languages is likely to be slower than just using the temporary variable. Because compilers and processors can see through the latter better than the XOR.