The XOR trick is not useful on modern CPUs for swapping memory blocks, because on modern CPUs the slowest operations are the memory accesses and the XOR trick needs too many memory accesses.
For swapping memory, the fastest way needs 4 memory accesses: load X, load Y, store X where Y was, store Y where X was. Each "load" and "store" in this sequence may consist of multiple load or store instructions, if multiple registers are used as the intermediate buffer. Ideally, an intermediate register buffer matching the cache line size should be used, with accesses aligned to cache lines.
Hopefully, std::rotate is written in such a way that it is compiled into such a sequence of machine instructions.
You want that, but can be tricky because the from and to regions may have different alignment.
Also, the XOR trick introduces data dependencies. That slows down pipelined CPUs.