Basically you are asking for DCAS (atomic compare exchange on two memory locations). There is a reason the last architecture (IIRC) to support it was the 680x0, the required extensions to the coherence protocol are hard to get right without deadlocking the machine and even harder to make it perform.
AMD internally had a design of a restricted HTM that would have guaranteed exclusive access to up to 7 cachelines (using an ad-hoc arbiter), enough to implement ana atomic and dequeues between two lists; they ended up documenting a variant that would only have guaranteed best effort and and are yet to implement anything so far.
Interestingly Intel CPUs do lock 2 cachelines in the case of an unaligned RMW spanning two cachelines. It takes thousands of clock cycles to complete and stalls the whole system; it exists only for compatibility with legacy software and it's use is strongly discouraged.