But then those accesses have to be protected by a lock in order to enforce invariants between "start" and "end". The whole point of representing both with a single atomic value is that this makes it trivial to preserve those required invariants.
In fact the other side value is normally cached in a local memory location so to amortize the cost of roundtrips over multiple reads and writes.
Cmon AMD/Intel, give us hardware queues already :)