This was on a wide variety of intel, AMD, NUMA, ARM processors with different architectures, OSes and memory configurations.
Part of the reason is hyper threading (or threadripper type archs) but even locking to groups wasn’t usually faster.
This was even moreso the case when you had competing workloads stealing cores from the OS scheduler.