If this is all there is to it, why do they have the high frequency and high l3 cache? Those seem to be optimizing for something, not just a “good enough” configuration for a part that is not the bottleneck
Data augmentation in CPU-space is often compute-light, but requires rapid access to memory. There are libraries (like NVIDIA's Dali) that can do augmentation on the GPU, but this takes up GPU resources that could be used by training. Having a multi-core CPU with fast caches is a good compromise.