A consumer CPU like a 285k caps out around 130 GB/s of memory bandwidth.
Each of its 24 cores can do two 8 wide FMA ops per cycle. Lets say holding a continuous 4 GHz clock speed.
This works out to over 1.5 trillion 32 bit floating point multiplies per cycle.
If you are doing vector matrix multiplies (like in single token no batching). `xW` then each weight loaded sort of gets used in 1 multiplication and 1 addition.
Doing the math you can clearly see even if each weight were just 1 byte you can at most load 130 billion of them in a second from memory.
But in the same timespan you could have done over 1.5 trillion multiplications.
So you are still memory bound.