Which would you rather have: some fixed-function unit shared between all cores (load balancing? what if you're suddenly doing crypto stuff on many cores?), or the general-purpose tools for running any algorithm on any core?
The number of crypto units would be SKU specific depending on the workload. A server box doing service mesh would need one per concurrent flow presumably.
The thing that accelerator offload gets you is an additional thermal budget to spend on general purpose workloads. If you know that you will be doing SERDES, and enc/dec, offloading those to an accelerator frees up watts of TDP (thermal design power) to spend other places. This is also why we see big/little architectures, the OOO processors suck up a lot of power. In-order cores are just fine for latency insensitive workloads.
And most encryption is to use "something" to get an AES key and then us that to decrypt data.
And that "something" (e.g. RSA/ECC based approaches) is what is threatened by quantum computing. But it's also not overly problematic if that "something" becomes slower to compute as it's done "only" once per-"a bunch of data".
AFIK the situation is similar for signing, as you don't sign the data itself normally but instead sign a hash of it and I think the hashing algorithms mostly used are not threatened by quantum computing either.
As to quantum, it looks like practical serial or parallel application of Grover's algorithm might still be decades away. But that is with current knowledge, and who knows what other breakthroughs will be made.
NSA: "You are correct!"