AVX512 is the future, but maybe not Intel's future, ironically.
AVX512 is the future, but maybe not Intel's future, ironically.
If I need to crunch large amounts of data in a hurry, I'll send it to a GPU. The CPU has no chance to compete with that.
I honestly don't understand who/what AVX512 is really for, other than artificial benchmarks that are intentionally engineered to depend on AVX512 far more often than any real-world application would.
I'm sure those cases exist, but they don't justify such a large chunk of silicon. And they definitely don't justify slowing down the rest of the CPU.
Inspector says that this web page, we're talking on is 48kB in size. That's small enough to fit inside L1 cache. Every single connection to the Web Server needs its own AES-key for encryption sake. The 48kB of data (such as my posts above, or your posts) are hot in Cache, but if 100 visitors to this webpage need to get it, HTTPS needs to happen 100x different times with 100x different keys.
So this 48kB text message (consisting of all the comments of this page) are going to have to be encrypted with different, random, AES keys to deliver the messages to you or me. AES operates on 16-bytes at a time, and AES-GCM is a newer algorithm that allows for all 48,000+ bytes to be processed in parallel.
AVX-512 AES instructions are ideal for processing this data, are they not? And processing them 4x faster (since 4x instances of AES are occurring in parallel, since AVX512 can work on 64 bytes per tick / 16 bytes per AES instance), is a lot better than just doing it 16-bytes at at time with the legacy AES-NI instructions.
----------
Despite being a parallel problem, this will never be worthwhile to send to the GPU. First, GPUs don't have AES-ni instructions. But even if they did, it would take longer to talk to the GPU than for the AVX512-AES instructions to operate on the 48kB of data (again: ~40,000 clock ticks to just start talking with the GPU in practice). In that amount of time, you would have finished encrypting the payload and have sent it off.
Probably more important, there's a 256 bit version of that instruction. You can get half of that extreme throughput without AVX-512.
http://nabstreamingsummit.com/wp-content/uploads/2022/05/202...
And a surprising amount of it was in TLS optimizations, in particular, offloading TLS to the hardware (Apparently Mellanox ConnectX ethernet adapters can do AES offload now, so the CPU doesn't have to worry about it).
Since Mellanox ConnectX adapters are trying to solve the AES problem still, I have to imagine that its a significant portion of a lot of server's workloads. Intel / AMD are obviously interested in it enough to upgrade AES to 4x wide in the AVX512 instruction set.
I can't say its particularly useful in any of _my_ workloads. But it seems to come up enough in those hyper-optimized web servers / presentations.
It takes literally 1 to 10 microseconds (10,000 nanoseconds) to talk to a GPU over PCIe.
In the 40,000 clock cycles, you could have processed 2.5 MB of data with AVX512 instructions *BEFORE* the GPU even is aware that you're talking to it. Then you gotta start passing the data to the GPU, the GPU then has to process the data, and then it has to send it all back.
All in all, SIMD instructions on CPU-side are worthwhile for anything less than 8MB for sure, maybe less than 16MB or 32MB, depending on various details.
----------
That's one core. If your one-core machine can talk to the other 32-cores or 128-cores of your computer (see 64-core EPYC dual-socket computers), you can communicate to another core in just 50-nanoseconds or so (~200 clock penalty), and those other cores can be processing AVX512 as well.
So if you're able to use parallel programming on a CPU, its probably closer to 1GB+ of data before it actually is truly an 'obvious' choice to talk to the GPU, rather than just keeping it on CPU-only side.
---------
Example of practical use: AES-GCM can be processed using AVX512 in parallel (each AES-GCM block is a parallel instance, the entire AES-GCM stream is in parallel), but no one will actually use GPUs to process this... because AES-instructions are single-clock tick (or faster!!) on modern CPUs like Intel / AMD Zen.
That's just going to happen whenever you go to an TLS1.2 or HTTPS instance, which is pretty much all the time? Like, every single byte coming out of Youtube is over HTTPS these days and needs to be decrypted before further processing.
Which would you rather have: some fixed-function unit shared between all cores (load balancing? what if you're suddenly doing crypto stuff on many cores?), or the general-purpose tools for running any algorithm on any core?
The number of crypto units would be SKU specific depending on the workload. A server box doing service mesh would need one per concurrent flow presumably.
The thing that accelerator offload gets you is an additional thermal budget to spend on general purpose workloads. If you know that you will be doing SERDES, and enc/dec, offloading those to an accelerator frees up watts of TDP (thermal design power) to spend other places. This is also why we see big/little architectures, the OOO processors suck up a lot of power. In-order cores are just fine for latency insensitive workloads.
And most encryption is to use "something" to get an AES key and then us that to decrypt data.
And that "something" (e.g. RSA/ECC based approaches) is what is threatened by quantum computing. But it's also not overly problematic if that "something" becomes slower to compute as it's done "only" once per-"a bunch of data".
AFIK the situation is similar for signing, as you don't sign the data itself normally but instead sign a hash of it and I think the hashing algorithms mostly used are not threatened by quantum computing either.
As to quantum, it looks like practical serial or parallel application of Grover's algorithm might still be decades away. But that is with current knowledge, and who knows what other breakthroughs will be made.
NSA: "You are correct!"
The CPU has an overhead of about ~10us to enable the AVX512 units.
It also dramatically reduces the clock on other cores.
For more information, see: https://stackoverflow.com/a/56861355
Timing information: https://www.agner.org/optimize/microarchitecture.pdf
Its all about AMD Zen4 or Xeon Ice Lake+, which has no clock reduction and no overheads.
On Alder Lake (pg 172).
> The reader is referred to the timings for Tiger Lake and Gracemont.
On Tiger Lake (pg 167):
> Warm-up period for ZMM vector instructions
> The processor puts the upper parts of the 512 bit vector execution units into a low power mode when they are not used.
> Instructions with 512-bit vectors have a throughput that is approximately 4.5 times slower than normal during an initial warm-up period of approximately 50,000 clock cycles.
I'm not saying you are wrong. I just haven't heard about that.
> Since 512-bit instructions are reusing the same 256-bit hardware, 512-bit does not come with additional thermal issues. There is no artificial throttling like on Intel chips.
At least for Zen4, there's no worries about throttling or anything really. Its the same AVX hardware, "double pumped" (two 256-bit micro-instructions output per single 512-bit instruction). But you still save significantly on the decoder (ie: the "other" hyperthread can use the decoder in the core to keep executing its scalar code at full speed, since your hyperthread is barely executing any instructions)
Which does look like hilarious waste of silicon, it could cut all of the float/division transistors off and still be plenty useful
PCIe is real bottleneck that affects bandwidth and latency. It only really works if your data is already resident on the GPU, your kernels are fixed and the amount of data returned is small.
With HBM and VCache, we will see main memory bandwidth over 1TB/s for consumer high perf cpus in the near future, at those rates, GPUs won't make sense. GPUs are basically ASICs that can take advantage of hundreds of GB/s of memory bandwidth, when the CPU can do scans at that same rate, the necessity of a GPU is greatly reduced.
Latency is a hell of a thing. Anything over a millisecond is an absolute eternity, and I know the GPU imposes more than that for most practical applications - especially gaming.