Parsing JSON faster with Intel AVX-512
lemire.me
lemire.me
> Of course, to get these new benefits, you need recent Intel processors with adequate AVX-512 support
AVX-512 support can be confusing because it’s often referred to as a single instruction set.
AVX-512 is actually a large family of instructions that have different availability depending on the CPU. It’s not enough to say that a CPU has AVX-512 because it’s not a binary question. You have to know which AVX-512 instructions are supported on a particular CPU.
Wikipedia has a partial chart of AVX-512 support by CPU: https://en.wikipedia.org/wiki/AVX-512#CPUs_with_AVX-512
Note that some instructions that are available in one generation of CPU are can actually be unavailable (superseded, usually) in the next generation of the CPU. If you go deep enough into AVX-512 optimization, you essentially end up targeting a specific CPU for the code. This is not a big deal if you’re deploying software to 10,000 carefully controlled cloud servers with known specifications, but it makes general use and especially consumer use much harder.
I know you can do this yourself, but last time I looked it was a heavily manual process— you had to basically define a plugin interface and dynamically load your selected implementation from a separate shared object. What are the barriers to having compilers able to be hinted into transparently generating multiple versions of key functions?
This LWN article has some additional information: https://lwn.net/Articles/691932/
Clang added it in 7.0.0: https://releases.llvm.org/7.0.0/tools/clang/docs/AttributeRe...
A nice presentation on it: https://llvm.org/devmtg/2014-10/Slides/Christopher-Function%...
https://gregoryszorc.com/blog/2022/01/09/bulk-analyze-linux-...
Even Intel’s Clear Linux does not bother to patch individual libraries. They just use the glibc library multi-versioning feature to load from different directories depending on the cpuid.
In my opinion, GCC and Clang could make the whole thing more ergonomic. Ideally you declare a function as "interesting for multi-versioning" in the source using an attribute, then in the command line define what -march's to actually clone for. Kinda like ICC’s /Qax. (On second thought preprocessor defs are sufficient, duh.)
For example, ripgrep's dependencies dispatch dynamically at runtime by querying CPUID, but nothing uses GCC's "IFUNC" thingy. So it's likely that much more software is utilizing target specific code than not. Still, it's probably less than one would like.
(I think this is less of a response to you and more of a response to this entire thread. It seems like some folks are conflating "IFUNC" with "all forms of dynamic dispatching based on CPUID.")
[1] https://www.singlestore.com/blog/a-programmers-perspective/ comments https://news.ycombinator.com/item?id=28179111
This has been broadly overplayed. There are two main families: consumer AVX-512 and server AVX-512. Server gets some additional VNNI instructions and BFLOAT16 support, consumer gets VPOPCNTDQ and IFMA/VBMI. That's basically all you need to know
The chart looks confusing in Wikipedia because it's ordered by date, not by "series" and generation, and some intrepid wikipedia editors have munged up the features even then. This is what the feature set looks like to me:
https://i.imgur.com/idAjB1X.png
Xeon Phi is sorta its own thing for sure, they did a weird 4-wide architecture and it got special 4-wide VNNI instructions to match. And since it's been abandoned it never was updated. But there's zero reason for average consumer/server software to target Xeon Phi anyway these days - it's abandoned and it was almost entirely a supercomputing/HPC product even when it wasn't.
The other instructions are (almost) a straight superset within their series, so you just need to know if you're targeting server or consumer. Laptop gets things a bit earlier because Intel started 10nm first on laptops, and 14nm was stuck with older architectures and then a backported architecture, so it's not quite a strict series in terms of dates, but generation-on-generation they haven't ever regressed a feature within a given "series".
The only exception is Alder Lake doesn't get anything because lol Intel - they disabled AVX-512 entirely, but if you got a model that supports it then it's also a superset of all previous consumer generations.
Again, all you really need to remember is, consumer gets better FMA and server gets VNNI/BFLOAT, that's the major difference. Big deal, who cares.
But "it's just strict supersets of features within server/client" wouldn't get the clicks/exasperated outcry of "gosh isn't this just an intractible mess!". It's not, and they've never regressed a feature within a series.
Also for the record, AVX2/AVX1/SSE/MMX look really shitty if you split them up in this same fashion. There were like 4 different "sets" of SSE4 feature sets alone and some of them were abandoned and never used again, and both AMD and Intel adopted them all at different dates and sometimes dropped them back out, so if you made a "feature set vs architecture" table it would be a giant fucking mess for SSE as well, let alone if you did it by year instead of by architectural series.
Yet I never hear anyone bring up how awful and how confusing AVX and SSE are, just AVX-512. Feature sets are messy!
https://en.wikipedia.org/wiki/SSE4
Eventually, AVX-512 support will converge around the sets of features that are useful, and some of the less-useful features may eventually be deprecated, just like the SSE4a instruction set did. It's gonna be fine.
AMD should be implementing AVX-512 on their own cores soon as well. Once Armv9 (with SVE2) becomes dominant, we'll pretty much be in a golden age of SIMD.
[0]: https://chipsandcheese.com/2022/04/30/examining-centaur-chas...
We already are in the golden age of SIMD. NVidia and AMD GPUs are easier and easier to program through standard interfaces.
Intel / AMD are pushing SIMD on a CPU, which is useful for sure, but always is going to be smaller in scope than a dedicated SIMD-processor like A100, 3060, AMD Vega, AMD 6800 xt and the like.
SIMD-on-a-CPU is useful because you can perform SIMD over the L1 cache as communication (rather than traversing L1 -> L2 -> L3 -> DDR4 / PCIe -> GPU VRAM -> GPU Registers -> SIMD, and back). But if you have a large-scale operation that can work SIMD, the GPU-traversal absolutely works and is commonly done.
AVX512 and ARM SVE2 bring the CPU up to parity with maybe 2010s-era GPUs or so (full gather/scatter, more permutation instructions, etc. etc.). But GPUs continued to evolve. Butterfly-shuffles are the generic any-to-any network building block, and are exposed in PTX (NVidia assembly) shfl.bfly, and AMD DPP (data-parallel primitives).
Having a richer set of lane-to-lane shuffling (especially ready-to-use butterfly networks) would be best. It really is surprising how many problems require those rich-sets of data-movement instructions, or otherwise benefit from them.
NEON and SVE had hard-coded data-movement for specific applications. The general-purpose instruction (pshufb) is kinda like permute/shfl from AMD/NVidia. A backwards-permute IIRC doesn't exist yet on CPU-side.
And butterfly networks are the general-purpose solution, capable of implementing any arbitrary data-movement in just log(width) steps. (pshufb / permute instructions would be the full-sized butterfly network, but some cases might be "easier" and faster to execute with only a limited number of butterfly swaps, such as what inevitably comes up in sorting)
--------
Still, all of these operations can be implemented in AVX2 (albeit slower / less efficiently). So its not like the "language" of AVX2 / AVX is incomplete... its just missing a few general-purpose instructions that could lead to better performance.
Which to me is reasonable for ISA compatibility. (Also considering that having to deal with ISA incompatibility across active cores is not reasonable at all.)
If the OS would give me the tools to deal with it, I would find it completely reasonable to write (eg) both avx2 and avx512 kernels to be run concurrently on the same machine.
So the only reasonable options for a big little system are:
1. Not have a big little system. But little cores are quite useful for power efficiency in laptops…
2. Don’t support AVX-512 on any cores. Everything beyond AVX2 becomes a dead extension outside of server chips.
3. The little cores support AVX-512 as best as is reasonable. Then thread director can weight AVX-512 usage even heavier than it already weights AVX2.
Also AVX2 performance on Alder Lake is already different between the cores enough that optimal implementations can be different.
On the other hand, the tooling for turning high-level languages into SIMD code is not there yet, ISPC refuses to support ARM, and is still kind of a novelty tool.
Additionally, 512-bit wide vectors are just too big - the resulting vector units take up too much die space even on big cores, and the power consumption causes issues causing said dies to downclock. Probably it won't be viable on small cores.
This is no longer true, citing [0]:
> At least, it means we need to adjust our mental model of the frequency related cost of AVX-512 instructions. Rather than the prior-generation verdict of “AVX-512 generally causes significant downclocking”, on these Ice Lake and Rocket Lake client chips we can say that AVX-512 causes insignificant (usually, none at all) license-based downclocking and I expect this to be true on other ICL and RKL client chips as well.
And we still have to see AMDs implementation of AVX512 on Zen4 to know what behavior and limits it may have (if any).
[0] https://travisdowns.github.io/blog/2020/08/19/icl-avx512-fre...
What is meant by "relatively recent C++ processors"? Is that supposed to be "compilers"?
At what point is it saner to use something like flatbuffers or capnproto style message encoding instead.
Isn't it better to have all options open?
I’m torn. I’ve worked at shops where we aim over time to reduce response time while serving business logic and using statistical models that get iterated on. Even there I haven’t seen a blatant need for non-JSON rpc. But I know my experience doesn’t mirror everyone’s. And I like seeing and learning about instruction sets. I’m currently taking a course in parallel computing and I just used a avx2 for the first time in a toy program to subtract one vector from another in a single instruction which while not particularly useful is a window into more interesting things and is still SIMD.
I think on the whole making json parsing for a large enough fraction of processors is probably a huge win for the environment. But who is parsing json in C++?
Well, Facebook for one! Folly has lots of utilities for this (see folly::dynamic[1]). We make extensive use of this at my (non-Facebook) job.
[1] https://github.com/facebook/folly/blob/master/folly/docs/Dyn...
https://nee.lv/2021/02/28/How-I-cut-GTA-Online-loading-times...
> Original online mode load time: ~6m flat
> Time with only duplication check patch: 4m 30s
> Time with only JSON parser patch: 2m 50s
> Time with both issues patched: 1m 50s
>
> (6 * 60 - (1 * 60+50)) / (6 * 60) = 69.4% load time improvement (nice!)
I would have loved to see someone swap out the parser for simdjson, especially without the code. That would be amazing. Maybe the 1m50s can be beaten.
Just thinking about that core heating up for 6m one can see that this sort of code improvement helps reduce electricity usage, heat generated, and so on. It's just less computation. I see now that people need to parse JSON in C++ as they would any other language, and someone has to do it, and it should be done as fast and correctly as possible. It's a different issue from defining a performant wire protocol with good developer ergonomics (which to be frank I am not sure why my brain even went there to begin with; I had protobuf on the brain).
Makes sense.
But people seem to think of C++ as only systems language only ignoring the decades of desktop usage and existing software. Plus web services running are using small fractions of the resources that systems like node use. I know in some examples I have done where a C++ web service used about 3MB of ram, plus usage depending on size of the request, whereas Node started at 300MB. With cloud costs, that's a lot more bang for buck.
But JSON is the lingua franca of networked I/O these days, need to interop
What's cost/opportunity to optimizing for a specific platform/instruction set ? At what point is it worth doing, when isn't it worth doing ? AVX-512 strikes as something... "ephemeral".
Sure AVX-512 is only applicable to specific workloads, and even many of those workloads the cost/opportunity of optimizing for AVX-512 might not be worth it. But there clearly ARE usecases that would benefit, and it might be worth it for more consumer applications to optimize for AVX-512 - but only if it can be used.
The way I see it is that the benefit of optimizing for AVX-512 is far higher if it becomes normal for consumer CPUs to have it. A 28% improvement is pretty decent, but it's only worth implementing if enough people can utilize it.
The only thing they've done is disable it in the hybrid Alder Lake cores, presumably because the E-cores couldn't support it (while the P-cores could), and they didn't want to deal with the headaches of ISA extensions being supported only on some cores in the system.
There are 0 current generation consumer CPUs of neither Intel nor AMD that have it
> The only thing they've done is disable it in the hybrid Alder Lake cores
Which happen to be all the current generation Intel CPUs
That is incorrect. You can buy Alder Lake CPUs that only have one type of core (the i3 series only has P-cores, for example), and those do not support AVX-512 either. They're not "hybrid" in any way.
Some of their motherboard partners initially allowed you to access AVX-512, but Intel has put a stop to this and the feature is disabled on all Alder Lake CPU SKUs, period.
Newer Alder Lake chips have AVX-512 fused off in silicon, if the firmware blob disabling it isn’t enough: https://www.tomshardware.com/news/intel-nukes-alder-lake-avx...
> Intel is not dropping support for it on a lot of CPUs.
That seems like a pretty questionable statement. Intel might keep AVX-512 around for Xeon, but it seems extremely dead on the consumer market. If Intel decides to bring it back for the next generation, that would be strange and very poor planning.
Even comparing common deserializarion libraries like gson and thrift, thrift is faster despite being much older than protobufs.
When upgrading vectors from SSE to AVX, the speedup is very rarely by a factor of 2. More often it's within 30-70%.
I never programmed AVX512, but I would expect for real-life code AVX512/AVX2 performance improvement to be much closer to AVX/SSE than to SSE/scalar.
Another one is bandwidth. On computers with dual-channel DDR4, aligned SSE vectors are delivered from main memory with a single transaction. For ideal scaling from SSE, AVX512 would need octa-channel DDR4, and 4x as much caches.
Another one is waste heat. On modern processors, sustained performance is often limited by thermals. For ideal scaling, AVX512 would need twice as much cooling compared to AVX.