Transcoding Unicode with AVX-512: AMD Zen 4 vs. Intel Ice Lake
lemire.me
lemire.me
The fact we are still using UTF-16 still irks me to this day. UTF-16 (which is actually two different encodings, not one, hence the need for a BOM) is basically a way to salvage all those platforms that hurried on the UCS-2 (aka, the original "UNICODE") bandwagon in the '90s hoping that by just doing s/char/wchar_t/g all their internationalization problems would be solved.
It did not go well, to say the least. UTF-16 and 32 are objectively worse than UTF-8, because they still are multibyte encodings, you still have to do normalization, ... while also having to deal with "char16_t" and the likes (I will not enter into the whole "TCHAR" fiasco).
Spoiler alert, the world is still full of allegedly "UTF-16 compliant" platforms out there are not, indeed, UTF-16 compliant, they just use 16 bit chars and hope for the best.
The whole idea "1 char = 1 character" is arguably a terrible idea in Unicode, though. You not only can have multibyte characters, but you can also have multirune characters, where multiple codepoints are normalized into a single displayed character (just think about ` + e = è). It's a mess and it's bound to be broken, and there's no real way to "fix that up" - ASCII's assumption of "1 value = 1 char" was the broken concept here, and it unfortunately flawed how every developer (me included) thinks about strings. Sigh.
https://manishearth.github.io/blog/2017/01/14/stop-ascribing...
Windows, unfortunately.
This is unlike RISC-V V extension, where the same code will run and utilize the hardware, regardless of vector unit width.
However, a lot of software is compiled on one machine to be run on potentially many possible architectures, so they target a very lowest common denominator arch like x86-64. This will have some SIMD instructions but (I don't think) AVX-512.
So if a developer wants to ensure those instructions are used if they're supported, they'll write two code paths. one path will explicitly call the avx512 instructions with compiler intrinsics and then the other path will just use the manual code and let the compiler decide how to turn it into x86-64 safe instructions.
``` void myNotOptimizedThing(my_data* d){ _SPECIAL_CPU_MANUFACTURER_0X3D512(d); } ```
edit: and include some header from the manufacturer most likely?
In my experience, clang unrolls too much, so you end up spending all your time in the non-vectorized remainder. Using smaller vectors cuts the size of the non-vectorized remainders in half, so smaller vectors often give better performance for that reason. (Unrolling less could have the same effect while decreasing code size, but alas)
I am not sure the really interesting AVX-512 instructions have intrinsics yet. For those it's asm or nothing.
E.g. https://www.intel.com/content/www/us/en/develop/documentatio...
Ah, but this repo mentions that the GCC 11 implementation apparently also works with clang: https://github.com/VcDevel/Vc. Thanks!
I'd imagine outputting optimized avx code from an existing C for() loop would be much easier than going from a "write me a python code that..." prompt.
Compiler autovectorizers also aren't very good at producing fast AVX512 code, so most of the benefit would probably come from using optimized libraries like Intel's MKL or simdjson.
Any installation of Gentoo is, presumably. (Otherwise, what's the point of compiling it all yourself?)
More interestingly, possibly all OEM firmware-installed copies of ChromeOS are -march=native builds as well, given that ChromeOS is based off of a Gentoo upstream.
Clear Linux is probably a more practical alternative. I used it a couple years ago, and found that they had a lot of avx2 and avx512 versions of random libraries built, with the appropriate ones presumably being loaded based on the hardware.
Random glibc math function calls, for example, were much faster on Clear Linux than Arch or Fedora. But development of Clear seems to have stopped, libraries like llvm aren't being updated anymore so the toolchains are outdated. I'd wanted to avoid the blood and sweat of managing my own toolchains, and ironically being on bleeding edge distros (Arch,Fedora,etc) was the way to keep that to a minimum. Next time I reinstall an OS, I'll look at Clear again. Or maybe Guix or Nix. Or maybe use spack for package management on top of some other distro.
https://github.com/openzfs/zfs/pull/14234#issuecomment-13345...
A bug report has been filed with GCC for one of the issues. LLVM is much better here, but not perfect, or at least that has been my experience when trying to have the compiler generate assembly for an explicitly vectorized fletcher4 implementation.
That said, many of the avx512 instructions are simply extended width AVX2/avx2 instructions. The interesting things about it are really the increased width and the additional registers. Not many of the new instructions that are not bitwidtg extended versions of the old ones are particularly interesting since Intel had already implemented most of the interesting things for smaller vector widths.
Here is a quite recent introduction to RISC-V Vector[0].
0. https://erikexplores.substack.com/p/grokking-risc-v-vector-p...
The former is nifty, though intended for single-issue machines, and the latter seems redundant because masks can also do that.
If you otherwise want vector (V extension), "right now" would limit you to pre-1.0 V extension implementations.
If you need to license hardware IP, there are several very high performance implementations as of RISC-V Summit[0]. Actual hardware will pop up throughout 2023.
So yeah, even in RISC-V V, vrgather has explicitly different per-element operation depending on VLMAX, which obviously depends on the HW's VLEN. So depending on the table size, you have to assume constraints on VLEN or execute different permute sequences.
Um... Ice Lake shipped over three years ago. I mean, there's a real question as to whether or not "senselessly wide SIMD" is a good or bad feature in a datacenter part, and how or whether AMD should attempt to implement it and within which market sectors. And surely there's discussion to be had about the design tradeoffs to be made chasing after this nonsense.
But, no, performance parity with chips that are nearing end of life has to be viewed as table stakes here. It's certainly not "remarkable".
Comparing apples to oranges is not how you determine how good a peach is.
Those differ only in the clock frequency and in a few details that are irrelevant for this particular benchmark.
So the comparison normalized by clock frequency, as done here, is legit, no other better comparison is possible for now.
Even if someone had a Sapphire Rapids sample, they would not be allowed to publish any benchmark yet.
I never looked into it in detail, so I could be mistaken.
This is also why it’s impressive that this is only AMD’s first attempt: it appears to work really well, where it took Intel multiple attempts to get it working well.
AVX-512 is great. The new instructions are incredibly useful for a wide variety of applications, but the fragmentation and segmentation by Intel has made it a total mess.
But hey this time it literally gave AMD time to catch up...
As others mentioned, throttling is basically nonexistent on Icelake (and AMD Genoa). It can hurt on Xeon Silver (so let's not use those?) and if you only sporadically use SIMD instructions (again, don't do that).
I claim that just about any reasonable code which sustains SIMD instructions over several milliseconds would still be a net win even with throttling.
Un-nuanced concerns about throttling are outdated and unhelpful. Perhaps I'll write up a paper on this.
I understand that dark silicon is helpful, but am not so sure that fixed-function HW is the way to go. Perhaps video _de_coding is the most convincing from your list; codec generations are 5+ years, so enough time to benefit from HW. Encoding, on the other hand, tends not to be impressive unless perhaps there is also a software component.
For the rest, programmability and deployability (can we rely on it?) is a major issue. Software has often been the limiting factor.
Another big concern is the 'hardware lottery'. The algorithms we develop and get are selected for, and tuned to, the current hardware. Perhaps this gets us 5x energy efficiency vs CPU/SIMD. But by painting ourselves ever further into the corner of dense linear algebra, which is definitely not the way that nature implements intelligence, we are missing out on far larger opportunities. For example: spiking nets or memristors have the potential to be 2 or 3 orders of magnitude better. Or actual sparsity, not the fixed-pattern thing (now that is a prime example of an irrelevent benchmark, because AFAIK algorithms haven't yet been able to use them well).
> A general purpose CPU however should specialize on unpredictable memory access.
Should it really? I think rather we should avoid such accesses whenever possible, because their energy cost now dwarfs that of computation.
> This means AVX-512 is somewhat misplaced on a CPU and probably only exists because it served Intel to create nice numbers in irrelevant benchmarks.
I have difficulty understanding how a reasonable person can come to such a conclusion. Lemire (the author linked here) has a long series of results showing nice speedups from AVX-512. I personally have seen gains in image compression, string processing, cryptography, linear algebra, integer coding, hash tables, databases, sorting, and compression.
[Opinions are my own.]
The applications are very niche. Compilers are usually not smart enough to utilize SIMD, it is a hit or miss. And in order to implement properly efficient SIMD algorithms you need experts that are rare. Furthermore many algorithms that work great with SIMD work even better as compute shader on your run of the mill cheap iGPU.
The application of this article is the best example how irrelevant SIMD really is: How many Terabytes of UTF8 are you converting to UTF16 per day? probably zero.
> in order to implement properly efficient SIMD algorithms you need experts that are rare
Some truth to this, but many algorithms can be implemented once and then reused, like a standard library.
> many algorithms that work great with SIMD work even better as compute shader on your run of the mill cheap iGPU
Also agree to some extent, except that you'd have more concerns about availability, vendor lock-in, and performance portability.
> best example how irrelevant SIMD really is: How many Terabytes of UTF8 are you converting to UTF16 per day? probably zero.
First, how does one example of a SIMD-enabled algorithm show that SIMD itself is irrelevant? Second, have you considered that some databases store UTF-16 and want to convert it for interoperability (or vice versa)? IBM apparently has dedicated instructions for this. Would they have been added if there was no demand?
a 240+ entry reorder buffer has entered the chat.
Users started dumb, develooers became dumb, now you want CPUs dumb too? Who will deal with the consequences??