Knight’s Landing: Atom with AVX-512
chipsandcheese.com
chipsandcheese.com
40% of 2.93 mm2 per core is AVX-512, so 1.14 mm2. This is a large fraction, but as the article says the core is basically a minimum wrapper around the vector unit, with rather weak L1i/branch predictor/store buffers.
Let's put that in the context of modern chips. 14nm density was 44.67 MTr/mm2 so that's 50.9 MTr for AVX-512. To compare with 5nm, let's use density of TSMC N5 (138.2 MTr/mm2) to get 0.37 mm2.
So that's about 10% of an Apple M1 Firestorm core to enable 5-10x speedups vs scalar code. Sounds worthwhile to me. We can now stop saying that "AVX-512 is a huge fraction of modern cores", and "give us more cores instead". Let's instead use the hardware we have :)
Personally, I bought HEDT (Skylake-X and Cascadelake) because I wanted 2x 512 bit AVX512. Glad it's cheap in terms of area, and I'm hoping we'll get more options with great vector performance in the future.
I share your hope for more focus on vectors. It's also up to us software devs, CPUs will not invest as heavily if we don't use it.
Maybe we'll even see quadruple pumping for AVX-512 some day? I'll be impressed if/when an Atom line CPU gets 4x 128bit fma units to match ARM's Cortex-X line or Apple's Firestorm).
I think these are good options, and can allow AVX-512 to sort of act like SVE, but with the benefits of a fixed size architecture (i.e., shuffles); compile one set of code and you're able to run it anywhere, with performance dictated by how much the vendor decided was worth investing into the vector units. And AVX512 can still help (like it does Genoa) by taking a lot of pressure off of the front end.
I'm also trying to do my part as a software dev! I wrote/maintain the JuliaSIMD ecosystem, and working on good loop vectorizers to let people take advantage of their vector units is my passion; LoopVectorization.jl has gotten great results on many benchmarks[0], and I'm rewriting it as an LLVM pass to try and address as many of its flaws and limitations as I can.
[0] For example, in a simple self-dot-product benchmark, LLVM's 256 bit code is actually faster than its 512 bit code when testing random sizes from 1-256: https://github.com/JuliaSIMD/LoopVectorization.jl/issues/446... However, LoopVectorization.jl's 512 bit code is close to twice as fast as either LLVM's 256 bit or 512 bit code. This is a trivial example; the difference can be much larger for more complicated code.
That makes a lot of sense. It's basically the equivalent of RISC-V's LMUL=4, with the big advantage of reducing instruction count as you say. That seems a better route than 4x128, which might actually be less in practice if there are resource conflicts.
> I'm also trying to do my part as a software dev! I wrote/maintain the JuliaSIMD ecosystem
That's awesome, congrats on the good result. Looks like your preference is to allow people to write high-level code without much worry about the arch details. Any thoughts on how we can spread awareness of the basics such as data-oriented programming (avoiding branches, optimizing for cache and contiguous memory accesses)?
That should be a pretty obvious conclusion--higher-clocked cores is better than wider cores, and wider cores is better than more cores.
Communication cost between concurrency domains for superscalar/supervector: min <1ns
Communication cost between concurrency domains for multiprocessors: min 10-100ns
I agree about the communication cost. Higher clocks are harder - quadratic increase in power. For shared-nothing problems, an array of wimpy (low frequency) cores is pretty good. It seems to me that wider, lower-frequency cores are a good compromise: beefy enough (thanks to vectors) to reduce the number of cores required, while still power-efficient due to both vectors and lower clocks.
The ones in Sapphire Rapids run 2x+ faster. Likely there is more pipelining and many more transistors used to reach those clock speeds.
If what you do can use a wide SIMD pipeline, then, by all means, get AVX-512.
If what you want is to have lots of branchy processes running serving different things, then more cores is a better idea.
I'd suggest the best of both worlds - asymmetric cores. Put a couple that have wide SIMD pipelines, alongside a couple others that don't, but may be smaller and more numerous.
AVX-512 tasks that are scheduled to non-AVX-512 trap and are scheduled to AVX-512 cores. Tasks that don't trap continue running wherever the kernel thinks is best.
> Intel spends nearly 40% of the core’s die area to implement wide AVX-512 execution, so Knight’s Landing gets some incredibly high throughput for a low power architecture. In fact, it almost feels like a small OoO core built around feeding giant vector units
Well said!
There's some interesting details here. I love how the article tests SMT1 (aka no SMT), SMT2, and SMT4, and shows how things change. The re-order resources getting subdivided as SMT scales up is obvious but a neat hack, a real neat hack, for repartitioning core resources- on a core (Atom) that wasnt doing reordering much at all, at that point.
Also a good reminder that this almost 10 year old chip was using DDR4. I'd forgotten that DDR4 has been around for so so long!
The chip has some really monstrous capabilities. Modern flagship consumer GPUs are hitting the 1Tbps mark in memory bandwidth. AMD's upcoming Navi 3 has 96MB of L3 cache good for 5.3TB/s. Meanwhile, here's a chip that's 10 years older that has 16GB of onboard DRAM that C&C gets up to 4.5TB/s. Damn ya'll. And with huge AVX-512x2 per core, as mentioned in the top quote... this thing was such a wonderful & fascinating beast.
I didn't realize the previous incarnation- which I bought used/hella-cheap a long time ago & wanted to toy with but never did- was P54C based, aka an original-ish Pentium. That's wild.
This article is just so good. On and on, with every little characteristic and quirk. This is like a perfect "Speaking for the Dead", out of Enders Game. Alas that the world just wasn't awesome enough to put this shit to use at real volume. It's like an ultra-flexible on-the-fly reconfiguring DSP?
A highly specialized processor that has very high computational throughput for specialized operations, but a quite limited scalar unit.
In both cases you really have to write software specifically for it to get it to perform with any reasonable speed. If you do that, you get great value, but any existing code will need significant modifications to perform.
Granted AVX512 is (now) more common than SPE code ever became. It is slightly better than the Itanium approach, but scalar performance (especially single threaded) will have limited the value you get from this CPU, unless you write your own software and it can utilize AVX-512.
The Phi was closer to the Sun Niagara family - lots and lots of simple, slow, cores, with the note that, in the case of the Phi, the weakling x86's had mighty SIMD abilities while the Niagara had more or less standard SPARC stuff.
Neither will have amazing performance unless you have at least as many threads running as you have cores, and most of the time, at least twice as many. For single-threaded code, they were on the slow side.
Still, I always suggested people use Phis to develop because they'd get a taste of future computers. Nowadays a decent laptop will have half a dozen cores and, unless your mail client has that many threads, it'll not feel as fast as it could be.
Unfortunately, I never saw used ones on eBay.
https://guide.handmade-seattle.com/c/2019/talks/lifecycle-of...
While they were quirky at the time (and I got some neat simulations out of them), they were a massive pain in the ass in every other way:
- Requires a motherboard that supports large PCIe BAR addressing, much more common now, not at release time (and good luck getting this working with multiple Xeon Phi coprocessors on a non-server motherboard) - You likely need enough host memory to cover each of the cards (so 8GB * number of cards) - You'll need to patch the Intel MPSS kernel driver to work with newer kernels - You'll need to patch the Intel MPSS tools to work on anything that isn't ~2015 RHEL - You'll want to buy the water cooling kit that alphacool made for them, it's the only way to keep temperatures and noise to a reasonable level
Solve all of those, and you get a (fairly cheap) coprocessor reasonably good at parallel yet branchy tasks. It's much easier to use the second generation of KNL (or third-generation KNM) which share the standard X86_64 ABI and can boot a modern Linux distro and you can use a normal toolchain).
Never found Knights Mill available anywhere. My understanding is that they were only available to integrators in the HPC space.
But it is very interesting to see that these CPUs’ optimizations for high throughput by adopting near-die MCDRAM. This strategy is very similar to what we see in Apple’s M1 series chips today.
Path tracing is usually bi-directional (from both light and camera) and uses something like Metropolis-Hastimgs to sample more efficiently than brute force. Our still requires many samples per pixel but the sampling nature means it can naturally support the random bouncing of diffuse reflections. Between bounces the is typically still ray tracing, though, unless it's some special volumetric material like Jade (the green jewel rock) or fog or (the contents of) a glass of milk.
Compare e.g. POV-Ray to LuxRenderer.