Even if they are not branded Xeon, but Atom, there already are a lot of Intel server CPUs with E-cores, e.g. Denverton Refresh, Snow Ridge, Parker Ridge and the recently launched Arizona Beach.
Now imagine Intel had that product and canceled it in 2017, and you will be living in reality: https://en.wikipedia.org/wiki/Xeon_Phi
Not only would something like 1024 cores with AVX512 be competitive with GPUs it would have the added advantage of being waaay more versatile and easier to program.
And I'm not sure it'll be that much easier to program - "few large cores" and "GPU waves" are now pretty well supported in the stack with mature tooling - trying to insert a new stack between the two would likely be pretty difficult as it doesn't really fit well with either established paradigm so likely needs something new to show it's benefits.
It was a big experiment, and it failed. Some of the brightest people have tried to push it to the limit. In my circle, I have never heard anyone missing Xeon Phi when it is discontinued.
The cores were in order. If you've never written code for in order cores, you may not appreciate exactly how heavily you're leaning on out of order to save your bacon. Things your used to just working suddenly don't.
The cores had a 2D layout, but Intel refused to tell you what it was and in practice it was impossible to optimize for. You ended up with variable (as in different on every node) latency for memory access that made programming again painful (and add this to the fact the cores are in order, so you really can't work around the latency).
Then there was the software stack that was just odd in some ways. I was researching I/O throughout, and things that just work on any other CPU (on any modern ISA) just didn't behave in a sane way.
So yeah, I was not at all sad when Intel cancelled it.
A chip with hundreds or thousands of E-cores should be ideal for any cloud provider to offer vCPUs with a high markup. For example, a company like Hetzner sells 1vCPU services for around 4€/month. Even if they don't overprovision and allocate 1 core per 1vCPU, a single chip can potentially earn them 500€/month, which leads to a break even point at around 2 years.
Why would they put a dig at Intel’s consumer chips in a slide deck for their server parts? The Intel Xeons don’t have E-cores. This doesn’t make any sense unless I’m missing something.
Also, you know AMD has recent consumer chips without AVX-512, right?
Intel server CPUs with E-cores that will use the Xeon brand, because they will have much more cores (6-times more, i.e. 144 vs. 24) than the current models, are announced for 2024.
Xeons with E-cores are planned: https://en.wikipedia.org/wiki/Sierra_Forest
And even if they weren't, it would still be a cheap PR win.
>Also, you know AMD has recent consumer chips without AVX-512, right?
All the current generation processors (Zen4) support AVX-512, both mobile and desktop. It may be confusing because AMD's numbering scheme is intentionally misleading and they sell previous gen chips with new model numbers.
I've been messing with SIMD optimizations recently in a datastructure library I wrote, and as tantalizing as the AVX-512 is, it'll be years before it can be used in production software on a real scale. Introduced in 2017, and still totally unusable in the wild.
as a counterpoint, i'm on a P14s Gen 2 Thinkpad (Ryzen 7 PRO 5850u) and have zero issues on EndeavourOS (Arch, KDE/Plasma). most of my time is in VSCode, SublimeMerge, Chrome, MPV.
i did have to rip out the realtek (or mediatek?) M.2 wifi card and swap in an intel one tho.
Requires a hardware ("pinhole" not power cycle) reset every time I want to switch between external monitor use and laptop only. My bug report/forum post here: https://forums.lenovo.com/t5/ThinkPad-Z-series-Laptops/Exter... -- other people are effected.
Last year they pushed out a BIOS update which caused the fan to run 100% of the time. Then a "fix" which caused it to overheat. And in-between there somewhere something was pushed that caused me to have to reformat the whole thing.
Lid-open-to-awake in Linux used to simply not work 75% of the time, required a hard reset. Now it works, but there's a 5-10 second pause before display turns on (Windows or Linux) vs almost instant on the work issued laptops I've had recently (MBP M1 and some HP Z series thing, running Linux)
Waste of cash. Bought this thing to be my dev workstation when I took a contract job last summer, wish I'd never done that. Software quality / support at Lenovo is a real problem.
According to AMD, the AMD Ryzen 7 PRO 6850H does not have USB type-C support.
This is not my experience. Do you have some justification for this statement? I've found myself far more impressed by autovectorization than disappointed by it. I've found that code most people think is autovectorizable actually violates scalar contracts. But if you write clear code whose scalar implementation won't introduce UB if autovectorized, the compiler is really good.
Here's my go-to example. Agner Fog's VCL is a well-respected library. It has vectorized versions of all sorts of useful mathematitical functions. For fun, I rewrote his `exp(x)` function using his exact algorithm, but in scalar code, and it autovectorized, and benchmarks the same.
And autovectorization is nice, but explicit SIMD intrinsics use usually wins if done right.
If I was getting paid $$ for this work, sure, I'd rent cloud instances or hardware to do that development. But it presents a dilemma for open source work.
Anyways, it's all griping. We'll either eventually all get AVX512, or it will die and some other more common wide vector extension will take its place, or we'll all be having the same gripe about NEON or RISC-V V extensions 10 years from now.
It's just frustrating some of the nice toys that are in AVX512, in particular nice support for bitmasking that would make the code I'm writing much nicer & faster.
> Anyways, it's all griping. We'll either eventually all get AVX512, or it will die and some other more common wide vector extension will take its place, or we'll all be having the same gripe about NEON or RISC-V V extensions 10 years from now.
I think we'll get a unified instruction set. Part of AVX512's difficulty is that it's not just AVX512 or not. It's AVX512F and/or AVX512BW and/or AVX512VL ...x10
Further to the Godbolt suggestion, one band-aid is that Highway fairly efficiently emulates some of the fancier AVX-512 instructions such as CompressStore. You could then develop on AVX2, then rent 1 VCPU-hour to build and verify it indeed works on AVX-512.
As for the subsets of AVX-512, we defined them into groups matching Skylake and Icelake; that works pretty well. Zen4 would also support the Icelake features, but it gets its own target so that we can special-case/avoid the microcoded and super-slow CompressStore there.
I've been using Highway for a couple weeks (I'm the guy who is writing the unroller feature). Highway is more limited in its breadth than raw intrinsics. I've noticed a few instances so far where you had to make a judgement call, and decided not to have Highway expose certain features (like 32 bit indexing into 64 bit type scatter/gather). And with more things that x86/ARM throw at us, the harder I think it becomes for a library to be the solution. From what I've seen of std::simd, I don't see how that possibly can be the solution. What do you think?
Agree about autovectorization. It is not even a true programming model, because we have only limited ability to influence results.
Also agree std::simd is far too limited (something like 50 ops, mostly the straightforward ones, vs >200 for Highway), and difficult to change/extend within the ISO process.
It is very difficult to get widespread traction with a new language, even given LLVM. Mojo, Carbon, Zig are also already potentially helpful.
I do believe a library approach (and in particular Highway, because considerable effort is required to maintain support for so many targets/compiler versions and AFAICS nothing else properly supports RISC-V and SVE) is the way to go for the next 5 years. Major compiler update cycles are something like 1.5-2 years and I don't think there will be a fundamental shift anytime soon towards RL, for example. After those 5 years, the future remains to be written :)
As to missing features: we are happy to add ops whenever there is sizable benefit for some app, and it doesn't hurt other targets. For mixed-type gather, x86 is the only platform that does this, so encouraging its use would pessimize other platforms. And I think apps can easily promote/demote their indices to match the data size. But always happy to discuss via Github issues :)
https://github.com/arduano/simdeez looks like it's trying to fit into this space, fairly promising.
Maybe modern tooling (e.g. Cargo) will lower the barrier so that it becomes less rare, but in C++ it's definitely not worth the effort for the vast vast vast majority of projects.
I admittedly focus more on libraries, but quite a large number of them are already vectorized. I would venture that a sizable fraction of CPU time, even in 'normal' non-HPC context, uses SIMD indirectly. Think image/video decompression, browser handshakes/rendering, image editing, etc.
Because it isn't just hot or autovectorisable loops that are faster in C++; everything is faster. Function calls, member accesses, arithmetic, etc. Even loops are normally not very hot and not autovectorisable.
You're right that things like audio/image processing, compression etc. benefits from SIMD but that is in the 1%. Those are libraries that have already been written. The vast vast majority of people are not writing audio codecs or whatever.
Written using SIMD.
> The vast vast majority of people are not writing audio codecs or whatever.
Or HPC, or finance, or json parsing, or PDE solving, or gaming....
All off these (and more) benefit from AVX512. Why are you going so far out of your way to be dismissive of this?