AMD EPYC 97x4 “Bergamo” CPUs: 128 Zen 4c CPU Cores for Servers, Shipping Now
anandtech.com
anandtech.com
Even if they are not branded Xeon, but Atom, there already are a lot of Intel server CPUs with E-cores, e.g. Denverton Refresh, Snow Ridge, Parker Ridge and the recently launched Arizona Beach.
Now imagine Intel had that product and canceled it in 2017, and you will be living in reality: https://en.wikipedia.org/wiki/Xeon_Phi
Not only would something like 1024 cores with AVX512 be competitive with GPUs it would have the added advantage of being waaay more versatile and easier to program.
And I'm not sure it'll be that much easier to program - "few large cores" and "GPU waves" are now pretty well supported in the stack with mature tooling - trying to insert a new stack between the two would likely be pretty difficult as it doesn't really fit well with either established paradigm so likely needs something new to show it's benefits.
The cores were in order. If you've never written code for in order cores, you may not appreciate exactly how heavily you're leaning on out of order to save your bacon. Things your used to just working suddenly don't.
The cores had a 2D layout, but Intel refused to tell you what it was and in practice it was impossible to optimize for. You ended up with variable (as in different on every node) latency for memory access that made programming again painful (and add this to the fact the cores are in order, so you really can't work around the latency).
Then there was the software stack that was just odd in some ways. I was researching I/O throughout, and things that just work on any other CPU (on any modern ISA) just didn't behave in a sane way.
So yeah, I was not at all sad when Intel cancelled it.
It was a big experiment, and it failed. Some of the brightest people have tried to push it to the limit. In my circle, I have never heard anyone missing Xeon Phi when it is discontinued.
A chip with hundreds or thousands of E-cores should be ideal for any cloud provider to offer vCPUs with a high markup. For example, a company like Hetzner sells 1vCPU services for around 4€/month. Even if they don't overprovision and allocate 1 core per 1vCPU, a single chip can potentially earn them 500€/month, which leads to a break even point at around 2 years.
Why would they put a dig at Intel’s consumer chips in a slide deck for their server parts? The Intel Xeons don’t have E-cores. This doesn’t make any sense unless I’m missing something.
Also, you know AMD has recent consumer chips without AVX-512, right?
Xeons with E-cores are planned: https://en.wikipedia.org/wiki/Sierra_Forest
And even if they weren't, it would still be a cheap PR win.
>Also, you know AMD has recent consumer chips without AVX-512, right?
All the current generation processors (Zen4) support AVX-512, both mobile and desktop. It may be confusing because AMD's numbering scheme is intentionally misleading and they sell previous gen chips with new model numbers.
Intel server CPUs with E-cores that will use the Xeon brand, because they will have much more cores (6-times more, i.e. 144 vs. 24) than the current models, are announced for 2024.
I've been messing with SIMD optimizations recently in a datastructure library I wrote, and as tantalizing as the AVX-512 is, it'll be years before it can be used in production software on a real scale. Introduced in 2017, and still totally unusable in the wild.
as a counterpoint, i'm on a P14s Gen 2 Thinkpad (Ryzen 7 PRO 5850u) and have zero issues on EndeavourOS (Arch, KDE/Plasma). most of my time is in VSCode, SublimeMerge, Chrome, MPV.
i did have to rip out the realtek (or mediatek?) M.2 wifi card and swap in an intel one tho.
Requires a hardware ("pinhole" not power cycle) reset every time I want to switch between external monitor use and laptop only. My bug report/forum post here: https://forums.lenovo.com/t5/ThinkPad-Z-series-Laptops/Exter... -- other people are effected.
Last year they pushed out a BIOS update which caused the fan to run 100% of the time. Then a "fix" which caused it to overheat. And in-between there somewhere something was pushed that caused me to have to reformat the whole thing.
Lid-open-to-awake in Linux used to simply not work 75% of the time, required a hard reset. Now it works, but there's a 5-10 second pause before display turns on (Windows or Linux) vs almost instant on the work issued laptops I've had recently (MBP M1 and some HP Z series thing, running Linux)
Waste of cash. Bought this thing to be my dev workstation when I took a contract job last summer, wish I'd never done that. Software quality / support at Lenovo is a real problem.
According to AMD, the AMD Ryzen 7 PRO 6850H does not have USB type-C support.
This is not my experience. Do you have some justification for this statement? I've found myself far more impressed by autovectorization than disappointed by it. I've found that code most people think is autovectorizable actually violates scalar contracts. But if you write clear code whose scalar implementation won't introduce UB if autovectorized, the compiler is really good.
Here's my go-to example. Agner Fog's VCL is a well-respected library. It has vectorized versions of all sorts of useful mathematitical functions. For fun, I rewrote his `exp(x)` function using his exact algorithm, but in scalar code, and it autovectorized, and benchmarks the same.
Maybe modern tooling (e.g. Cargo) will lower the barrier so that it becomes less rare, but in C++ it's definitely not worth the effort for the vast vast vast majority of projects.
I admittedly focus more on libraries, but quite a large number of them are already vectorized. I would venture that a sizable fraction of CPU time, even in 'normal' non-HPC context, uses SIMD indirectly. Think image/video decompression, browser handshakes/rendering, image editing, etc.
Because it isn't just hot or autovectorisable loops that are faster in C++; everything is faster. Function calls, member accesses, arithmetic, etc. Even loops are normally not very hot and not autovectorisable.
You're right that things like audio/image processing, compression etc. benefits from SIMD but that is in the 1%. Those are libraries that have already been written. The vast vast majority of people are not writing audio codecs or whatever.
Written using SIMD.
> The vast vast majority of people are not writing audio codecs or whatever.
Or HPC, or finance, or json parsing, or PDE solving, or gaming....
All off these (and more) benefit from AVX512. Why are you going so far out of your way to be dismissive of this?
And autovectorization is nice, but explicit SIMD intrinsics use usually wins if done right.
If I was getting paid $$ for this work, sure, I'd rent cloud instances or hardware to do that development. But it presents a dilemma for open source work.
Anyways, it's all griping. We'll either eventually all get AVX512, or it will die and some other more common wide vector extension will take its place, or we'll all be having the same gripe about NEON or RISC-V V extensions 10 years from now.
It's just frustrating some of the nice toys that are in AVX512, in particular nice support for bitmasking that would make the code I'm writing much nicer & faster.
> Anyways, it's all griping. We'll either eventually all get AVX512, or it will die and some other more common wide vector extension will take its place, or we'll all be having the same gripe about NEON or RISC-V V extensions 10 years from now.
I think we'll get a unified instruction set. Part of AVX512's difficulty is that it's not just AVX512 or not. It's AVX512F and/or AVX512BW and/or AVX512VL ...x10
Further to the Godbolt suggestion, one band-aid is that Highway fairly efficiently emulates some of the fancier AVX-512 instructions such as CompressStore. You could then develop on AVX2, then rent 1 VCPU-hour to build and verify it indeed works on AVX-512.
As for the subsets of AVX-512, we defined them into groups matching Skylake and Icelake; that works pretty well. Zen4 would also support the Icelake features, but it gets its own target so that we can special-case/avoid the microcoded and super-slow CompressStore there.
I've been using Highway for a couple weeks (I'm the guy who is writing the unroller feature). Highway is more limited in its breadth than raw intrinsics. I've noticed a few instances so far where you had to make a judgement call, and decided not to have Highway expose certain features (like 32 bit indexing into 64 bit type scatter/gather). And with more things that x86/ARM throw at us, the harder I think it becomes for a library to be the solution. From what I've seen of std::simd, I don't see how that possibly can be the solution. What do you think?
Agree about autovectorization. It is not even a true programming model, because we have only limited ability to influence results.
Also agree std::simd is far too limited (something like 50 ops, mostly the straightforward ones, vs >200 for Highway), and difficult to change/extend within the ISO process.
It is very difficult to get widespread traction with a new language, even given LLVM. Mojo, Carbon, Zig are also already potentially helpful.
I do believe a library approach (and in particular Highway, because considerable effort is required to maintain support for so many targets/compiler versions and AFAICS nothing else properly supports RISC-V and SVE) is the way to go for the next 5 years. Major compiler update cycles are something like 1.5-2 years and I don't think there will be a fundamental shift anytime soon towards RL, for example. After those 5 years, the future remains to be written :)
As to missing features: we are happy to add ops whenever there is sizable benefit for some app, and it doesn't hurt other targets. For mixed-type gather, x86 is the only platform that does this, so encouraging its use would pessimize other platforms. And I think apps can easily promote/demote their indices to match the data size. But always happy to discuss via Github issues :)
https://github.com/arduano/simdeez looks like it's trying to fit into this space, fairly promising.
could you tell more please? sorry, not in the loop on what a typical “scientific workflow” is like
[0] https://en.m.wikipedia.org/wiki/Message_Passing_Interface
A lot of power, but nothing really abnormal. Most rooms in a house won't be wired for it but that's about all. You probably wouldn't want this much thermal output in a normal room anyway, not to mention fan noise.
Honey, can you queue up a large simulation? Got a full load to dry again.
I don't know what things look like now, but I recall hearing many stories a decade ago about datacenters running out of power and cooling when the volume was closer to 1/3rd full.
https://semiwiki.com/forum/index.php?threads/tsmc-officially...
So, L3 cache didn't get smaller (in area or bytes), there's just less for each core. L1/L2 is relatively small, but they did use techniqued to make it smaller at the expense of performance.
I think the big difference really is the reductions in buffers, etc, needed to get a design that scales to the moon. This is likely a major factor for Apple's M series too. Apple computers are never getting the thermal design needed to clock to 5Ghz, so setting the design target much lower means a smaller core, better power efficiency, lower heat, etc. The same thing applies here: you're not running your dense servers at 5Ghz, there's just not enough capability to deliver power and remove heat; so a design with a more realistic target speed can be smaller and more efficient.
Teams is bad, yeah, but does not have an entire monopoly on being terribly inefficient.
> quad core
pick one
However, I wish to all 6+ core Intel laptop owners a very happy 65 decibel exhaust fan when opening your 3rd chrome tab.
So I think they've known how unpleasant Teams is!
https://www.microsoft.com/en-us/microsoft-365/blog/2023/03/2...
Take for example the whole business of screensharing, if you share a single screen in zoom you get black blocks for all the zoom windows. Yes it makes sense that I don't share the window with participants, but at least let me freaking close it. Similarly moving the controls from bottom to top of the screen reliably confuses early (and even more advanced users).
And don't get me started on quoting text or including math in slack and what is the whole threads section?!
Code blocks working in Teams is inconsistent, and it seems to add additional spaces when copying.
Quoting & mentions also sucks. The channels, chats, activity screens are always a mess, feel like UX bandaid.
Private channel limits are garbage.
Search and discovery sucks.
There's inconsistent options for the calendar between Outlook & Teams. Don't you dare expect functional compatibility between MS products. Scheduling assistant sucks, sometimes confuses itself when people are busy.
The UX seems to rely on a myriad of nested menus & modals that are XHR backed, so each one lags. Even just typing text feels slow. It's as fluid as molasses.
It regularly adds _fake spaces_ to code blocks. People paste working code/SQL into their things and all spaces were replaced with mysterious utf8 that are invisible like spaces.
Using something like Teams by comparison feels like mollases.
Put that on a napkin and give it to an unsuspecting intern.
In a lot of loads, it doesn't matter that much, because the application uses way more CPU than packet handling... But if it does matter, you really want to get things lined up as much as possible. Each CPU core gets one NIC queue pinned and one application thread pinned and the connections correctly mapped so the fast path never communicates to other cores. I'm not tuned into current NICs though, I don't know if you can get 128 queues on NICs now. If you have two sockets, then you also have the fun of Non-Uniform PCI-E Access...
When I was working on this (for a tcp mode HAProxy install), the NICs did 16 queues, and I had dual 12 or 14-core CPUs, so socket 0 got all its cores busy, and socket 1 just got a couple work threads and was mostly idle. Single socket, power of 2 cores is a lot better to optimize, but I had existing hardware to reuse.
In the subscriber only section it was mentioned there will be some Zen5 consumer parts using Zen5c as the equivalent of current Intel E-cores.
If there was nothing wrong with it OP wouldn't feel the need to use a throwaway account.
There is a lot more detail in the article on the possible 4c/5c uses, it's been speculated in a lot of contexts for the hybrid configurations (even by AMD directly) so don't think it's surprising to mention this at all.