Loading up on cheap cores with updated Milan Epycs
nextplatform.com
nextplatform.com
That's a big difference for 48 cores if you're paying for it yourself. The old ones are circa 2019, so not ancient by any means. Same net L3 cache as well.
So middle of the road, like most things for the general customer base.
Who's looking for the lowest performance per watt?
The article makes these processors look like specialty things.
Where they shone has always been highest single-thread performance per capex dollar (not necessarily per opex, which is highly influenced by power consumption).
When you add everything, the CPU is a relatively minor part of the server - a lot of energy is going to storage, memory, networking and other non-application activities, as well as memory and vector pipelines in GPUs (often a lot more than whatever goes into the CPU).
Being less burdened by compatibility, ARM is in a much better place WRT performance per watt. I expect RISC-V to be in a similar position soon.
I think this is a mistake from the author.
The slightly more recent Dell PowerEdge R720 can have CPUs with more cores and use a bit less electricity and noise but they are much more expensive where I live.
I've been on the lookout for a more modern platform with more cores per $100 but couldn't find any yet.
Ryzen 7000 CPUs are not the best only when the performance of the CPU is less important than having a lot of PCIe lanes and/or memory concentrated in a single box, instead of being distributed in multiple boxes (for the latter case multiple Ryzen 7000 boxes are cheaper than a single big server with the same total amount of memory and total count of PCIe lanes).
For a lot of memory and PCIe lanes in a single box, there are some Sapphire Rapids models with few cores that are quite cheap, but of course the Intel marketing has decided to cripple them by limiting their memory speed to DDR5-4400. The ubiquitous crippling of the cheap Intel SKUs is intended to force their customers to buy the SKUs with outrageous prices, but all that this has ever done to me was to force me to buy from their competition.
For the same purpose, AMD Siena CPUs are expected to be competitive, but only if they will become available at retail, which is uncertain.
Hetzner, OVH and Contabo are some of the only few I’m aware of who does.
The users certainly do not want such a speed reduction, especially taking into account that they will usually buy standard DDR5-4800 modules, which will never be used at the performance that they have been paid for.
It is very unlikely that the manufacturing cost reduction due to a relaxed specification for the memory controller is important. Cutting whatever features they could in the cheap SKUs, to annoy the customers (like disabling instruction subsets, and in Sapphire Rapids disabling the accelerators), has been the MO of Intel already for decades.
If anything I'd say that Epyc has support for 12 memory channels (win of the dedicated I/O die) is leaps and bounds more relevant. The Intel SKUs would need to run their 8 channels at 7200 MHz just to break even and here we are talking about 4400 MHz instead of 4800 MHz as a deciding reason to not buy an Intel server!
Every manufacturer has the MO of artificially cutting features to segment things. Intel, AMD, Nvidia, Apple - all of them. Now maybe the folks at Intel marketing were smoking something funny that day and genuinely though this limitation would actually create incentive to go to a higher tier, who knows. Even then, why would that particular (poor) segmentation be a driver be worth choosing to exclude a brand over when the meaningful segments from all players are all much larger?
Only Intel has always disabled features that have negligible impact on the manufacturing cost, but which force various expenses upon the customers, like the disabling of various instructions or of internal peripherals. Due to this Intel policy, there are still plenty of programs that run slower than they should on the latest CPUs, because they may have been built for instruction sets that have become obsolete more than a decade ago, like SSE.
The 10% memory throughput reduction is not negligible. There are 2 main scenarios for using a cheap Sapphire Rapids SKU, one would be to use it to host a very large number of SSDs, and the other would be to host some application that cannot be distributed over multiple servers and which wants a lot of memory. In the latter case, if the application has a performance dependent on the amount of memory, it is likely that it would also benefit from a faster memory, but with Sapphire Rapids if you want a 10% faster memory, which costs nothing extra in the price of the DIMMs, you may need to pay double or triple for the CPU.
10% memory bandwidth is actually extremely negligible, particularly for the lower core count SKUs in the tier. The desktop gear you mentioned this all for has 1/3 the memory bandwidth, 1/4 if you plan on putting lots of RAM in, and can get away with a great deal many more tasks than you're describing something with 10% less memory bandwidth as able to do. We're talking a maximum worst case difference, for going between the lowest end Xeon to highest end Xeon, of 10% for the "my workload is to benchmark memory bandwidth only" use case. The reasoning for why it's supposed to be bad seems more focused with the starting assumption anything Intel does must somehow be bad than actually explaining what actual workload impact is had vs other options, and that's the red flag for me here.
As much as I like AMD, I find this a significant advantage for a workstation / small server.
On the other hand, for any given performance at full load, the AMD CPUs have a much lower power consumption, as low as two thirds.
Which is preferable depends on the application. For a personal computer, idle power consumption may be the most important. For a server, the power consumption when used is normally more important.
I have several servers in my home, which are used intermittently. Their idle power consumption does not matter much, because whenever they are not used they are shut down. Wake-on-Ethernet is enabled on them, and whenever I use them Ethernet magic packets wake them immediately.
The Intel E-cores are quite good for mostly integer workloads, like code compilation, or for file servers or Web servers, so for such tasks the Intel CPUs are adequate (but ECC-supporting MBs for them are quite expensive).
On the other hand their floating-point performance is quite poor (and the P-cores are few and intentionally crippled), so for such workloads AMD Zen 4 is much better, especially for running non-legacy programs, which have been updated and optimized for the modern ISAs.
The wall plug idle power consumption that I have seen published for recent Intel-based desktops was usually around 30 to 35 W, about the same as of the older Kaby Lake or Coffee Lake based computers that I have.
The wall plug idle power consumption of Ryzen-based desktops, both as published and also valid for a couple of Ryzens that I use, is about 40 W to 50 W. (The idle Ryzen cores consume a negligible total power of less than 1 W, but the I/O die consumes 20 to 30 W, even when idle.)
The desktops that have GPU cards have higher idle power consumptions than these.
The (well-designed) small computers like Intel NUC or similar, which use mobile CPUs, have a wall plug idle power consumption of 5 to 10 W, regardless whether they use Intel Core or AMD Ryzen.
When the idle power consumption is a concern, the real solution is to use a small computer with a volume of less than 1 liter and a mobile CPU as the workstation (there are now models with Ryzen 9 7940HS or Intel Core i9-13900H, which are more powerful than most big desktops from just a few years ago), and to use big desktop computers only as servers that sleep when they are not used.
The multi-threaded performance of 14900K is good only for mostly-integer workloads (e.g. for software compilation it works fine), and even then only with the price of a much higher power consumption, e.g. by consuming 50 W to 100 W extra at equal performance.
For floating-point workloads or for other workloads that can exploit the vector instructions, 7950X can be much faster than 14900K.
If you need the I/O or full ECC of the bigger systems then you just need the bigger systems.
Unfortunately, there are still cloud providers who haven't rolled out mitigations either because their upstream vendors are slow or because they're afraid of the 15% performance hit.
If you can't run actual software faster than your competition, your theoretical benefits remain just that.
Obviously, this doesn't apply to every general purpose problem. Just noting there were some narrow cases where writing for Itanium produced something close to the promise.
I remember their Xeon Phi ones were very nice as well. I miss those machines.
No it can't. Itanium is a terrible design due to combining both static scheduling and pointless complexity. The worst of both worlds.
If you really insisted on static scheduling, you would implement an EDGE [0] architecture. The idea is quite simple. Programs consist of instruction blocks with a finite size. The CPU performs some very basic scheduling on the block level so the expectation is that the compiler optimizes one block at a time. Additionally, these blocks can reference other blocks and form a graph to make up a complete program. The processor then schedules the instruction blocks dynamically.
Tachyum built such a processor and it turns out they built it out of order anyway.
[0] https://en.wikipedia.org/wiki/Explicit_data_graph_execution
Isn't it "bring your friends"? I could be wrong (the last time this song played on the radio was in the 90s sometime).
Original: "Load up on guns, bring your friends ..."