Sizing Up Servers: Intel's Skylake-SP Xeon versus AMD's EPYC 7000
anandtech.com
anandtech.com
All in all, whhen you have a server that seems so close in performance to Intel for less money and consuming less power, I can't imagine that EPYC won't see broad adoption and Intel won't be squeezed.
I'm glad AMD is back and there is renewed competition in the server market!
Also note that Intel is refining their pesky market segmentation game: the full 2x 512-bit FMA/cycle is only available on high-end CPUs, only lower-end it's only 1 FMA/cycle and you still get the extra AVX512 clock throttle hit [2]!
I really like the fact that AMD stayed away from such devil-in-the-details feature-based market segmentation!
[1] https://software.intel.com/en-us/blogs/2013/avx-512-instruct... [2] https://www.servethehome.com/wp-content/uploads/2017/07/Inte...
Socket compatibility would be wise, I hope at least the Zen+
I'm hoping they soon get back to competing aggressively at high-end HPC too, perhaps they could spend some area on Zen+ on better FP SIMD, or even flexible width SIMD (though the latter would be more realistic on the 7nm successor).
1. The mesh interconnect looks like a big loser for the smaller parts. It's a big jump up in complexity (there's an academic paper floating around which describes the guts of an early-stage version) and seems to be a power and performance drain. I can't imagine they got the clock speeds they wanted out of it. Sure, it's probably necessary for the high-core-count SKUs, but the ring bus probably would have done a lot better for the smaller ones.
2. There's almost nothing in here for high-end workstations (which typically have launched with the server parts). Sure, AMD has Threadripper coming soon, but this looks like Intel's full lineup... so where are the parts? We've bought plenty of Xeon E5-1650s and 1660s around here, and it doesn't look like there's anything here to replace them. That's unexpected. The "Gold 5122" (ugh what a silly name) is comparable, but at $1221 is priced just about double what an E5-1650v4 runs.
Workstations are a bit of an interesting case because their loads look a lot more like a "gaming desktop" than a server: a few cores loaded most of the time with occasional bursts of high-thread-count loads. That typically favors big caches, fewer cores, and aggressive clock boosting. If you're only running max thread count every now and then, you can afford a huge frequency hit when you do. But since these are business systems we try to avoid anything that doesn't say "Xeon" on it (or "Opteron", in years past) as reliability is paramount. To see nothing here from Intel in this launch is discouraging, to say the least. I have an upgrade budget and it looks like it'll be heading nVidia's way at this point.
Indeed surprising given that they had a few years to learn from the KNL experience.
Got a link to that paper?
> I have an upgrade budget and it looks like it'll be heading nVidia's way at this point.
Just for curiosity: what workload allows you to go with NVIDIA instead of Intel (rather than AMD)?
> Just for curiosity: what workload allows you to go with NVIDIA instead of Intel (rather than AMD)?
The particular workload I'm thinking of is Abaqus FEA. Per-core licensing is a part of that, but it's starved for double-precision (FP64) FLOPS and thus runs great on an appropriate Tesla. We have a number of K40s (which last I looked was the best workstation FP64 Tesla) for that exact reason. And they cost less than a single "Gold 6154" Xeon that the sibling comment (rightly) singles out for good single-thread performance. Throw in Abaqus's horrid per-CPU-core licensing and it's a no-brainer. Intel is just charging too much money for anything launched today to be viable in a workstation.
[1]: http://www.anandtech.com/show/11550/the-intel-skylakex-revie...
[2]: MoDe-X: Microarchitecture of a Layout-Aware Modular Decoupled Crossbar for On-Chip Interconnects, DOI: 10.1109/TC.2012.203
Ugh, per-core licensing is pita. Tell the decision-maker at your work to pay for the cores or switch to some code with more sensible pricing model. Do the alternatives (like Ansys) charge per core as well?
Wasn't that launched last month with LGA 2066, http://www.anandtech.com/show/11550/the-intel-skylakex-revie...? Sure, those do not wear the Xeon name, but that platform has cpus that are comparable to Xeon E5-1650 and 1660s. And there are additional cpus with higher core count announced.
[1]: http://www.tomshardware.com/reviews/-intel-skylake-x-overclo...
> This pricing seems crazy, but it is worth pointing out a couple of things. The companies that buy these parts, namely the big HPC clients, do not pay these prices.
We in HPC would not touch these outside big memory systems which is even niche for us. The consumers of these are far more likely to be those with data warehouse style needs (a.k.a Oracle customers).
Much like the rest of the world 2 socket systems in HPC are by far the most common.
<disclaimer, I work for DO>
Looks pretty solid. Sure not everything scales linearly with corecount but if your task does, it looks like AMD might be worth considering.
"What does this mean to the end user? The 64 MB L3 on the spec sheet does not really exist. In fact even the 16 MB L3 on a single Zeppelin die consists of two 8 MB L3-caches. There is no cache that truly functions as single, unified L3-cache on the MCM; instead there are eight separate 8 MB L3-caches."
Also:
"AMD's unloaded latency is very competitive under 8 MB, and is a vast improvement over previous AMD server CPUs. Unfortunately, accessing more 8 MB incurs worse latency than a Broadwell core accessing DRAM. Due to the slow L3-cache access, AMD's DRAM access is also the slowest. The importance of unloaded DRAM latency should of course not be exaggerated: in most applications most of the loads are done in the caches. Still, it is bad news for applications with pointer chasing or other latency-sensitive operations."
I was kind of expecting this, but it's still disappointing to see. Looks like if you need a lot of L3, Intel is still the best/only option. Not to say that AMD hasn't made massive improvements though - and it's also worth noting that while AMD's memory latency is generally worse, throughput is also typically better than Intel.
So a dual socket AMD has 8 zeppelin chips and 16 8MB L3 caches. I'd be quite surprised if intel could match the bandwidth of those 16 L3 caches. Additionally if there is enough cache misses AMD has a 33% advantage in both outstanding memory references (16 at a time in a dual socket) and bandwidth.
Basically both architectures are HUGELY complicated. Even minor things like which compiler/which compiler flags can make a big difference. Now more than ever it's important to benchmark your workload, any simple rule of thumb is likely to be useless.
https://software.intel.com/en-us/articles/introduction-to-ca...
This strikes me as either a bug or a benchmarking glitch, though other benchmarks seem to imply that the situation is real.
Assuming it's legit, this gives AMD a great opportunity for a boost in their first respin of the 8C/16T die.
How can other clients get in on this? My CS department (UK university) is wanting to replace servers (and extend ML capacity) and I've been making them wait until AMD availability becomes clearer (but I can only do that for so long...).
There's definitely a market here - if anyone from AMD happens to be reading this and would like to demo... though I'm not sure our volume would be quite the same scale as the above.
Wondering other than some sparse matrix applications known to be memory bandwidth bound, what kind of performance impact this is going to cause. Is there any real memory bandwidth bound applications other than ML/AI stuff used by those Internet big names?
More and more applications are drifting into the memory-bound regime, especially with the wider SIMD instruction sets increasing arithmetic throughput while memory throughput lags behind.
My back-of-the-envelope calculation (with a guesstimated AVX512 clock) gives a 12 FLOPS/byte for a big Skylake chip like the 8176 while this was around 9 FLOPS/byte for Broadwell. I'm not entirely sure about the instruction throughput of Zen, but it looks like the 7601 should be around 4-5 FLOPS/byte (that's worst-case with mixed FMA+ADD workload based checked Agner F's manual [1] IIUC).
Of course this does not consider NUMA and other effects, but given the above a lot of applications will benefit from the great bandwidth advantage of EPYC.
AFAIK, GitLab's server proposal included them. They will probably not use 128GB TSV LR-DIMMs immediately though. I think the price gap between 32GB RDIMMs and 64GB LR-DIMMs are falling right now right?