AMD Rome – is it for real? Architecture and initial HPC performance
dell.com
dell.com
Because Intel paid Dell and others to not use AMD[0][1]. Dell officially ended their Intel exclusivity 13 years ago when they agreed to sell servers with Opterons[2].
[0] https://en.wikipedia.org/wiki/Advanced_Micro_Devices%2C_Inc.....
[1] https://business-ethics.com/2010/07/23/0901-dell-inc-agrees-...
Big discounts, kickbacks, "deals" and "programs" are being deployed and more on the way.
[0] https://s21.q4cdn.com/600692695/files/doc_financials/2018/An...
[1] http://ir.amd.com/static-files/438f4934-2883-4c85-9193-d5218...
'Big discounts, kickbacks, "deals" and "programs" are being deployed and more on the way.' are generally perfectly legal.
The practices that have resulted in Intel being sued/settling/fined by companies/states/FTC/EU/Japan.
HPE on the other hand tends to be a little more supportive.
And I could only guess there are lots of push from Intel to ask their customer not to offer any AMD machines.
No kidding ;)
"Initial performance studies on Rome based servers show expected performance"
Which is true, but misses the real point.
We're a VMware shop so I'd be hesitant to deploy non-Intel just because I don't want to get caught holding the bag in case AMD can't keep their momentum up for the next 5 years when we'd most likely looking to lifecycle the environment.
Was there any of this at play recently at AMD?
He moved to Intel during last transfer season
Apart from SiByte/Broadcom pretty much a greatest hits list of disruptive architectures over the last 20 years.
When Broadcom uses the term SOC, they mean "Switch on a Chip" to differentiate their single chip designs from older multi chip architectures.
SiByte was MIPS-based, more like an NPU than an (Broadcom term) SOC.
Obviously, people have different ideas about what they mean when they say "Moore's law is/is not dead", but I appreciate Jim's point that even if the current innovation curves we've been riding are slowing down, there are other innovation curves that we can take advantage of, and that the combination these curves mean that there's still plenty of innovation and improvement to be made in processor design.
Of course, the question of whether the x86 architecture will be the foundation for that future improvement remains to be seen. ARM and RISC-V are both eating x86 from below, and the fact that both of the processor architectures allow for greater competition among processor implementations suggests to me that one (or both) may catch up with x86 at some point in the future.
That Sunny Cove architecture slide shows a massive increase in the number of execution units.
It may be that Intel's next architecture, combined with memory speed improvements, is competitive with AMD.
But AMD also isn't sitting still. AMD's contributions (from what I can tell) include:
* Core counts have blown past 6 cores / 12 threads
* CPU prices have just been cut by half
* PCIe lanes have blown way past 8
* ECC ram is being offered in consumer PCs
https://www.statesman.com/business/20160904/amid-challenges-...
[1] At least that's how I think they work. One decode and FPU per pair of integer cores, right?
220 watts!!!
Some say Jim Keller's work at AMD lines up nicely with AMD's resurgence.
With no malice, I'm not aware if anything of significance came from his time at Tesla.
It will be interesting to see if AMD can keep things going past Zen 3 next year (which is likely to be the end of Keller's direct influence), and also what Intel can do with his guidance, which I would expect to release sometime in 2021-2023.
This dynamic is the reason competition is essential. Without external force compelling action, every organization trends toward stasis.
You can't build a new thing right from scratch, but you can market intermediate imperfect results to cover the losses somehow; I suppose that's what AMD did a few years ago.
Now they have finally hit the 10, and reap the benefits.
The key, as you say, was surviving an unpredictable amount of time till the tinkering paid off.
Intel also helped by botching their move to 10 nm.
All this coupled with really good execution from AMD, new architecture that has turned out to work really well and not really having to deal with Spectre/Meltdown regressions.
So yes AMD is doing very well, Intel's current problems makes them look even better. How much of this is hiring and how much is just trucking along? Who knows.
If you don't want to comment publicly and are still interested in helping me you can reach me at: gfair @ uncc.edu
Speculative execution is not really a "shortcut". Intel did make one shortcut but that resulted in Meltdown, which is easy to fix (conceptually anyway - I'm sure it was a lot of work!)
Any idea why I've been downvoted so much for saying something that surely everyone here remembers (wasn't that long ago) and is easily verified on Wikipedia:
https://en.m.wikipedia.org/wiki/Spectre_(security_vulnerabil...
> As of 2018, almost every computer system is affected by Spectre, including desktops, laptops, and mobile devices. Specifically, Spectre has been shown to work on Intel, AMD, ARM-based, and IBM processors.
The context of the discussion was cpuguy83's suggestion that compared to AMD, Intel has been "suffering from shortcuts leading to Spectre/Meltdown and the performance regressions due to patching those".
cpuguy83's comment was very concise, so let me elaborate that the distinction between Spectre and Meltdown is essential here. Both are security flaws in CPUs that were published at the same time. Spectre is an industry-wide problem that also hit AMD, but "Meltdown" is the result of an Intel-specific implementation choice that I think one might fairly describe as a "shortcut". (Even IshKebab later agrees: https://news.ycombinator.com/item?id=21342048) Meltdown was also more immediately dangerous, and the software workarounds that were necessary to mitigate it in existing CPUs cost a lot of performance.
Given this context, I'll quote the relevant item from cpuguy83 again and then IshKebab's reply:
> cpuguy83: Intel [... has been ...] suffering from shortcuts leading to Spectre/Meltdown and the performance regressions due to patching those.
> IshKebab: AMD chips suffer from Spectre too (which is the hard to fix issue). They didn't really take any fewer "shortcuts" than Intel. And they weren't "shortcuts".
Note how IshKebab carefully ignores the Meltdown part of cpuguy83's comment to be able to claim that Intel hasn't been doing any worse than AMD, and that there were no shortcuts. For Spectre in the stricter sense, that's technically true. It's not true in the context of the entire Spectre/Meltdown event, which was cpuguy83's argument.
Meltdown arguably does fit your description, but AFAIK the cost of checking permissions in the right place is almost zero, so it’s arguably better described as “we never realized it would be dangerous to not check permissions here” than “we skipped the check for performance’s sake”. (AMD processors were not vulnerable to Meltdown.)
Spectre, on the other hand, is sort of an inherent flaw of speculative execution (not related to permissions checks). Speculative execution itself is definitely a shortcut, but it’s a shortcut that’s crucial to the performance of all modern high-performance processors, with the result that nobody really knows how to deal with Spectre. Intel was apparently hit harder than AMD by side channel mitigations collectively, apparently because Intel was doing more aggressive speculation – but those mitigations are only partial. Both vendors’ processors are still vulnerable to Spectre attacks even with mitigations applied [1], and that will remain the case even on future processors, for the foreseeable future.
AMD essentially got lucky on this one - their neural-network based branch predictor is difficult for an attacker to train to follow specific code paths, which is a necessary component of Meltdown/Spectre style attacks. Pretty much every other processor that does speculation is affected.
The potential for cache timing to serve as a side-channel leak was not widely appreciated in the industry, although it was theoretically described as far back as the early 90s.
Yes, the response is technically correct. But...
I'm really not trying to defend Intel (I understand that it may come across that way), just how we look at the situation. Indeed you could be spot on, just that the situation at Intel is likely multi-faceted.
Who cares about management vs engineers, we are customers and want fine reasonably priced products. Amd delivers on this front much more than Intel. Rest are details.
AMD's success has nothing to do with Intel's failure.
AMD engineered a beast of a processor architecture that runs circles around the competition at a price that renders competing options obsolete. Ryzen's performance track record would not be any different even if Intel did not sold processors with security problems or succeeded shrinking die sizes.
It boggles the mind how some people try to spin AMD's outstanding progress as something that Intel did instead of a collosal technical achievement by AMD.
I said AMD is doing very well, and they look like they are doing even better because Intel is having problems.
I recall an anecdote from an AMD engineer (from some article) "we expected to be competing against Intel's next-gen architecture" (paraphrased).
e.g. https://www.wsj.com/articles/amd-to-license-chip-technology-...
Edit:
https://en.wikipedia.org/wiki/AMD%E2%80%93Chinese_joint_vent...
I don't know about µarch details but a bunch of bigger-picture things have contributed to AMD's run:
Spinning off their foundry ops led to AMD getting to use TSMC, who turned out to have a great process node at 7nm. Indirectly, it probably helps that other huge customers mean TSMC can amortize their process development costs across all of them.
The chiplet approach with a separate I/O die has various advantages:
- The Zen 2 chiplet is identical from the cheapest client CPU to the most expensive server CPU. Surely reduces complexity, and silicon that won't work in a server part might work in a client chip or whatever.
- Relatedly, binning 8-core chiplets is way more forgiving than binning huge monolithic CPUs like Intel's doing; with AMD lots of cores and high perf doesn't require a huge, uniformly near-perfect die, just enough chiplets that meet the spec that you can glue together.
- The I/O die is on a very mature, probably cheaper GloFo process, which may have made it less costly for AMD to offer interesting features on the I/O side (PCIe 4 and 128 lanes even for the cheapest server part, AES-128 in the memory controller, etc.).
Won't happen this gen, but I kind of expect Intel to eventually use some version of chiplets for their larger parts. They're talking about their advanced stacking/packaging tech, so not totally outlandish. Short-term they do have a two-die Cascade Lake mega-CPU planned, but I mean more broadly.
One cost of chiplets is in higher DRAM access latencies. AMD's done things to try and mitigate that, including just using lots of L3 cache. From benchmarks it seems to work.
The Intel analysis also mentions their shift to higher-margin parts as important.
I don't know how critical it is to the story but AMD also did some deals that may've given them resources to invest in their designs, like the GloFo spinoff, the deal to sell Epyc 1 clones in China, and the deals to produce Xbox and Playstation chips.
Finally, Intel would ordinarily have their own progress that would make AMD look comparatively worse. But 10nm stalled an incredible amount of time and Intel's post-Skylake core design, Sunny Cove, depended on it.
Additionally they were first to hypertransport (Intel followed with QPI) and knocked it out of the park with chiplets. Having two dies per package was pretty common, even way back at the 66 MHz pentium pro. But HT allows AMD quite a bit of flexibility. They can switch lanes between hypertransport (now updated and called Infinity Fabric) between pci-e and IF.
The this helps them on multiple fronts. It decouples the CPU and memory interface and pci-e standards. It raises their yield by using smaller chips. It also allows them to use different processes for different chips, so now the I/O die can use an older process.
One big impact is now only does the yield increase, but also the increased number of chips lets AMD amortize their R&D over more dies, and also customize for different produces without having to spend the extra R&D on numerous different dies. The low end desktop/laptops get 2 chips (1 cpu + 1 I/O). Higher end chips get 3 chips (2 cpu + 1 I/O). The high end servers get 9 (8 cpu and 1 I/O). So they can go from under 65 watts and under $200 to over 250 watts and over $5,000 all based on the same chiplets.
The AMD first generation Epyc did expose some weaknesses in OS/Applications that didn't like the high variations in latency to main memory and I/O. 4 chips inside a single socket had their own pair of memory channels and would have to use Hypertransport/IF to get to the other 6 memory channels. The design was reasonable, but many apps were NUMA aware and ran poorly.
In the second generation AMD moved all memory channels to the I/O chiplet and now the socket is a single NUMA domain and all chiplets see identical latency. The NUMA tweaks, 10-15% improvement in IPC, big improvement on the floating point side, and 1.5 x more cores (depending on the model) means that for many real world codes that AMD Epyc generation 2 chips (rome) are twice as fast at real world codes as the previous generation, which Intel is still trying to match. Meanwhile Intel is still trying to get a Xeon shrink they promised in 2017 working.
So generally ability to switch fabs and how well HT/IF works with chiplets caught Intel at a really bad time and for once AMD seems to be actually executing well and producing non-trivial volumes into multiple market segments.
So if you want to compare performance fairly I'd use gcc (or at least a non-intel compiler) and one of the MKL like libraries (ACML, gotoblas, openblas, etc). AMD has been directly contributing to various projects to optimize for AMD CPUs. They used to have their own compiler (that went from SGI -> cray -> pathscale or similar), but since then I believe have been contributing to GCC, LLVM, and various libraries.
If shopping I'd compare the highest end Ryzen + motherboard and the lowest end Epyc single socket chip and motherboard and try to guesstimate that price/performance for your workload.
Generally the Threadrippers seem like a much lower volume product and the motherboards are often quite expensive (for the current generation). Both Ryzen and Epyc enjoy significantly higher volumes.
Keep in mind that Threadripper has twice the memory bandwidth of Ryzen, but half the memory bandwidth of Epyc.
MI60 has 7.4 TFLOPs double precision at TDP=300W.
We are one step/generation away from running BGP IPv4 routing in PC CPU L3 cache."256MB L3 cache." I believe one needs 512MB L3 cache to fit the current routing tables in cache enabling very fast route lookups on generic PC hardware.
For forwarding, you do not need most attributes, but you may need better data structures for best-match lookups.
Yes. Though you have also the option of running a newer kernel (through EPEL but maybe not necessarily)
My best guess was that I didn't fully understand the architecture of the IO die. Eg, there was some benefit to being local that I didn't understand fully.
IIRC the next generation of Ryzen will introduce CCD or even I/O die caches. I think one of Intel's Broadwell chips had L4, so it's the same idea.
Can't wait to build a beast as my Valve Index VR rig.
*edit: shitty as in pixelated; not the content