AMD 3rd Gen EPYC Milan Review
anandtech.com
anandtech.com
https://www.anandtech.com/show/16548/interview-with-amd-forr...
> IC: While AMD increases the performance on its processor product line, the bandwidth out to DRAM remains constant. Is there an ideal intercept point where higher bandwidth memory makes sense for a customer?
> FN: I think you’re absolutely right, and really at the top of the stack, depending on the workload, that can be the performance limiter. If you’re comparing top of the stack parts in certain workloads, you’re not going to see as much of a performance gain from generation to generation, just because you are memory bandwidth limited at the end of the day.
> That’s going to continue as we keep increasing the performance of cores, and keep increasing the number of cores. But you should expect us to continue to increase the amount of bandwidth and memory support. DDR5 is coming, which has quite a bit of headroom to DDR4. We see more and more interest in using high bandwidth memory, for an on-package solution. I think you will see SKU’s in the future from a variety of companies incorporating HBM, especially for AI. That will initially be fairly specialized to be to be candid, because HBM is extremely expensive. So for most the standard DDR memory, even DDR5 memory, means that HBM is going to be confined initially to applications that are incredibly memory latency sensitive, and then you know, it’ll be interesting to how it plays out over time.
And I cant wait to see Netflix plays around with these in their FreeBSD box.
It is interesting the biggest upgrade from Zen 4 won't actually the Core part but the iOD. As I am expecting Zen 4 to be some tweaks and a Die Shrink to 5nm. TSMC also somehow unexpectedly announce doubling their 5nm capacity with new Fab capacity being built. My guess would be an aggregate demand from Apple and AMD exceed certain threshold to be worth doing.
I/O die sends signals to the rest of hardware. Wires are way longer, often with multiple connectors in them. I/O chiplet needs to produce and consume more of these milli-amps of electrical current.
One can produce larger transistors with finer processes e.g. by adding more fins to FinFETs, but since they need larger transistors anyway to handle more current, I’m not sure there’s much value in upgrading process of the I/O chiplet.
> What’s really interesting is that inter-socket latencies have also seen very notable reductions. Whereas the Rome part these went up to 272ns, the new Milan part reduces this down to 202ns, again a large 25% improvement between generations.
Good I/O is going to have a cost in terms of die area and power. Granted what Anandtech observed is interesting for 2P systems. Maybe there's a NPS4 vs NPS1 difference at work here, or AMD just isn't doing as much power gating on the IOD as it could in one or either case.
I'll take a 7702P at $2-3K over a 7713P at $5K ten times out of ten.
I think you're right: that buying a generation old or so can offer gross cost savings. But that's only true for the time-period where those chips are available.
That may seem odd, but a lot of safety-critical applications (e.g. medical, military, aerospace, etc.) require spending tens of thousands, hundreds of thousands of dollars, or even millions of dollars (not to mention months of time) re-validating a system with any substantive change.
Even for less critical applications, spending $2000 extra on each CPU is a bargain compared to re-validating a system.
If AMD wants to be a credible presence in those markets, and I'm pretty sure it does, it needs to chips with many year lifespans before EOL.
Some companies manage this by having a subset of devices or of software which is LTS.
In contrast: artwork, houses, and a few other goods (Magic: The Gathering cards?) seem to appreciate... or increase in price. These stuff you want to buy before they get more expensive.
You need this exact model of X, Y, and Z, because those have been validated as a working combination through a huge one-time capital expenditure; and changing anything would entail going through that again.
Vendors tend to not mark down prices for components, but rather just sell them for MSRP for as long as they can, and then immediately stop producing the item. The margins on each component were usually thin enough that there's no way to sell the component any cheaper but "make it up in volume." It's either "sell it at the MSRP" or "don't produce it at all."
As such, the only place to get the component after the original vendor ceases production, is the secondary market, e.g. new old stock. This is where you see governments buying old 5.25" floppy disks for huge markups.
In that case, like I said, why not just buy all the 5.25" floppies you'll ever need in advance? They know they'll be using them for 50 years, because that's the decomissioning schedule for the system they just built, and they'll be unlikely to ever get budget to replace/upgrade it before then. They also know how long a floppy lasts in use. So, do the math, buy that many floppies early, and toss them all in an airless vault. Right?
And the same thing would be true here with particular-model CPUs, no?
Old process nodes (such as 130nm process, which is 100x larger than today's 12nm or 7nm processes) are still useful today. They're much cheaper to operate, because all the equipment is old and better understood (while 7nm and even 12nm processes are highly proprietary). Case in point: the Perseverance Rover that just landed on Mars is literally built from the ancient 130nm process. (POWER architecture)
Instead of ordering all the chips you need, you instead want to just ensure that the 130nm (or 45nm process, or 22nm process or whatever) will remain useful for 50 years into the future.
Lets say you base a modern government project (a spaceship, a jet fighter, a naval super carrier, etc. etc) off of a modern 14nm process node today.
1. You can order all the chips you'd "ever need" today, which is expensive today because tons of people are still using 14nm nodes.
2. Alternatively, you can order the company to continue to support the 14nm node for the next 50 years. This seems like the superior option. You order the chips you need later, when you need them. And 10 years from now, people will be on 3nm or maybe even 2nm process, so it'd be cheaper to order 2020-era 14nm chips at that point.
You might not know how many you need. Suppose one of your data centers burns down and has to be replaced or you have triple the expected customer demand.
Also, time value of money.
https://ir.amd.com/news-events/press-releases/detail/993/amd...
256MB L3 (or really, 8 x 32MBs L3) and 24-cores suggests the bottom-of-the-barrel 3 cores active per 8-core CCX.
8x CCX with 3-cores. The yields on those chips must be outstanding: its like 62.5% of the cores could have critical errors and they can still sell it at that price.
EDIT: My numbers were wrong at first. Fixed. Zen3 is double-sized CCX (32MBs / CCX instead of 16MBs/CCX)
---------
In contrast, the 28-core 7453 is $1,570. I personally would probably go with the 28-core (with only 2x32MB L3 cache, or 64MBs) rather than the 24-core with 256MBs L3 cache.
For my applications, I bet that having 7-cores share an L3 cache (and therefore able to communicate quickly) is better than having 1 or 2 cores having 32MBs of L3 to themselves.
There are also significant price savings, as well as significant power / wattage savings with the 28-core / 64MBs model.
Which is cheaper than the 24c 7443 and 7413 but not the 16c 7343 and 7313.
And it only has half the L3 compared to its siblings (1/4th compared to the 7543 top end), a lower turbo than every other processor in the range (whether lower or higher core counts), well as an unimpressive base frequency, and a fairly high TDP by comparison (as high as the 7543).
The 74F3 has no such discrepancy, it has the same L3 as every other F-series and slots neatly into the range frequency-wise: same turbo as its siblings (with the 72 being 100MHz higher), 300MHz base lower han the 73, and 250 higher than the 75.
28-cores for $1570 seems to be the "cheapest per core" in the entire lineup.
It all comes down to whether you want those cores actually communicating over L3 cache, or not. Do you want 7-cores per L3 cache, or do you prefer 4-cores per L3 cache?
4-cores per L3 cache benefits from having more overall cache per core. But more-cores per L3 cache means that more of your threads can tightly-communicate cheaply, and effectively.
---------
More L3 cache probably benefits from cloud-deployments, Virtual Desktops, and similar (since those cores aren't communicating as much).
More cores per L3 cache benefits from more tightly integrated multicore applications.
EDIT: Also note that "more cores" means more L1 and L2 cache, which is arguably more important in compute-heavy situations. L3 cache size is great of course, but many applications are L1 / L2 constrained and will prefer more cores instead. 24c 7443 with 2x32MB L3 is probably a better chess-engine than 16c 7343 4x32MB L3.
People who don't save a boatload by getting the license-optimized CPU will invariably choose to buy the 24-core one, which helps AMD by making it easier for them to keep up with the demand for the 16-core variant, and the 16-core variant gets an unusually nice profit margin. Win win.
This is not the first time AMD or Intel have offered a weird inverse-pricing jump like this... I highly doubt it is a typo.
My other comment reiterates some of these points a different way: https://news.ycombinator.com/item?id=26469182
75F3 32-core 2.95GHz $4,860
7543 32-core 2.80GHz $3,761
7513 32-core 2.60GHz $2,840
74F3 24-core 3.20GHz $2,900
7443 24-core 2.85GHz $2,010
7413 24-core 2.65FHz $1,825
73F3 16-core 3.50GHz $3,521
7343 16-core 3.20GHz $1,565
7313 16-core 3.00GHz $1,083
Source: https://ir.amd.com/news-events/press-releases/detail/993/amd...“ Users will notice that the 16-core processor is more expensive ($3521) than the 24 core processor ($2900) here. This was the same in the previous generation, however in that case the 16-core had the higher TDP. For this launch, both the 16-core F and 24-core F have the same TDP, so the only reason I can think of for AMD to have a higher price on the 16-core processor is that it only has 2 cores per chiplet active, rather than three? Perhaps it is easier to bin a processor with an even number of cores active”
As I said below the article in the comments:
> If I were to speculate, I would strongly guess that the actual reason is licensing. AMD knows that more people are going to want the 16 core CPUs in order to fit into certain brackets of software licensing, so AMD charges more for those to maximize profit and availability of the 16 core parts. For those customers, moving to a 24 core processor would probably mean paying significantly more for whatever software they're licensing.
This is the more compelling reason to me, and it matches with server processors that Intel and AMD have charged more for in the past.
"Even vs odd" affecting the difficulty of the binning process just sounds extremely arbitrary... definitely not likely to affect customer prices, given how many other products are in AMD's stack that don't show this same inverse pricing discrepancy.
Case in point: AMD Zen2 has 2048 TLB entries (L2), under a default (in Linux and Windows) of 4kB per TLB entry. That's 8MBs of TLB before your processor starts to page-walk.
Emphasis: Your application will pagewalk when the data still fits in L3 cache.
------------
I'm looking at some of this lineup with 3-cores per CCX (32MBs L3 cache), which means under default 4kB pages, those cores will always require a pagewalk to just read/write its 32MBs L3 cache effectively.
With that being said: 2048 TLB entries for Zen2 processors. Maybe AMD has increased the TLB entries for Zen3. Either way, you probably should start looking at hugepage configuration settings...
These L3 cache sizes are absurd, to the point where its kind of unwieldy. I mean, with enough configuration / programming, you can really make these things fly. But its not exactly plug-and-play.
So having 2MB (or even 1GB) hugepages is a big advantage in memory-heavy applications, like databases. No, 1GB pages won't fit in L3 cache, but it still means you won't have to page-walk when looking for memory.
1GB pages might be too big for today's computers, but 2MB pages might be good enough for default now. Historically, 4kB was needed for swap purposes (going to 2MB with Swap would incur too much latency if data paged out), but with 32GBs RAM + SSDs on today's computers... fewer and fewer people seem to need swap.
There might be some kind of fragmentation-benefit for using the smaller pages, but it really is a hassle for your CPU's TLB to try to keep track of all that virtual memory and put it back in order.
---------
While there is performance hits associated with page-walks, the page-walk process is fortunately pretty fast. So most applications probably won't notice a major speedup... still though, the idea of tons of unnecessary page-walks slowing down untold amounts of code bothers me a bit for some reason.
Note: ARM also supports hugepages. So going up to 2MBs (or bigger) on ARM is also possible.
But Phoronix has synthetic benchmarks from systems quite a few people use (e.g. databases, rendering, simulation). SPEC is nice, but I wouldn't know where to look if I would want to apply that to e.g. postgres. With Phoronix' results, at least there is an idea of how I could apply the results of the benchmark to my usage.
So I find this link quite useful.
That's not to say that I think Anandtechs benchmarks are not valuable (I love reading the nitty gritty details like cache timing specifics), but I can't apply them to anything other than a generic CPU workload, not even slightly in the direction of my specific workload.
So one can maybe argue that those workloads don't match your use case. But... SPECINT is the "standard server benchmark" for a reason.
I think there's been some suggestions that SPEC benchmarks are a bit small these days and maybe bigger code would be more realistic. But the actual programs they use are decent.
That I didn't know.
SPEC doesn't really do much in regards to giving descriptive names to benchmarked items, so its hard to determine what the numbers indicate (other than "this benchmark, which emulates some workload, now has X higher performance"). Some names are descriptive-ish (I think I recognise 7 out of 22 names) but it's all gibberish without a link to why those benchmarks are run and what they are based on (at least not in a way similar to how SPECjbb is explained). It might just be me failing to find it, but searching SPEC in the page (or their website) doesn't seem to give me relevant links.
Does Anandtech have a page detailing the rationale behind their current (CPU) benchmark suite?
Among other benchmarks. But Perl is the first one in their suite.
EDIT: SPECjbb is the Java benchmark suite: https://www.spec.org/jbb2015/docs/designdocument.pdf
SPECjbb is a ~2 hour benchmark simulating a high-throughput Supermarket backend.
* Computex 2017 is when Zen Threadripper was announced.
* Computex 2019 is when Zen 2 Threadripper was announced.
* Computex 2020 is when Zen 2 Threadripper PRO was going to be announced. After the show was delayed, and then officially canceled, AMD did their own announcement shortly after.
So odds are it'll be Computex 2021 on June 1st. This fits in with the usual release timing, as Threadripper tends to follow EPYC by a few months. I'd expect retail channel availability for Threadripper as early as late July, as there's usually a few months between announce and channel availability.
Trust me, I've been waiting to build a 5970X as well. :)
INVLPGB New instruction to use instead of inter-core
unterrupts to broadcast page invalidates, requires
OS/hypervisor support
VAES / VPCLMULQDQ AVX2 Instructions for
encryption/decryption acceleration
SEV-ES Limits the interruptions a malicious hypervisor may
inject into a VM/instance
Memory Protection Keys Application control for access-
disable and write-disable settings without TLB management
Process Context ID (PCID) Process tags in TLB to reduce
flush requirements
Interruptions (Instructions) and Unterrupts (Interrupts) aside (the article obviously was pushed out as fast as AT could lol) - these additions seem like they would help with performance when it comes to mitigations of all the speculation vulnerabilities in an hypervisor env?[1] I may be wrong here but the TR looks to me like an EPYC chip with the mult-CPU stuff all pulled off. It would be interesting to have a decap with the chiplets identified.
(... really need to replace my 2×E5-2630v2, it's getting quite old... but I guess a 2×7282 will do.)
[Ed.: nevermind, hadn't read to the "idle consumption" graphs yet... nope, I'll stick with Rome ...]
But, maybe their home server's singular purpose is to have a lot of RAM?
For a 7282 box I'd spend about as much on the RAM as on the CPU again; 16GB DDR4-3200R sticks are about 80€ each, so filling one socket is 640€. The 7282 currently sells for ~700€ currently so that lines up all right.
(This would give me 2×128GB total; 32GB RAM sticks are unfortunately indeed still a bit too expensive, especially at 3200MHz rating. But maybe I should go with 32GB and half fill each socket & fill it out later when it's cheaper...)
Two 7282s is pretty much the best perf/power/money option as far as I can see.
Also since Xeon dies are monolithic, unlike AMDs chiplet design, means that increasing the size of certain components on the die, like cache for example, increases the risk of defects wich reduces the yields, making them unprofitable.
I wonder if we could get numbers for L3 misses and cycles spent waiting for main memory under realistic workloads.
I suggest learning to read performance counters, so that you can get information like this yourself! L3 cache is a bit difficult for AMD processors (because many cores share the L3 cache), but L2 cache is pretty easy to work with and profile.
General memory-reads / memory latency is pretty easy to read with various performance counters. Given the amount of latency, you can sorta guess if its in L3 or in DDR4.
Thus, 'ye olde' would be pronounced 'the old'. 'Ye' is just a different form of 'the' from a time when 'y' represented 'th'.
So, 'the ye olde' would be read as 'the the olde'.
As such, AMD can make absolutely huge amounts of L3 cache (well, many parallel L3 clusters), while other CPU designers need to figure out how to combine the L3 so that a single core can benefit it from it all.
The first generation of Epyc had a complicated hierarchy that made latency quite hard to predict, but the new architecture is simpler. A CPU can talk to a cache in the same package but on a different die with reasonably low latency.
(I don't have numbers. Still reading.)
Think of the MESI messages that must happen before you can talk to a remote L3 cache:
1. Core#0 tries to talk to L3 cache associated with Core#17.
2. Core#17 has to evict data from L1 and L2, ensuring that its L3 cache is in fact up to date. During this time, Core#0 is stalled (or working on its hyperthread instead).
3. Once done, then Core#17's L3 cache can send the data to Core#0's L3 cache.
----------
In contrast, step#2 doesn't happen with raw DDR4 (no core owns the data).
This fact doesn't change with the new "star" architecture of Zen2 or Zen3. The I/O die just makes it a bit more efficient. I'd still expect remote L3 communications to be as slow, or slower, than DDR4.
As a comparison, an IBM z15 mainframe CPU has 10 cores and 256MB per socket.
Well, that's eDRAM magic, isn't it? Most manufacturers are unable to make eDRAM on a CPU.
> My question is why they chose to have just 32MB for up to 80 cores when AMD can choose to have 32MB per 8-core chiplet.
From my understanding, those ARM chips are largely I/O devices: read from disk -> output to Ethernet.
In contrast, IBM's are known for database backends, which likely benefits from gross amounts of L3 cache. EPYC is general purpose: you might run a database on it, you might run I/O constrained apps on it. So kind of a middle ground.
> those ARM chips are largely I/O devices
Anything Neoverse-N1 is pretty good at general compute and databases. There's definitely a lot of Postgres running on AWS Graviton2 instances already :)
That said, there are two major reasons:
1. Epyc is on a chiplet architecture. Large chips are harder to make than small ones. Building two 200mm^2 chips is cheaper than building one 400mm^2 chip. Since Epyc has a chiplet architecture, this means they can put more silicon into a chip for the same price. This means that Epyc can be just plain bigger than the competition. This comes with some complexity and inefficiency but has, so far, paid off in spades for them.
2. Epyc is on a newer process. This means AMD can fit more transistors even with the same area. Intel has had serious problems with their newer processes, so this is not an advantage AMD expected to have when designing the part. The use of a cutting-edge process was, in part, enabled by the chiplet architecture. It is possible to fabricate several small chips on a 7nm process even though one large chip would be prohibitively expensive, and AMD has been able to use a 14nm process in parts of the CPU that wouldn't benefit from a 7nm process to cut costs.
The first point is serious cleverness on the part of AMD. The second point is mostly that Intel dropped the ball.
[1] https://www.anandtech.com/show/16021/intel-moving-to-chiplet...
[2] https://www.anandtech.com/show/16051/3dfabric-the-home-for-t...
The mobile market cares more about efficiency than easily scaling up to much bigger chips, so the M1 and other ARM chips are probably going to ignore this without much consequence for smaller chips.
Intel still tops sales because of non-perf related reasons like refresh cycles, distrust of AMD from last time they fell apart in the server space, producing chips in sufficient quantities unlike the entire rest of the industry fighting over TSMC's capacity, etc.
Data got weird.
I suspect I still have the record for most expletives used in front of the headmaster at my old school because someone turned a few kW speaker on while I was wiring something under the stage - i.e. I wasn't pleased.
My 5900X at idle draws around 22W. Modern Intel desktop CPUs (which are basically all Skylake refreshes right now) are 8-10W at idle.
Lots of airflow.
High-density systems are usually 2U, but with four dual-socket systems per chassis.
Overall heat density on chips has increased due to litography changes so the chiplet architecture in a way is only a stopgap.
Intel or AMD or SPARC or ARM the setup is all the same for rackmount hardwares, pure copper passive heatsink and high power axial fans
Really powerful desktop and server-grade processors are available with fully open source firmware and hardware (even fully open BMC). It's not that expensive, comes with PCIe 4.0 and other modern features [1].
They are POWER9 processors, not x86_64. But many software work on it, checkout 'talos-workstation' on IRC to chat with the users [can access IRC through Element Matrix app] - someone is running Android Studio on it (yeah so much for privacy but still).
These systems are currently being used by businesses and developers alike, including Google.
The desktop-grade Power9 gets rekt by a Core i3: https://openbenchmarking.org/vs/Processor/POWER9%204-Core,In...
The advantage of POWER9 is fully open firmware and hardware while being really powerful. They won't beat the latest AMD in performance. But AMD can never be open. The RaptorCS products are FSF RYF (respects your freedom) certified.
https://ryf.fsf.org/vendors/raptor
Some benchmarks from 2019: https://www.phoronix.com/scan.php?page=article&item=rome-pow...