AMD's Future in Servers: New 7000-Series CPUs Launched and EPYC Analysis
anandtech.com
anandtech.com
So AMD pours everything into a single die and manages to hit the desktop (Ryzen with 1 die), workstation (thread ripper with 2 dies), and server (Epyc with 4 dies). All with a single wafer of silicon, and as a bonus the memory bandwidth, max memory capacity, and total number of cores scales to fit all 3 markets.
Pretty crafty. This is far from new of course, various Power CPUs, Intel cpus as far back as the pentium pro, and of course the previous generation Xeons.
Intel's strategy has been one die for most of their low end/low power chips (max 2 core/4T) and a larger die for their desktop (max 4c/8t) that use the same socket (LGA1151).
Then Intel targets HEDT (High end desktop) and server with the same chip, same die, same socket, just marketing to differentiate the x-servies chips and the regular single/dual socket chips that share chipsets, sockets, and a lga2011 socket.
AMD seems to have scared Intel pretty bad. They have held off on the skylake xeons, only shipping them to cloud providers while they wait for the AMD release before releasing skylake xeons to the masses.
I'm glad to say AMD performance seems pretty good, SpecINT (using GCC) is around 50% faster than the similar Intel chip and SpecFP is even better. Seems fair unless you use the Intel compiler to compile all your binaries anyways. In fact the fastest AMD + gcc-6.2 is faster at SpecFP than the fastest Intel + Intel's compiler (1330 vs 1090 respectively).
You can see some info about them here, http://www.anandtech.com/show/11550/the-intel-skylakex-revie...
(Bottom table)
I imagine the intra-processor latency between dies must be better than the latency between sockets.
So I think EPYC is a pretty compelling solution for the bulk of the market.
I think AMD reusing the same dies is quite elegant (everything* scales as you add more cores). The inter-die bandwidth/ latency shouldn't really be a problem, because if more than 8 cores are dependant on the same data - locking issues would make any additional cores useless anyway. If they can get NUMA tuned correctly things should work nicely.
*it's strange that they all have the same amount of L3 cache.
Does anybody have contacts with AMD's marketing/ engineering department who can provide test units?
10 years later, Intel sticks to single-die chips, and AMD has no qualms stuffing a gazillion dies on a chip. Looks like they've learned a lesson or two :)
From my understanding though chip A in socket 1 is directly connected to chip A in socket 2. So the worst case would be chip A in socket 1 to chip B, C, or D in socket 2. I don't think that number is public yet.
* Epyc uses socket SP3 https://en.wikipedia.org/wiki/Socket_SP3
* Threadripper uses socket TR4 https://en.wikipedia.org/wiki/Socket_TR4
* Sockets SP3 and TR4 have the same number of pins (4094 pins) and they have the same cooler bracket mount (see https://www.overclock3d.net/news/cases_cooling/noctua_showca... )
* However they are still two separate sockets so you shouldn't expect to be able to use Epyc on TR4 or Threadripper on SP3
I guess it would be hard as there are to many ways to scale out what you run - how many VMs, how many containers, what are you running in them? It would be an interesting benchmark matrix to sort for.
It would be interesting just to see how many containers you could start, run lighttpd and each server a static web page? Maybe 1/2 with the page and 1/2 with an application that builds the page? Who knows...to many variables.
I think we will just by a system when we can and try our workload on it. Oh, well.
VMs - now with processor virtualisation technology I'm sure the different processor architectures do make an interesting difference there.
I am familiar with Calico and it would not have solved it.
$6k vs $500k... Let that sink in :)
https://www.phoronix.com/scan.php?page=news_item&px=Ryzen-Co...
Last response from AMD: "The vast majority of users using Ryzen for Linux code and development have reported very positive results. ... A small number of users have reported some isolated issues and conflicting observations. AMD is working with users individually to understand and resolve these issues."
I still recommend it to people who are serious about diving into Linux.
These are awesome times we live in.
I recall, back in the long long ago, running through `make menuconfig` and disabling what I didn't need to get a smaller kernel + shorter build-times.
https://www.bloomberg.com/news/articles/2017-06-20/amd-serve...
I said exactly that, "things haven't changed much since then". It ran well and all, and never overtly irritated me.
But the newest round of affordable mid-high end CPUs is a huge upgrade. Same with SSDs, even compared to ones a few years old. It's one of those things you don't really realize how big the difference is until you use them back to back.
Do it!
this is very very different from back in the 90s, in those good old days 18 months old PC is basically useless.
I had a Q9300, and than upgraded to some i7 because I needed more slots on the motherboard. I was very surprised by how much quicker the whole system was, even if I was under the same impression, that in the last years things didn't change a lot.
CPUs get faster by around 20% per year, but in 8 years this can get compounded a lot.
Optane isn't interesting, and while I have a fascination for AVX512 from discussions with a friend who does HPC, I know I concretely won't use it most of the time. The only unique thing that Intel offers me right now is Thunderbolt 3.
Threadripper has half the memory channels and half the PCIe lanes.
EPYC is available with up to twice the cores.
The main advantage I see for Threadripper is that at 16 cores EPYC will have half the cores disabled, so for problems that fit in L2 you lose some performance from the reduced L2 sharing. That and it should be priced better than server chips, with I suspect the 12 - 14 core being the sweet spot.
Based on leaks I think Threadripper might boost a few 100 MHz higher, with base clock up to 3.6GHz.
If this applies to multi die solutions like Threadripper, there will be no 10 and 14 core parts as they would require uneven CCX combinations.
If the Epyc lineup is an indication this might be true. Especially since there is no 12 core part for Epyc. Since it has 4 dies you could produce 12 cores by using only 3 cores of each die, but this would require uneven CCX (1+2 or 0+3).
IOW core #13 is the 5th core in the second die, aka the 1st core in the second CCX on the second die. So it would have a direct link to the 1st core in the first CCX on the second die, as well as a link to the 5th core in the other 3 dies as well as a link to the 13th core on the second chip if in 2P.
So it seems quite likely that all 16 CCX's in a 2P server must have the same number of cores.
Is there any info on cache architecture for Zen?
Zen has 512K L2 per core and 8M L3 per CCX (two CCX per dice). L3 is a victim cache iirc, unlike previous generations where the L3 was inclusive.
Intel usually went with a similar scheme in the last few years, where the L3 is partitioned into slices assigned to cores; accessing the local slice is faster than a non-local slice. Skylake-SP deviates from this (significantly), for better ... or worse.
That doesnt make sense to me. I can't find any good info on this.
[0] https://www.reddit.com/r/Amd/comments/6icdyo/amd_threadrippe...
I think AMD also stated there will be no dual CPU Threadripper setups. so the CPU interconnect stuff also won't need to exist on threadripper motherboards.
See eg. or Socket 'L' vs. 'F' in AMD's history (hint -- they're the same thing), or the various Socket 2011 Intel iterations across server/enthusiast markets.
You may lose out on some features (eg. RDIMM support with server CPU, overclocking support in either direction, etc.)
Untimately, I hope, the cross-compatibility of EPYC CPUs on enthusiast chipsets will be a decision for motherboard manufacturers to make, as it has been in the past. 64 PCIe lanes should be enough for... at least some of us ;)
ETA:
As parent mentioned, the situation this time is more complicated.
Threadripper in an EPYC board is likely to be problematic -- the CPU doesn't have enough PCIe/Infinty Fabric connectivity to allow CPU links and PCIe to be active simultaneously. Even in a 1P system, only half of the board's potential PCIe lanes would be available. Due to these issues, it's likely such a setup may not be supported. It would be a strange thing to try in the first place, however, especially if equivalent EPYC parts end up similarly-priced to Threadripper.
EPYC in a Threadripper board is the news I'm hoping (and expecting) to be better -- the parent talked about powering the extra dies... an interesting consideration, though I expect (and hope) the only pins which won't be broken out on a Threadripper board will be the extra PCIe lanes.
https://news.ycombinator.com/item?id=14598660
https://www.servethehome.com/amd-epyc-7601-dual-socket-early...
eish
The Ryzen supports the AVX2 instruction set. 256-bit AVX and AVX2 instructions are split into two µops that do 128 bits each.
AVX2 increased register width from 128-bit (AVX) to 256-bit, yet Ryzen cores can only process them 128-bit at a time. There is more to AVX2 than just width but that means in comparison to Intel processors, which can do the full 256-bit in a µop, the Ryzen throughput will suffer in tests that heavily emphasize AVX2 instructions (think video encoding).
Granted, that wasn't a server chip.
[1] https://www.amd.com/en/technologies/sense-mi
There are number of research papers that attempt to use neural networks for combination optimization, but no interesting results.
- Branch Predictor
- Cache Replacement Policy
- Memory Prefetcher
- Schedulers
As for the design of CPUs, they do some automated layouts and some things like that, but I think that might not be what you mean.
https://cloudplatform.googleblog.com/2017/04/quantifying-the...
Seems reasonable to me, I'd consider getting a quad system in 2U amd system with single sockets if it beat price and perf of a quad system in 2U intel system with dual sockets.
Anybody have any info on things like L0 to L2 size, type, latencies, etc?
http://www.anandtech.com/show/11170/the-amd-zen-and-ryzen-7-...
Also, note that the numbers have changed significantly for Zen since its original launch due to updates from AMD/Manufacturers, so many of the numbers you see in older reviews are no longer accurate.
If you read reviews of the new Intel processor today, you'll see latency numbers have increased for Intel with their new architecture:
https://www.pcper.com/reviews/Processors/Intel-Core-i9-7900X...
In the end, workload-based performance metrics tend to be far more meaningful than synthetic benchmarks or simplistic latency measurements.
So for whatever cores are enabled on each die, you get the L1/L2 caches for each core as per the Ryzen launch. Additionally, you get all of the shared L3 cache, irrespective of the number of cores disabled per core complex. This pattern follows across all four dies in each socket.
https://www.theregister.co.uk/2017/06/20/amd_epyc_launch/
Various L3 sizes in the article.
Despite this, apparently Ryzen beats Kaby Labe on SPECfp, so theorethical max throughput is only part of the story. It does not help that Intel needs to heavily reduce their boost speeds when using the full AVX unit.
The review: http://www.phoronix.com/scan.php?page=article&item=ryzen-kab...
I had to do Xilinx in my CE classes. It was terrible software, we joked that the CE and EE people code that software. Crashed all the time, made me so paranoid that to this day I would often ctrl+s every few minutes just in case my IDE crashes.
Strangely, I've not seen much on HN, or elsewhere, make mention of AMD's software support. Is this because it doesn't exist, or because compilers are less "sexy" than shiny new hardware?
My take is that AMD's newest offering will be very well received by everyone who has a relatively tight budget but needs a small supercomputer on the desktop. This means data analysts and people doing all kinds of structural analysis work. As some optimization algorithms fit the definition of embarrassingly parallel, the expected turn-around time of anyone doing that sort of work will benefit greatly from the extra speed, bandwidth and core count of AMD's Ryzen/Threadripper/Epyc line.
So, quad-CPU is faster than dual-CPU? Not surprising.
IIRC, Windows Server is sold per-core however. So lots of cores may rise the total-cost of ownership in the case of Windows Server.