The Booming Server Market in the Wake of Skylake
nextplatform.com
nextplatform.com
Broadwell was never a huge leap for general purpose server performance. At least it didn’t get perceived as such. It seemed to concentrate on being good at the low power end to compete with ARM and more virtualization features iirc.
For some, rather like phones these days: they’re good enough not to need a refresh as often as they might have been.
For others Skylake is the first real jump / change in performance (the cache hierarchy and core interconnect topology changes are particularly interesting to me).
For yet others, they delayed refreshing for longer than they used to: “we might as well wait until the next release now”. And once Skylake was out in all its thousand-different-sku glory they bought it.
Edit: typo
- 6 memory channels (instead of 4)
- new CPU cache architecture that should show big gains for things like databases
- new vector ISA (AVX-512) that is significantly more useful than AVX2, in addition to being twice as wide
The first two should be instant wins for things like databases. AVX-512 isn't going to be used in much software yet but it is arguably the first broadly usable vector ISA that Intel has produced. This should enable some significant performance gains in the future as code is rewritten to take advantage of it. (Not idle speculation on the latter, we're queued up to get some of this hardware for exactly this purpose. Previously vectorizing wasn't worth the effort outside of narrow, special cases but AVX-512 appears to change that.)
The cache structure more closely mirrors the data locality intrinsic in recent high-performance database engines, which don't share data pages across cores and which now commonly use page sizes (256k) that don't fit in L2. The increase in cache size from 256k to 1M is particularly important because it allows you to store multiple pages in L2 which should make a number of multi-page and complex query on single page operations significantly more efficient. Should be great for join kernels. Similarly, the non-inclusive and smaller shared L3 makes more sense in that the amount of state shared across cores is actually pretty small. In short, it redistributes cache resources in a way that is very useful to the way database engines are currently designed.
There are not as many huge enterprises, but they are a huge amount of the market in terms of volume, and at those scales a 1% savings is worth years of engineering expenses, and you can bet that they will be doing the studies and choosing the most cost-effective option regardless of previous decisions.
And most of these Skylake order were placed well in advance, when they / the market didnt even have time to evaluate EPYC.
I do think they will move to EPYC in some form, just to keep Intel's price in check. After all EPYC offers much more value.
"Microsoft Announces Azure VMs with Dual 32-core AMD EPYC CPUs" - https://www.anandtech.com/show/12116/amd-and-microsoft-annou...
Wouldn't be surprised to see an equivalent AWS announcement soon, either.
It would be interesting to hear from someone who has taken delivery of one of these units.
Qualcomm Amberwing is also available to buy, although at a somewhat eye-watering price.
Edit: Teaser picture: https://rwmj.wordpress.com/2017/11/20/make-j46-kernel-builds... As usual because of NDAs I'm not allowed to publish benchmarks.