AMD server CPUs capture highest market share gains from Intel in 15 years
hothardware.com
hothardware.com
Without serious reform to the design of the Intel X86 chip to eliminate and refactor what basically amounts to a performance-before-safety feature, Intel is going to see the lions share of performance hits. Eventually people will tire of writing backflip code to sidestep pitfalls in the Intel HT design and simply return the responsibility to Intel, where its existed since day one.
It could be argued Intels mouthpiece has already lost its ability to convince major datacenter and cloud customers that HT is even remotely safe as a continued investment. Intel needs a new X86.
My Ryzen has 6 cores and 12 threads. Isn’t that pretty much the same?
Still exploits, but almost impossible to do in the field last I checked, where Intel has been demonstrated on live systems handling production load.
[0] https://en.wikichip.org/wiki/amd/microarchitectures/zen#Simu...
Hyperthreading was Intel's branding of SMT, if I understand this correctly.
Mine has 12 and 24. :)
Hyper-Threading was Intel's branding of SMT, if I understand this correctly.
My Ryzen has 12 and 24. Simply amazing what they did. :)
In practice this allows some attacks to attach themselves to the same pipeline with the target process and nibble required information slowly, but surely.
IIRC FreeBSD disabled HT out of the box at the kernel level for this reason.
I'm tired and it's late here. I may worded some stuff wrong or plainly misremembered it. Please feel free to correct.
% sysctl -d machdep.hyperthreading_allowed machdep.hyperthreading_allowed: Use Intel HTT logical CPUs
All good so far, for (say) HPC where you have people with guns guarding the machines you can turn on all the performance. The issue is that as CPUs have got more complicated, it's increasingly easy to attack the internal state of the processor. This is already possible in a single thread (i.e. the prototypical spectre and meltdown implementations), however imagine what you can do next door to a completely different "core". The processor maintains separation of structure but not of state, so you can use (say) timing side channels to extract information about what your neighbour is up to.
FWIW these issues are more of a problem for Cloud vendors than you or I as they usually require pretty specific knowledge of the hardware being attacked and even then are not easy to pull off.
As an example, in our test suite, we found that we get almost a 80% boost by turning on HT.
The right rule is “HT may help or hurt your performance depending on workload. Test and decide”.
Ignoring the performance improvements for real workloads, the #1 driving factor is performance per watt. By going with AMD that metric improves ~2x over comparable Intel offerings.
There are facilities where physical space to expand is available but the building has effectively run out of additional power that can be provided by the utility company.
Scaling up on performance per watt is the last frontier.
The AWS M4 EC2 instance uses "2.3 GHz Intel Xeon® E5-2686 v4 (Broadwell) processors or 2.4 GHz Intel Xeon® E5-2676 v3 (Haswell) processors" [0] so about 4-6 years old and may approach 7 years. A comparable processor Intel® Xeon® Processor E5-2670 v3 costs about $ 1600 [1] in bulk. Average US home electricity price is ~ 0.13 $/month [2], this would probably be much lower for a data center operator. The mentioned processors probably consume ~ 150 W at full load, which will probably not be the case most of the time. Anyway, let's see:
CPU cost over its life time may be on the order of 1600 USD, which is quite cheap as today's top processors cost usually between 4000 and 8000 USD. [3] The electricity cost would probably be: 24 hours * 365 days * 7 years * 0.13 $/kWh * 0.15 kW ~ 1200 USD which seems about right.
We can observe the electricity cost per CPU to be significantly lower than the purchase price of a rather cheap processor even assuming very generous CPU load and electricity price, while assuming rather low CPU cost. This of course doesn't include the mainboard, RAM, NIC, power conversion, cooling and other hardware that is needed to operate a modern data center. On the other hand, the cost of this hardware wasn't included in the calculation either.
Of course, I am not the only one to do the calculations. James Hamilton from AWS has done them too: https://youtu.be/kHW-ayt_Urk?t=333
[0] https://aws.amazon.com/ec2/instance-types/ [1] https://ark.intel.com/content/www/us/en/ark/products/81709/i... [2] https://www.eia.gov/todayinenergy/detail.php?id=46276 [3] https://www.anandtech.com/show/16529/amd-epyc-milan-review
Both my work and all my side projects have moved to AMD instances where available, except for some legacy on-prem stuff.
Personally I switched to AMD hardware a couple of years ago and haven't looked back, but corporations don't do that.
But those are on-prem customers, what about cloud customers? In the SOA world, surely greenfield stuff won't be using any of that proprietary Intel software.
And if AMD get the feeling they are being used this way, offering quotes with no margin will make for painful days at Intel.
Electricity, water and even just plain packaging of the cpu will likely cost more than the difference in silicon use between 14 and 7 nm
2. If the constraint on the number of packages you can sell is the number of chips you can produce, then the packaging cost of the chip is not so relevant (assuming packaging is not a constraint on production). If you can halve the chip area on the same production node, you can double production of packages, which can make a huge difference to profits (assuming Intel is a high margin business with high demand and that demand elasticity is in their favour etcetera).
Disclaimer: I am not in the industry, but what you say just seems wrong without even arguing that the cost of the silicon for Intel dominates packaging costs.
193i steppers are also dirt cheap, and most of their old fabs can also be modified for 14nm/10nm production, whereas EUV tools are 180 tonne behemoths that require overhead cranes and/or physical disassembly of the plant to move.
Cloud customers may be taking volume quotes to Intel from AMD to see what they can/will do for them on price. I don’t see why they wouldn’t do that, what with their (way out of the normal range) buying power.
I suspect AMD is getting a bigger chunk of a smaller pie as ARM makes headwind.
We moved stuff to AMD cores on Google Cloud and saw a roughly 15% reduction in utilization (and thus cost) for the same work load. Those are just Zen2 too, not Zen3 yet.
> Intel said its gross margin, the percentage of revenue remaining after deducting the cost of production, was 55.2%, down more than five percentage points from the same period in 2020. This is a key indicator of the strength of its manufacturing and product pricing. Intel has historically delivered margins above 60%.
https://www.msn.com/en-us/news/technology/intel-falls-most-i...
High clocks come at the detriment of power efficiency, so they are avoided when possible.
while an M5a is an "AMD EPYC 7000 series processors with an all core turbo clock speed of 2.5 GHz" ($0.172 for xlarge)
which is only about 10% cheaper. But the claimed clock is 20% lower. Different IPC or sustained clock might shift the balance a bit, but it seems unlikely that AMD wins decisively on performance/cost.
The 8175M's base clock 2.5Ghz, with a Turbo (all core) of 3.1Ghz and (one core) 3.5Ghz.
> assuming the IPC is similar to Intel
Using your assumption as a criteria... the Platinum 8175M should perform the same as an AMD EPYC 7763, whose base clock is also at 2.5Ghz with 3.5Ghz top Turbo speed. This is the only EPYC part that keeps nearly the exact same clocks as the Platinum at every stage.
But we know the IPCs aren't equal, so I don't even know why you'd mention that when discussing your comparison when it's so fatally flawed from the start.
Even using something as rudimentary as Passmark highlights the difference.
8175M Single Thread Rating on Passmark: 1903
EPYC 7763 Single Thread Rating on Passmark: 2639
So, clock for clock where the speed stages area identical, AMD wins.
Going through the list of every Eypc 7000 series part I could find, the one that turbos at or near 2.5Ghz is the 7551... first gen part and only hits 2.55Ghz when it's all core turbo.
EPYC 7551 Single Thread Rating: 1813.
Performance difference ratio of single thread rating between EPYC 7551 vs Platinum 8175M: 0.95270625328
Cost difference ratio of EPYC 7551 vs Platinum 8175M: 0.89583333333
Looks like AMD is still the more cost effective solution here.
On the other hand, if I take the all core passmark result divided by cores (since you pay per-core), it's 26659/24 = 1110 vs 27445/32 = 857.
They seem close enough that it's not possible to predict which one is better without actually benchmarking on AWS, since it'll depend on the clock rate they're able to sustain in AWS's setup.
In any case, I don't see any significant cost savings potential for AMD on AWS, which was the point of my original post.
[0] https://jan.rychter.com/enblog/cloud-server-cpu-performance-...
It's not just about CPU bound workloads but also things that are heavily I/O dependent (typically on bare metal hardware that you own, rather than instances you rent somewhere). There's lots of networking and storage things that would be performance bottlenecked on an intel cpu with less PCI-E lanes. Having a 16-core CPU around $950 that has 128 PCI-E 4.0 lanes is very useful for many things.
And not just for EPYC but also the single socket threadripper parts, which are used in both higher end workstations and some types of server.
This may even apply to high end gaming machines. Soon as you have 2 SSD's or video cards you will exceed link budget on most Intel CPU's and everything slows down.
It also happens to routers. 10 gig NICs plus attached SSD storage, over link budget again.
Intel's stupid market segmentation is biting them in the rear
https://www.intel.com/content/www/us/en/products/docs/networ...
In calculating the bandwidth and pci-e bus throughput needed, a single 100GbE port is full duplex, so one has to budget about 210Gbps per port.
The funny thing is that some of the best 100GbE NICs for x86-64 servers on the market right now are Intel, but are best used on an AMD platform...
We've been quite happy with Mellanox and Chelsio 100GbE NICs. The latest from each can do in-line HW TLS offload, which is a killer feature for us. No Intel NIC can do that.
IMHO the last good Intel NIC was the 10GbE "ixgbe" NIC. The design of the NIC was so tight as to be almost beautiful.
Recent 40GbE (and 10GbE based on the 40GbE chipset), and the new 100GbE NIC have the feel of being designed by a committee with endless features of questionable value stuffed in and consuming power and chip area.
If I had to make a perhaps overly broad generalization, I see more Chelsio and Mellanox NICs used in end point servers, and more Intel used in DIY whitebox network equipment.
Do you see any benefit from the fancy features? Can it source/sink min sized frames at 100GbE? (144Mpps) ?
also, cumulative number of pci-e 3.0 or 4.0 lanes in a system is a big consideration if you want to have, for instance, four dual-port 100GbE NICs all talking to one CPU. Or some mixture like three dual-port 100GbE NICs + one or two 4-port 10GbE NICs.
Where the lower end of the Intel server CPU offerings really falls flat is not having anything close to 64 or 128 PCIE lanes at a reasonable price.
Since you're limited by the 100G negotiated speed, you nic will never send more than 100G with iperf on a single port.
Yes, the 3.0 version needs two PCIe x16 slots.
By having clients who can't run away from you, you often limit yourself to the tech you can sell to them, instead of pursuing new superior alternatives.
Cisco, Oracle, SAP, IBM — all great examples of this.
Regardles, when third generation epyc Milan rolls out to AWS this year or next, the wave of movement to AMD will be massive
For c5a in us-east-1, it's simply not offered in all the AZs we've been assigned
Nothing. Unless you're explicitly targeting AVX-512, you don't miss anything. Just move the systems and continue where you left.
Furthermore, first Epyc sold so fast that, it was virtually impossible to buy in large quantities since Dropbox, FB and Google just bought the whole production out, IIRC (we weren't able to buy it, and no big IT vendor sold it under their generally available server lines).
Licensing is a part. VMware moved to per core licensing a while ago and away from per socket.
The other part is some limitations with mixed architecture clusters. Things like DRS and vmotion will be gimped.
The third part is lead times.
VMWare workstation 16 Pro is still has a flat price, regardless of core count [1].
[1]: https://store-us.vmware.com/vmware-workstation-16-pro-542417...
https://4sysops.com/archives/vmware-moving-to-per-core-licen...
Workstation isn’t used in the enterprise. That’s gonna be vsphere and esxi.
1. TSMC leading-edge process is the hottest (coolest, really ;) process and just not enough capacity to meet all the demand for mobile, high end desktop, GPUs, etc. This predates the main shortage we're talking about. Cryptocurrency mining is a contributor to this one, too.
2. There's a broad shortage of parts made on less cutting-edge processes. Causes: disruption to production from COVID, disruption from automakers churning orders, increased demand for consumer products, and speculation/hoarding.
Of course, #2 made #1 even worse.
But yeah, of those that have it in stock there's like 1-5 units at each place, and those out of stock don't have a firm date on next delivery...
Both CPU/GPUs from AMD are hard to get a hold of, at least you can get CPUs sometimes. Most of production seems to go to consoles and server CPUs.
I have been able to buy everything I've looked for at normal retail from European stores
Antonline[1] has them in stock for shipping in a bundle with a Lenovo gaming monitor.
[0]: https://www.microcenter.com/product/630282/amd-ryzen-9-5950x...
[1]: https://www.antonline.com/Lenovo/Computers/Computer_Displays...
(Once you're willing to spend $2000, you might as well just get a 3970 and have twice as many cores, if you can tolerate higher latency in exchange for higher throughput. I have a 3970 and definitely benefit from the throughput more than I would benefit from lower latency, even if it's quite noticeable. For example, some games are bottlenecked by the CPU, which is annoying. But, if you want consistent 360 fps in every game, you're spending an infinite amount of money for no financial gain anyway, so the cost-based reasoning goes out the window.)
I also love my 3970X, but this really depends how much of what you do is limited by single-thread speed and how much can actually make use of > 32 threads for sustained periods of time. Remember the 5950X is 20-30% faster in single-thread benchmarks.
If AMD wanted the latest process they had to deal with the limits TSMC said they could achieve. Intel's chip team could simply push back on their manufacturing team to improve reliability via leadership.
If Intel had gone with chiplets would 10nm have been delayed?
At no point in the past 20 years did we even come close to that.
You're probably thinking of Moore's Law, which refers to transistor count, not performance.
For the past 5 years, it's been doubling about every 3 years.
> 2005: 770
> 2006: 1540
You are doubling every year, not every two years.
https://www.cpubenchmark.net/compare/AMD-Ryzen-7-PRO-4750U-v...
The top CPU choice for my current Lenovo T460s was the i7-6600U. Exactly 5 laptop generations later the top CPU choice for the T14sG2 is the Ryzen 7 5850U. The ratio of performance between those two is 5.82. That's 42% per year or almost exactly 2x every 2 years on the same form factor and TDP. Last year's 4750U was at 4.4x after 4 years.
Intel has only been able to do 1.6x every 2 years in the same comparison. Their single core performance is comparable but they've only been able to deliver 4 cores versus AMD's 8.