AWS Graviton2
perspectives.mvdirona.com
perspectives.mvdirona.com
I've been doing some initial M6g tests in my lab, and while I'm not able to disclose benchmarks, I can say that my real-world experience so far reflects what's been claimed elsewhere.
Graviton2 is going to be a game changer. It's not like the usual experience with ARM where you have to trade off performance for price, and decide whether migrating is worth the recompilation effort. In my lab, performance of the workloads I've tried so far is uniformly better than on the equivalent M5 configuration running on the Intel processor. You're not sacrificing anything by running on Graviton2.
If your workloads are based on scripting languages, Java, or Go, or you can recompile your C/C++ code, you're going to want to use these instances if you can. The pricing is going to make it irresistible. Basically, unless you're running COTS (commercial off-the-shelf software), it's a no-brainer.
Is that some theory of economical model ?
Or is that a way of saying the final price of a product varies greatly because the commodity part with the product is only a small percentage of the TCO / BOM ?
I care because it's an interesting question regarding future trends in technology. More broadly, people care about a lot more than just what has direct short term applications to their jobs.
It does have an impact on operational cost, and cost is a very important factor.
Pretty big presumption.
> Because there are so many power sensitive applications where ARMs are used, much has been invested in power minimization and management and they do very well. It’s easy to get remarkably better power consumption with an ARM part. But, in this particular case, our focus was more on server-side price/performance and, with that focus, our power consumption isn’t really materially better the alternatives.
Most customers don't measure, and don't optimize cache usage for the actual tradeoffs (cache, tlbs) to matter.
(Hi Michael!)
“Lies, damned lies, and benchmarks.” I can say that the qualitative experience is superb; everything “just works” as you’d expect it to, and performance is stellar.
It's possible to design an x86 chip with much more priority on throughput per square centimeter, with many more simple cores working together, but I have no idea how it would work out.
The vast bulk of the effort in modern JVM compiler optimisations is more at the program structural level: removing allocations, merging methods together so they can be optimised as a whole, removing abstraction, and so on. All that stuff is CPU independent.
Instead of comparing the 7nm Graviton2 processor against an 14nm Intel processor, I'd like to see its performance compared to an AMD Epyc 2 processor, which would be a more apples-to-apples comparison as both are "7nm" parts. Unfortunately Epyc 2 processors aren't available from AWS yet (but are already announced: https://aws.amazon.com/de/blogs/aws/in-the-works-new-amd-pow...).
Smaller processes used to mean higher frequency switching, lower power and increased density. With the death of Dennard scaling, we mostly just get the latter. This means that the benefit is now largely economic; you get largely the same chips, you can just pack them more tightly on the wafer.
If you're one node behind, you still price the chips at a price the market will bear, they just cost you a bit more to produce. And maybe not even then; mature last generation nodes perform pretty damn well against immature next generation nodes once you take yield and performance in to account.
Intel's 14nm transition yielded Broadwell Xeon [1], barely any improvement over the Haswell chips. Haswell itself, however, gave us a ~50% performance boost on the same process node. This is the difference between a new process node and a new architecture in today's world.
The reason Intel is in trouble is because of the 10nm fiasco, but not because of the lack of a die shrink. Their shrinks have been working like a well oiled machine for decades, and there was no contingency in place for a large delay. All post Skylake chips were being developed tightly against their 10nm libraries, with no possibility of a back port. It's not the lack of a Haswell->Broadwell analogous die shrink that's hurting Intel, but a Haswell->Broadwell->Skylake die shrink + new architecture.
How do you know this is true? Because Intel switched gears and is now decoupling future architectures from die shrinks. If they did this earlier, you'd be seeing Ice Lake (or maybe Tiger Lake) on 14nm++ as an answer to Zen 2, and it would be a pretty good chip. Instead they're doing whatever minor tweaks they can to so many variations of Skylake I'm not sure I could list all the codenames from memory.
Zen 2 is a seriously formidable chip, but most of the benefit came from cleaning up nasty edge cases in performance, like cross core communication being slower than a spill to DRAM. You can't disentangle the shrink from the architecture, because they happened simultaneously.
[1] https://www.anandtech.com/show/10158/the-intel-xeon-e5-v4-re... [2] https://www.anandtech.com/show/8423/intel-xeon-e5-version-3-...
AMD seems to disagree. In their "Next Horizon Gaming Tech Day General Session" last year they claimed that ~40% of the Zen 2 performance improvements came from "Design Frequency and 7nm Process", while the remaining ~60% are from "IPC-Enhancements" ([1] slide 13). As the frequency is directly related to the process it's obvious that moving to TSMC's 7nm process played a pretty important role for the performance improvements.
My point being that if you want to compare a state-of-the-art ARM CPU, you should compare it to a state-of-the-art x86 CPU and Intel's CPU's are simply not state-of-the-art at the moment.
Intel would be doing just fine with a 14nm Ice Lake.
A customer doesn't care about nm. They care about what's available. The apples to apples comparison is The best x86 available vs the best ARM available.
AWS offers instance types with Epyc CPUs
Their blog says the instance type will be called C5a: https://aws.amazon.com/blogs/aws/in-the-works-new-amd-powere...
Right now, according to the official C5a page, they are "coming soon".
You can buy Epyc CPUs right now even from Amazon, and you can even use Epyc CPUs in EC2 instances.
> Amazon EC2 M6g instances are currently in preview and will be generally available soon.
* ARM Servers have been inevitable for a long time but it’s great to finally see them here and in customers hands in large numbers.
How efficient is this? Can different cores have different encryption keys, so that different VMs under a hypervisor can't benefit from breaking the hypervisor's protections?
What’s going on here is that “different keys for different VMs” does not actually improve isolation without a considerable amount of hardware or microcode enforcement. AMD has this type of tracking of which VM is which. Intel does not. I don’t know what AMD does.
In any case, exception makes little difference. Cores aren’t bound 1:1 to VMs, so the core can access any VM’s data if it wants. And actually clearing the key on a context switch would require flushing caches and require that there is no cache shared between cores. The performance hit would be extreme.
Pinning a VM to a set of cores when encryption is enabled would make sense, and could be a feature cloud users would be willing to pay for.
The underlying issue here is that encryption is fast but not fast enough. So no one encrypts cache — instead, plaintext is cached and data is encrypted on its way to DRAM. So the actual isolation is in the access controls that the CPUs apply to which process or VM can access which pages, and this has little to do with encryption.
It’s worth noting that Intel has been very bad lately at protecting cache contents from side channels, while AMD has done just fine. You can turn fancy encryption on, but those side channels leak plaintext.
SEV was broken once, completely (at least on EPYC) in such a way that it could not be fixed. From what I understand.
So I'll give Intel a break here. Their performance is much better than AMDs.
The whole point of SGX is that people tried making an entire VM the security surface. That was the prior generation of tech (Intel LaGrande/TXT) and it didn't work. There's far too much code in an entire OS like Linux to make it secure or auditable (and without auditing none of these schemes mean anything).
Enclaves are a design idea that says, shrink the amount of code you have to trust and read to the smallest size possible. Only then do you have a chance of security.
It's unfortunate that this lesson has been learned and is now being lost again.
As far as I can tell, it’s only “successfully renewed” if you have HT off. If HT is on, SGX is dead.
It doesn't matter if that means designing Graviton2 or challenging Fedex by trying to build the biggest delivery network in the USA.
They could buy ARM processors available in the market, but I doubt they will be able to get them as cheap as AWS who builds their own.
Google had the software stack ready for internal workload longtime ago, PowerPC was used
https://www.forbes.com/sites/patrickmoorhead/2018/03/19/head...
It feels like Google has been directing their in-house designs on ML/TPUs while Amazon went all in on ARM. It will be interesting to see how those bets pay off.
Huawei makes their own silicon and servers with that silicon — also only internal, not available on huaweicloud :(
The only other player is Scaleway who bought first gen Cavium ThunderX's way back when. And Packet of course but that's bare metal only, no cheap small VPSes.
- Reliability (performance and availability) could be below Azure standards
- Supply chain maturity - they may have difficulty scaling procurement and deployment to meet orders
- Lock in - major cloud providers typically provide product guarantees with forenotice on the order of years before a deprecation. It's a big commitment to launch a product externally.
- Business case - maybe the TCO doesn't make sense when compared with Azure's data on demand and price point
I'm excited to see the price drop when Elasticache moves to ARM.
No, but they've been working closely with Qualcomm since the Windows Phone 7 days (10 years ago). Their recent Surface Pro X runs a customized Snapdragon 8cx dubbed "Microsoft SQ1".
I wonder if it could help bringing ARM to Azure.
It is not like Google or Microsoft does not have the expertise in house for these task. The Core and Interconnect on Graviton2 are licensed from ARM based on Neoverse. It is Fabbed on TSMC 7nm.
While there are still a lot going on with customisation, I would not be surprised if ARM have have a few solution on hand already.
The cost advantage of fabbing its own CPU is so huge, it is only a matter of time Google or Microsoft make their own CPU to compete.
I'd considering x86 in their environment but never anything I can't immediately port somewhere else.
Stock ARM maybe. Anything boutique? Nope.
Everything else has been flawless.
When I was trying to shift all my current infrastructure onto a couple of RPI's, many of the Docker containers didn't support ARM (qeumu and buildx aren't reliable) and other software didn't support ARM either.
Unless there's a good way to go from AMD to ARM, I'm not entirely sure how great Graviton or other competitors will get.
All the testing of open source stack Amazon uses internally will support ARM. That is all of their Hosted Open Sources offering. This will kickstart all software support. AWS ARM offers cost advantage which proprietary software now has an incentives or their customer will request ARM support.
All these will create a positive feedback loop into the ecosystem.
That is why it was mentioned as the fall of x86 on Servers.
AWS services like Amazon Elastic Load Balancing, Amazon ElastiCache, and Amazon Elastic Map Reduce have tested the AWS Graviton2 instances and plan to move them into production in 2020.
Normally I try to find Primary sources rather than secondary like Zdnet [1]. But I think those exact wording was quite widely reported at the time.
They say they are not Anti-Intel or AMD. Which is true. ( They are only Anti x86. ) And they say the same to UPS and Fedex at the time.
[1] https://www.zdnet.com/article/aws-graviton2-what-it-means-fo...
There are CPU HW features that Intel doesn't have which can benefit JVM workloads but they're all pretty obscure and aren't really Java specific.
What does he mean here?