Reinventing virtualization with the AWS Nitro System
allthingsdistributed.com
allthingsdistributed.com
Speaking of which, if anybody who works on EC2 is reading this it would be great if you could continue exposing more MSRs and bare metal features. rr in particular would like MSR_INTEL_MISC_FEATURES_ENABLES to be available[1], which would enable trace portability for traces recorded on EC2.
[0] https://github.com/mozilla/rr [1] https://github.com/mozilla/rr/issues/2667
For the bare metal instance configurations (where the instance size type in ".metal"), there is no hypervisor running on the processor.
Approximately 30% of the x86 CPU cycles went into running the hypervisor, then they designed some new hardware, moved the virtualization functionality over to the new hardware, and apparently (in this telling) stopped accounting for the power / cycles consumed by that entirely new hardware they added to the system.
Maybe this makes accounting more clean, and it's great that it safely enables bare metal instances. However, the tone of the article hinted that under Nitro, a higher percentage of the cost / environmental cost goes into running the user's workload and then failed to account at all for the monetary or environmental cost of the added hardware. (If there's no cost savings to the customer or environmental savings for us all, why does the reader care about lower virtualization overhead? The article isn't phrased as satisfying pure intellectual curiosity.)
Taken to an extreme, they could have used dynamic translation to run all of the user workloads also on custom hardware and claimed negative infinity virtualization overhead: billions of x86-CPU cycles-worth of work done, zero x86 CPU cycles spent.
Am I misunderstanding this analysis? I don't think it's all accounting tricks, and wouldn't be surprised if dedicated virtualization hardware/firmware is more efficient at virtualization than general-purpose CPUs (and likely more cost-efficient, based on Intel / AMD's markup on CPUs). I also understand their internal hardware costs and energy costs for this new hardware might be sensitive information, but it would be nice to see some kind of accounting for this new hardware instead of apparently treating it as zero CapEx / zero OpEx / zero wattage.
As far as development and deployment goes, this is a great example of how large companies should use CapEx to build competitive barriers. They can do something hard and expensive—5 years of engineering and then the expense of custom fab—then use that to make their unit economics better (significantly more utilization of the server by customer apps) and differentiate their products. Vogels says that development began in 2012; left unsaid is that AWS knew by then that competitors were gunning for them, so how do you protect your lead? Nitro seems like one answer.
One thing to consider in the balance is the potential savings from not relying so much on Xen. There have been Xen vulnerabilities; those haven’t been fun for AWS. It’s a complex piece of software and AWS’s version was likely heavily patched, requiring a dedicated team of senior devs. The reduction in operational complexity from using Nitro plays a role, too.
Finally, you have to consider one of AWS’s big market pushes, getting Big Enterprise to transition internal business applications away from in-house hardware/data centers and onto AWS. Few in-house IT teams could likely match the performance of EC2 with Nitro; killing off managing data centers makes the CFO happy, and having improved performance on AWS makes users happy.
When I talk about CapEx, I'm talking about the portion of the whole server's CapEx represented by the Nitro hardware cost, including amortized Nitro development cost. CapEx + OpEx is the normal way to account for "total cost" of the new hardware, which can be compared against CapEx + OpEx fraction of the server's cost that can be attributed to the Intel / AMD CPUs.
Likewise for OpEx, I'm talking about the portion of the whole server's OpEx represented by the Nitro hardware operating cost.
They've offloaded 30% of the work to this new hardware, and there are two obvious ways to look at if this was a good idea: (1, evaluating monetary savings) is the cost (CapEx + OpEx) of this Nitro hardware less than that of the 30% of the box's CPU resources freed up? (2, evaluating environmental savings) is the wattage used by this Nitro hardware less than the 30% of the box's CPU resources freed up?
It also seems like an easier engineering problem to be able to give 100% vcpu to the customer rather than adding a fancy middleware that tries to fairly allocate the requested number of cpus while transparently hiding the cost of virtualization. The efficiencies of running on dedicated hardware might be great but just being able to separate is in itself pretty amazing.
Making it simpler or making the jitter better is a valuable benefit. But the claim that they are running more efficiently was completely unsubstantiated.
I’d wager that new hardware could have a huge energy efficiency gain for the specific Nitro workloads mentioned. The 30% resource utilization on Xen could be wasting a lot of time waiting for locks, context switch’s, etc between cores that could potentially be removed entirely from custom hardware optimized for handling multiple memory contexts simultaneously just by removing cache coherency at certain points.
I agree it was probably a win for them, I'm just pointing out that the way it was presented in the article was misleading.
Now a VM provider can provide the claimed amount of computing resources to the customer more accurately adhering to the advertised customers. That's offered by the part of moving shared services out of main CPUs.
Then next, people now can have bare-metally services. For example, VMWare wantted to have such system so that they can manage the whole machine with their VM software, you can imagine, Nitro make such system more easier to run on AWS machines, while at the same time, still have full control over security, IO, networking, etc.
Meanwhile, AMD uses a closed-source Arm coprocessor (PSP) for SEV features like VM memory encryption, inside their x86 CPUs. Intel has upcoming hardware with dedicated x86 silicon to run an Intel-signed TDX (Trust Domain Extensions) hypervisor for VM security features, https://www.phoronix.com/scan.php?page=news_item&px=Intel-TD...
Kudos to Annapurna for blazing the Nitro trail. Their founder has since pioneered NVME-over-TCP storage virtualization, with optional FPGA acceleration from Lightbits Labs, code upstreamed to Linux. Hopefully the next few years will bring more open-hardware interposers for storage & network paths, for academic research and commercial prototyping.
Would make sense to fit 2 c5.4xl alongside 4 more c5.2xl, for example.
That's great but what are the approved ways? This does not prevent access to customer data. Is there and built-in audit functionality to see accesses that were approved and done/attempted? This would also need to be implemented in all levels of the stack.
This basically means that AWS closed a compliance issue through technical control at the lowest level.
it depends on what you’re talking about. in the context of this article, brendan gregg measured the overhead on nitro as less than 1%.
http://www.brendangregg.com/blog/2017-11-29/aws-ec2-virtuali...
https://www.techempower.com/benchmarks/#section=data-r19&hw=...
I also don't work on Azure which is used in that benchmark, and the link I provided was specifically about benchmarking AWS.
Xeon Gold 5120: 14 cores
Seems like Nitro is radically changing AWS' hardware story.
[0] https://twitter.com/ogawa_tter/status/1108767124476981248/ph...
Edit: Here's a paper on it https://ieeexplore.ieee.org/document/9167399 and a mini-discussion on twitter: https://twitter.com/_msw_/status/1297223835519815681
http://www.brendangregg.com/blog/2017-11-29/aws-ec2-virtuali...
It shows how various subsystems improved as virtualization evolved at AWS.
As the article makes plain: It's an entire hardware stack specifically designed for virtualization from the ground up, and a corresponding software stack to utilize and manage it. There is no traditional HAL in Nitro because the hardware itself handles virtualization and resource isolation. You can't do that with traditional PC server hardware where a NIC is just a NIC and a SAS card is just a SAS card.
Xen, on the other hand, is designed to use standard PC hardware, with all the virtualization handled in software. The only VM-specific hardware support in Xen is the hypervisor support built in to modern CPUs, and occasionally some helper functions in NIC firmware. But you still have a 100% software-based HAL to provide the isolation and bare-metal simulation.
Imagine a NIC, for example, that instead of presenting itself to the OS as a single card with however many ports, can generate new virtual NICs, in hardware, and present them to the OS for assignment to a VM. The card itself manages bandwidth allocation, aggregation, encryption, and communication with cards in other hosts to create and manage VPCs.