EC2 Instance Update – C5 Instances with Local NVMe Storage
aws.amazon.com
aws.amazon.com
The disk is exposed as a /dev/nvme* block device, and as such I/O goes through a separate driver. The earlier versions of the driver had a hard limit of 255 seconds before I/O operation times out. [0,1,2]
When the timeout triggers, it is treated as a hard failure and the filesystem gets remounted read-only. Meaning: if you have anything that writes intensively to an attached volume, C5/M5 instances are dangerous. We experimented with them for our early prometheus nodes. Not a good idea. Having the alerts for an entire fleet start flapping due to a seemingly nonsensical "out of disk, write error" monitoring node failure is not fun.[ß]
If you run stateless, in-memory only applications on them (preferably even without local logging), then you should be fine.
0: https://bugs.launchpad.net/ubuntu/bionic/+source/linux/+bug/...
1: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/1729119
2: https://www.reddit.com/r/aws/comments/7s5gui/c5_instances_nv...
ß: We handle nodes dying. The byzantine failure mode of nodes suddenly spewing wrong data is harder to deal with.
In any event, the instance storage is unlikely to run afoul of any such timeouts since it’s more or less directly attached (albeit virtualized) and there’s no SAN involved.
Presumably the same applies to i3.metal as well, though I have not yet personally checked.
I can imagine some crazy random access, small-record, vectored IO from large thread pools. But that's not exactly common because most software that is IO-heavy tries really hard to avoid these things.
It was actually filed about EBS block disks using NVME (on these new instances, there is a hardware card that presents network EBS volumes as a PCI-E NVME device). In certain failure cases since this is a network block store, they can fail for some period of time exceeding this timeout.
The idea of this change is to ensure once they come back the machine is left in a usable state.
Of course, I would not expect to see this on Local NVME disks which is what they announced - that you can now get such instances with local disk as well as EBS.
Some more information from the original announcement: https://blog.ubuntu.com/2017/04/05/ubuntu-on-aws-gets-seriou...
(Disclaimer: I work at Canonical/Ubuntu, if that matters)
https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/nvme-ebs...
No need for a kernel with the default set - just put it in your kernel cmdline options in grub. Make an AMI with the change after that, and you're good to go.
Edit: My apologies. I didn't read your comment carefully enough, and didn't realize you had specifically called out needing to make these changes, and that older kernel versions had the 255 maximum. Leaving the above for posterity, and so people can notice I'm a bad reader :)
Did you see these issues even with it set to 255? That seems like it should have been enough timeout if everything is working normally. Perhaps small GP2 volumes that are out of credits with a large queue depth might see this? I can't say I've run into it yet.
Oh yes we did. To make sure we didn't just suffer from a single fluke, we read through docs and changed the value to 255 after the first time. Then waited. Didn't have to wait for long, the thing broke again in less than 3 weeks.
The workload was a pretty aggressively tuned prometheus. At that point we went into compaction after two weeks, so it would have been doing _very_ heavy I/O for a few days.
Some extra details about our monitoring setup here: https://smarketshq.com/wait-what-is-my-fleet-doing-2e7b1b06f... (We have 10s scraping intervals for everything and 1s for our most critical, highly latency-sensitive services.)
There's a lot to love about NVMe and timeouts are not actually part of the NVMe specification itself but rather a Linux driver construct. Unfortunately, early versions of the driver used an unsigned char for the timeout value and also have a pretty short timeout for network-based storage.
As mentioned elsewhere in the thread, recent AMIs are configured to avoid this problem out of the box.
That being said, I think one of the less compelling parts of this is that it'll probably vary per instance type quite a bit, being limited to C5 to start, so if you have a workload needing way better disk I/O than CPU performance, you might have to waste. That's one thing GCP really does have on AWS, better granularity.
I3 and F1 instances also have NVMe SSD instance store volumes https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ssd-inst...
Here's the full list of instances with instance storage: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/Instance...
Either way, I use both GCP and AWS but I am finding that I'm spending less on GCP. I don't believe every workload will be cheaper in GCP, but it is clearly competitive at least.
- What OS image(s) were used? This would aid in reproducibility.
- What tool was used for benchmarking? Preferably with the configuration so I could reproduce it.
- What GCE instance is used? I assume you will bottleneck on something other than I/O if comparing a very small GCE instance to a very large AWS instance.
I will probably benchmark this on my own at some point now that my interest is piqued, but if you have already done the work to benchmark this I'd love to hear more. So far though it's too vague for me to use meaningfully. Is there a blog post or other material that is publicly accessible? I've Googled for such before but come up surprisingly empty handed.
> dd if=/dev/zero of=/dev/disk/by-id/google-local-ssd-0 bs=8192
> 63356186624 bytes (63 GB, 59 GiB) copied, 36.6496 s, 1.7 GB/s
> dd of=/dev/zero if=/dev/disk/by-id/google-local-ssd-0 bs=8192
> 63356186624 bytes (63 GB, 59 GiB) copied, 17.5434 s, 3.6 GB/s
...But I'm perplexed by why it cuts off at 64 GB.
On Amazon I couldn't figure out how to get decent speeds at all. Comparable to my Google Cloud box, I set up an i3.xlarge, which had 4 CPUs and 1 SSD. While the dd operation was working fine, it slowed down to a paltry 430 MB/s quickly and never recovered. I doubt this is the true sequential performance of the drive, so I gave up.
I would've attempted to do multiple SSD configurations, but those seemed even harder. mdadm-based RAIDs were clearly bottlenecking somewhere other than hardware, because they were slower than writing or reading directly.
tl;dr: It turns out these new-fangled NVMe instance store setups are too complicated for me to understand.
This isn't the case only on new-fangled NVMe devices...
https://cloud.google.com/compute/docs/disks/performance#type...
Is Google being disingenuous? I haven't benchmarked it, but I will say that anecdotally the GCP local NVMe SSDs feel very fast for things like building software compared to other setups I've done on other clouds (AWS, UpCloud.)
EBS volumes are great and all but not for database where the dataset is many multiples of the working set.
It’s totally safe to use local storage if you build it right. But those raided EBSs caused a lot of problems. In short, when one gets slow the whole volume gets slow because software raid isn’t hardware raid.
The main advantage of RDS is that they take care of the mundane redundancy for you.
As a database operator I treat safety on i3 similarly where I have multiple hot replicas of my data so that if any fails I’m good to go. Additionally, there isn’t any reason you couldn’t have a EBS replica of an ephemeral node.
What we typically do with i3 is mirror the data locally, replicate it, have an EBS replica, and take backups. This is probably overkill but the data needs to be both accessed quickly and secure so that’s where we are at.
Is it automatic or manual?
On infrastructure I handled from top to bottom, I used VIPs with keepalived (only the vrrp part, with a weight linked to success/failure of a check script).
But in AWS, I'm wondering how to do it properly, maybe DNS records with low TTL (like 1 second).
As you mentioned, it is limited by instance size, but for a DB that fits it works great and has fewer moving parts. Knowing that your entire database is essentially ephemeral raises the stakes too and forces you to take replication, backups and restore testing seriously.
I don't think you should trust your data to a single disk, whether or not it's a physical device in your own datacenter or an EBS in AWS. Everything fails eventually.
> Encryption – Each local NVMe device is hardware encrypted using the XTS-AES-256 block cipher and a unique key. Each key is destroyed when the instance is stopped or terminated.
Does anyone know if the existing i3 EC2 instances NVMe drives are also encrypted in this fashion? I can't find any articles stating this...
Thanks!
The documentation at https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ssd-inst... will be updated soon with the same information.
The bottom of the article mentions "PS – We will be adding local NVMe storage to other EC2 instance types in the months to come, so stay tuned!"
Local SSD is just orders of magnitude faster if you need a fast place to put temporary (or replicated) data.
io1's top out at 32K PIOPS per volume and a 32K PIOPS 225GB volume would cost around 2100/month.
But an entire c5d.2xlarge with 225GB of local SSD will only cost around $283/month, and (based on i3 performance) will give you around 180K write IOPS.
jenkins workspaces