I’m really more concerned about the absolutely miserably shitty performance of snapshots myself which always goes unmentioned and unmaintained. So I’ve got a 4TiB volume online attached to a node which is blasting out fragmented writes galore (Jenkins - ugh).
So we need to back up this volume for DR. Snapshots right? Well not when it takes 28 hours to snapshot that volume. So off to rdiff-backup it is. Back in the last infrastructure cult of virtualisation this was a 20 minute turnaround.
Good job it’s 5 nines now. That means we only risk a 3 day recovery window 2000x less often. Actually that hurt writing that.
I’m sure someone will now immediately sell me EKS to run Jenkins now which is ten fresh layers of eye gouging pain between idea to delivering software, something which I barely remember being able to do efficiently at this point.
However really the software industry is all about state and where to keep it because that’s where the trolls collect their gold coins.
So what can we do?
Here's an old rant mentioning broken container filesystems, a bit old from 2016 but storage didn't change much https://thehftguy.com/2016/11/01/docker-in-production-an-his...
Really? Can you share details? I'm asking because I know people who run databases in Docker for years and have no issues. Just store data on a volume.
> Here's an old rant mentioning broken container filesystems, a bit old from 2016 but storage didn't change much
Maybe not, haven't had a storage issue with Docker myself since maybe 2017. That article feels like a big rant but yes, it's ... old.
The issues come when you move into high availability (specifically in a Kubernetes context). K8s may decide to move your database to a different node (server) on a whim, and suddenly your data is gone. You have to solve this by having mounted, block-level storage available to your cluster (eg AWS's EBS; the product being discussed) so you can use persistent volumes. That comes with added complexity, and usually the need to incorporate db operators into your stack.
As a result, a lot of k8s admins (myself included) choose to pay crazy amounts for RDS and S3 to avoid the complexities of persistent data.
My comment was specifically discussing the challenges of keeping block storage-dependent, stateful apps within k8s (as opposed to using RDS or rolling your own HA cluster on EC2 or elsewhere). Apologies to GP if I incorrectly assumed they were discussing container volumes more broadly.
We run over 900 very stateful MySQL database clusters and several hundred zookeeper clusters in kubernetes on ec2 with ebs successfully. Sure, we've hit our share of problems over the years, but storage itself is solid.
Lots of people run their databases on Kubernetes, but it’s best to remember that Pod ephemerality makes it a more complicated orchestrator than them on VMs. That’s said, the market seems to have decided that’s what they want.
There is a gap between liking a technology and migrating a hundred critical production databases onto it.
People caught in tech fads get obsessed with mechanism over outcome.
I am paid hourly and I have literally, in the past week, billed ~14 hours to save ~$150,000 per month. We work for a client and we strive to make sure we're not spending more money in labor than we're saving in AWS costs. I don't know exactly the rate I'm billed at that the client sees, but I know for sure that on aggregate, aside from the actual productive work that I know adds value, I have justified every penny I've been paid in the last ~2 years through my cost savings work alone.
Maybe some folks overcharge or undersave, but I take exception to your comment that you're "sure is a net loss."
Edit: our monthly AWS spend is in the 7 figures. We go through cycles where our priorities change, but when we're in cost savings mode, we regularly discover cost savings measure that are a) implementable in double digit hours and b) save six figures per month.
Imagine if pricing was transparent, calculable and easy to rationalise the cost of a change. But it isn’t unless you pull a third party in.
My comment was not about the pricing of io2 in particular, but about the pricing of EBS in general.
io1/io2 is about $12/month per 100 GB, and $60/m per 1000 IOPS.
So 200 GB and 10,000 iops would be $624
With the caveat that (io2) cannot support more than 500 IOPS per GB.
gp2 is $10/m per 100 GB, and you get 3 IOPS/GB (with potential to burst to 3000 IOPS for volumes under 1 TB).
GP2 IOPS max out at 16,000 per volume.
io1/io2 IOPS max out at 64,000
Basically, use io2 if you need more than 16,000 IOPS. or you need a large amount of IOPS for only a small amount of storage.
io2 also has a 0.001% annual failure rate, while the others have a .2% failure rate.
E.g., to get 10,000 IOPS with gp2, you would get a 3.33 TB volume, which would cost $333/m.
Local storage is far, far faster (up to millions of IOPS!)...but not as reliable, and does not have all the EBS features. https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/storage-...
Is failing once every 5 years really different from failing once every 1000 years? You should have disaster recovery plans anyway and configuration errors are the most likely cause of failure, not the underlying storage system.
Thank you for the correction!
For standard volumes you pay by IO operations (absolute calls). For the newer types you pay by provisioned IOPS (a max rate). That's a particular difficult decision to make if you move in that area. The less advanced competition has simpler pricing, but then you might be more unsure about what you get. Again for high requirements you know that if someone offers you 25 or even 50 GB for 1€ per month you know it's too good to be true. For modest requirements that's something you need to consider.
But smaller/cheaper competitors might have more downtime (hard to quantify in advance), problems to keep their kernels maintained, less predictable capacity and product lifetime. So it's not without risks going for cheaper options.
Launched 11/16 and somewhat reasonably priced since 08/18. Still more expensive than some others I also use, but at these prices it might be an easier decision if you want to avoid risks with the more aggressively priced competitors.
Edit: Even with a terraform provider, nice!
However, as a person having connections to several countries I use to spin up an instance in a suitable country to work around geo-blocking of TV stations occasionally. That way you very quickly accumulate a couple of GB. Looks like for this use case this lightsail should be much cheaper, even if my bill is not a problem in absolute terms.
https://aws.amazon.com/ebs/previous-generation/
Also easy to see on the pricing calculator
Unless you're specifically looking for high-IOPS or massive HDD volume, you should just default to "gp2" SSD volumes these days.
True. When inspired by this discussion checking out what I pay for a moment I had this oh shit experience for a moment, I pay twice as much as the cheapest option. When trying to convert the first volume I noticed that the minimum size for that is 500 GB, that's 100 times my volume size, 50 times the prize, so I really don't need it.
However, as someone pointed out somewhere else in this discussion there might be a more suitable option on the low end called AWS lightsail.