FRA1 Block Storage Issue
status.digitalocean.com
status.digitalocean.com
We run 26 services in production on DigitalOcean. Every single VPS in our setup uses block storage feature as a persistence layer (system logs, app logs, databases, etc).
Thanks to this architecture we can rebuild machines at will. The _function_ and _state_ are nicely separated.
The downside is that we are now fucked.
Databases like low-latency local storage (the newer nvme instances on aws are very good) and logs and whatnot are aggregated with other systems (fluent, logstash, etc). I actually do not miss EBS much at all really -- if I have a problem with a VM, they're disposable and redundant and I've designed out most of the SPOFs.
As for the architecture, in DigitalOcean the external block storage is the only way to separate function from state. We need this to be able to rebuild VPS-es regularly. This is to avoid configuration drift and to prove/enforce reproducibility.
With Ceph you can distribute data across hosts, switches, racks, rows, fabrics, rooms, datacenters. If you design correctly, you can have a very resilient storage system.
Source: I come from enterprise storage and have been running Ceph in production (successfully) for ~5 years.
Sure, you're insulated from hardware failures, but not from bugs in the underlying system. Data loss is rare, but having a cluster go unavailable does happen.
No argument.
> radius of failure absolutely includes the entire cluster
Completely agree, same goes for any other storage system, network-based or otherwise.
I'm simply disagreeing with a blanket statement that all network block storage is unreliable, and that is simply not true.
Don't do silly things with Ceph, don't just have a single network fabric, don't buy cheap switches that drop packets on the floor at medium load, don't test things in production.
Can’t understate the need for good network, especially with the traffic amplification of replication.
I was running 2x10GE and had plans to scale up the cluster network to 2x or 3x10GE.
My biggest issue was the MTTR of a failed cluster. If one is RBD mirroring to a different cluster, recovery time may be hours or days.
That said, I had one cluster run continuously for a couple years with basically zero administration. It had great performance and good reliability other than that one week-long service impacting incident as a result of a bug.
I’d definitely run it again.
Everything has issues given a long enough time span though. Each layer has an accounted level of risk, and it definitely is a balancing act.
I wouldn't throw the baby out with the bathwater. Across Google, we rely almost exclusively on networked storage (Colossus) unless we need extremely high-performance local flash. Being able to separate compute from storage is a huge part of our ability to both scale out, as well as to do live migration for GCE (where Persistent Disk, our equivalent of EBS is built upon Colossus).
Persistent Disk has never had an outage like the EBS one, but I attribute that to the Colossus and underlying teams having run this at Google for a really long time. Fwiw, the AWS folks have also massively improved EBS over the years. You can still worry, but I prefer to think about overall MTBF rather than consider networked storage in particular as the plague :).
Unfortunately unlike EBS, persistent disks (Colossus) in Google cloud share a network plane with the VM. To quote the docs.
"Each persistent disk write operation contributes to your virtual machine instance's cumulative network egress cap." https://cloud.google.com/compute/docs/disks/performance
If you haven't noticed the mtu in Google cloud is < 1500 vs AWS where you can get jumbo frames (9k). I have no reason to believe persistent disks are different.
Enabling live migrations? You mean the choice of having an instance terminated with zero notice or migrate with a 60 second warning if you subscribe to the right API? Oh yeah and this is happening constantly.
Live migrations (and by connection VM attrition), persistent disks, and network performance are my least favorite aspects of google cloud today.
That being said google cloud does have alot of advantages over ec2. These are just not them.
On google cloud platform your virtual machine usually gets migrated between machines without any downtime.
Aren't you talking about preemptible instances?
Yes I have verified all the scheduling flags for our instances.
Preemptible: false OnMaintenance: MIGRATE (only other option is terminate which happens far too often for any sizable shop to use and you lose the 60 second warning)
Key word being usually. Even if it's not a complete failure we have seen all sorts of fun from packet loss/latency spikes to seg faults.
I have no doubt things will improve but right now pd, network, and live migrations are pain points for our group in gce.
https://cloud.google.com/compute/docs/instances/live-migrati...
Can you say why you prefer explicitly separated egress caps? We let our networking egress be shared between all sources of traffic on purpose, because it lets you go full throttle rather than hardcap on "flavor". That is, why restrict someone that doesn't write to disk much by "stealing" several Gbps for the PD/EBS they aren't going to use?
Finally, it's true that our MTU is too damn low. But that also isn't particularly material for PD: when the guest issues a write, we handle it all behind the scenes (it's not like your guest sees the write get fragmented into packets).
Live migrations
To clarify "this is happening constantly". I meant we see live migrations happen frequently through the day. Going back the last 24 hours I see "hundreds" of migrations. We do have days where we will see 3 or 4x this number. The majority of these were successful and our logging, probes, and graphs show nothing exciting.
On the positive side of things, we have noticed a marked improvement in migration times, probe failures, and instance fatalities in the last 6 months. Where before we would regularly see live migrations take upwards of 15 minutes or longer. They are now at or under 2 minutes (With only a handful of exceptions barely worth mentioning).
I do appreciate the facility of live migration and the proactive approach googles takes to host maintenance. The 60 second notification window is just too damn short for some of our services to properly drain themselves. So instead, we hold on to our butts and hope for the best on those boxes.
If there was one improvement to live migration it would be to have the option of a 15 minute (or even 30 minutes.. am I being greedy?) notice.
Networking egress caps shared between instance and persistent disk
The edge case that hurts here is when you have a high bandwidth service that also writes a lot of data to a PD disk and the comcast style burst bandwidth throttling that happens on the instances (This is pure speculation and may have improved since last time this was investigated. The observation at the time was the throttling is a bit too efficient and hit the instance disproportionately). We have since migrated to either local ssd or tmpfs disks for these type of hosts in gce. Their sister services are still running fine in ec2 on ebs instances.
MTU
Yay! 1500 would be better, 9k would be great. This hurts when connectivity starts to see an increase in packet loss and the corresponding increase in packet re-transmission (and latency). Not to mention the overhead these extra packets incur (Those 20-60+ bytes just in headers add up quick).
TLDR wishlist
Longer notification window before live migration actually starts PD-optimized instances 1500/9k MTU
Local copies of puppet/ansible can be very useful as well.
For logs and metrics, I'm not aware of an out of the box solution that could replay from the last, successfully transmitted line; this is something rsyslog, graphite, etc. could certainly benefit from. (Please let me know if you are aware of these kinds of buffers.)
However, distributed network block storage is usually something unavoidable after a certain size; local filers are really expensive.
filebeat + kafka does exactly this.
Can anyone speak to quality/reliability of other object storage providers that have S3-compatible (including presigned URL) APIs? S3's pricing is absolutely ridiculous by comparison, but they have the reliability argument on their side...
However, improperly designed/architected you will end up with serious scaling issues.
It's important to actually test things.
There is no other software, open source or otherwise that works quite as well as Ceph for providing durability and scale.
ScaleIO gets high marks for block storage performance compared to Ceph. It's not quite as durable and lacks some other features, but people seem to like it.
These sound like... problems with Ceph.
Sane deployment, management, and troubleshooting are critical features of any distributed system.
If they’re not there the system isn’t “ready”.
Personally, I feel like Inktank was pretty close and when they were acquired by Red Hat, progress seemed to slow to a crawl.
I haven't run the last couple of versions. It could be much improved. Certainly bluestore is promising for performance.
Salesforce runs several large Ceph clusters, and they have a dedicated team to run it. If you can't invest in the employees, you should invest in commercial support.
Salesforce also commits a lot of updates and patches back to the Ceph community
This is the part that ruled it out for me
I also looked at B2 [0] once or twice. The price is great, but the traffic cost (egress from GCE) renders it unusable for us.
But, you can build a poor-man's CDN -- varnish caches on DO/Linode/whatever where you get multiple terabytes of bandwidth for a small VM. So, you use the best object storage provider, but move most of the bits cheaply using Varnish + Route53 geo-dns.
There are actually a few options in egress land for us if cost is your primary concern.
If you're doing serving over http(s), you should probably be using Cloud CDN with your bucket [1] or put one of our partners like Cloudflare or Fastly with CDN Interconnect [2]. Both of these get you closer to $.04-$.08/GB depending on src/dest.
If not, and you don't care that we have a global backbone, you can get a more AWS-like network with our Standard Tier [3] (curiously with pricing squirreled away at [4], I'll file a bug). The packets will hop off our network in a hot potato / asap fashion, so you're not riding our backbone as much.
[1] http://cloud.google.com/cdn
[2] https://cloud.google.com/interconnect/docs/how-to/cdn-interc...
[3] https://cloud.google.com/network-tiers/
[4] https://cloud.google.com/network-tiers/pricing#standard_tier...
It's still way too expensive. And Cloud CDN isn't an appropriate tool for my use case. I really do just need a bunch of egress from a single location that isn't insanely expensive. $0.085/GB is in that insanely-expensive tier, for me.
I'd be happy to talk further via email; it's not secret, just not public.
Your bandwidth pricing is a joke. Yes, you got a nice network and yes you pay premiums to get transit of providers that are "hard to work with". And yes, you have dark fiber between your locations, which is costing a lot of money, but even considering those facts you are still charging at least 10x as much as your bandwidth should cost your customers.
How have you even calculated those prices? "Let's look at AWS and make it even more expensive"?
Not by enough, but they are.
Well, unless your S3 buckets are in us-east-1. For some reason Amazon keeps having issues with S3 in that region.
Since the storage costs appear to be the same between Spaces and S3 ($0.02/GB/month) and neither charge for inbound transfer, I'm assuming your problem is with the outbound transfer pricing (S3 charges 9x what DO charges) and/or the per-request pricing. GCP's Regional Cloud Storage has the same storage pricing, even higher outbound transfer pricing, and the same request pricing. I haven't looked at any other providers, but if you want reliability, you're going to have to pay for it.
As this is mostly a side project, I think I can live with being a little adventurous.
For distributed object storage: I have also used MooseFS,LizardFS distributed object storage and MooseFS,Lizard runs very steady on production work loads. Steady as setup and then no ops issues.
Also to the short list is BeeGFS, BeeGFS is created by Fraunhofer is seriously fast distributed file system.
It has erasure coding as well. You could deploy on bare VM's with local storage in any cloud provider and have no dependency on network blocked storage.
With k8s 1.10 you get persistent local storage as well, as such you could probably build a fairly highly available system. Pro tip: do it in GCP as they have nice local SSDs you can attach to any instance. They're 375GB, 25k IOPS, and $.08GB, way cheaper than AWS I2 instances.
Right now the leading (uncomfortable) solution is probably DigitalOcean Spaces and a little bit of prayer.
[0] https://storpool.com/ [1] http://blog.inoreader.com/2018/03/success-story-inoreader-op...
It really doesn't speak well of the Ceph architecture. It is highly performant, but at what cost? Failures on this scale can ruin a business.
-----
Hello,
On 2018-04-01 at 7:08 UTC, one of several storage clusters in our FRA1 region suffered a cascading failure. As a result, multiple redundant hosts in the storage cluster suffered an Out Of Memory (OOM) condition and crashed nearly simultaneously.
We have identified that you, or your team account, were impacted by this incident and will grant an SLA credit equal to 30% of your entire Block Storage spend for April, not just usage in FRA1. This credit will appear on your account at the end of April, and will be reflected on your April 2018 invoice.
We apologize for the incident and recognize the impact this outage had on your work and business. You can read the full detail of our public post-mortem here: http://status.digitalocean.com/incidents/8sk3mbgp6jgl
Thank you, Team DigitalOcean