That seems... really bad, for the Core Core Service that Everything Depends On.
That seems... really bad, for the Core Core Service that Everything Depends On.
Note: The consent decree, the actual legal obligation of the Bell system, actually specified the percentage of times someone would pick up the phone and not get a dial tone or operator. It was surprisingly high -- IIRC something like 2%, which is why you can see people toggling the hook in old movies. The five minutes per decade constraint I know because some of my old customers made digital phone switches and they had to provide that SLA to their phone company customers, or else not get the order.
Stepping back a bit, it seems like the FCC would be in charge of establishing rules like this but that consent decree I found (which might be the wrong one), is with the house Antitrust committee. (After my own Googling failed I also asked GPT-4 and it isn't aware of this either)
However the 5 min figure comes from my customers like DSC (R.I.P), Ericsson and Nokia building POTS switches. These guys were deadly serious, like the folks who made spacecraft and medical devices, plus the Ericsson and Nokia folks were nice too.
In DSC’s case they were so paranoid that they paid us a massive amount to maintain a special tool chain for just for them. It was frozen in time (no upgrades) and when they reported a bug and we sent them an updated tool chain they diffed the binaries and made sure that every delta was due to the big fix and nothing else (that the dev hadn’t snuck in some other patch for some reason)! They did some other headstands with their hardware and software, but in the end it didn’t save them.
It’s cool that the consent decree is online. That arrangement with the Bell system was very clever, though it led to a lot of weird anomalies and distortions but I think it did end up with a better phone system than the PTT model. It’s also been a better model than what’s happened with the power and water utilities.
The FCC back then was a better regulator for phone customers than the (now obsolete) ICC had been.
We had one 18 minute issue in the ten years of use. It was the best facility we ever used.
These companies weren't idiots: they had replication for their databases and leased redundant transmission service because they knew this kind of thing could happen. The problem is you'd buy transmission from two providers...whose fiber turned out to be in the same conduit, or even both had rented bandwidth on the same fiber.
At least people are smarter these days, and with widespread cloud service there are fewer people who need to keep track of this stuff.
The equipment was trusted to not fail in it's lifetime.
Telco basically set the bar for "5 9s."
note, I worked on S3 2015-2017
So yeah you need more storage-optimized server racks and all the associated manpower and maintenance, you also need them to be distributed across different datacenters and zones which of course also impacts latency and your ability to provide some appearance of consistency, then you also need the same distribution for stateless services serving the data.
On and on and on and you might be nearly doubling cost to get to an extra 9 that almost all of your customers won't care about.
I worked on DigitalOcean's object storage for about a year not long ago. Makes sense that those of us who have been in this space would be interested in this article haha.
You're not wrong, this is really really bad, especially for Dropbox, storage is their business so I expected way better.
These stats are no different to S3 at all. All of this engineering and moving away from AWS and for so few gain in availability.
I was initially excited when they moved away from AWS and expected industry leading higher availability when they moved away, but I was wrong.
This is disappointing for Dropbox which this is their main business, storing files without any minor hiccups or outages for years.
I would expect Dropbox, a file storage company that proudly invests heavily in tech and infrastructure to achieve a better availability than what they already were on AWS (99.99%)
In terms of availability the change is pretty much 0 and as a business / enterprise customer I might as well choose a different service with similar or higher 9s or (if my needs are complex) choose S3.
https://aws.amazon.com/blogs/publicsector/achieving-five-nin...
Here is a service that has managed to achieve 5x9s of availability:
Guaranteed availability is a bet they're willing to make, a gamble they've been on top of so far, a risk that, should something fail, they will pay out on according to their SLA.
I would have thought that given Dropbox's engineering talent, they would have designed a system that would account for 5x9s and even making that guarantee for enterprise or mission critical customers.
Can't even find their SLAs anywhere for these customers, so I presume that Dropbox doesn't care about them.
Guess I was wrong and this is just disappointing and made the move not worth it.
What makes you think a gain in availability matters or is necessarily a motivation for the project?
If they can achieve the same availability at far lower cost, it’s a win for them, which is why they would (and did) do it.
This isn't a win for enterprise / business / mission critical customers. Governments and public services cannot use this at all.
> Emergency response systems is 99.999% or “five nines” – or about five minutes and 15 seconds of downtime per year.
https://aws.amazon.com/blogs/publicsector/achieving-five-nin...
I see hospitals will be using AWS and other reliable hosts to use S3 rather than Dropbox.
One can have a custom built or an experienced IT team using AWS and achieve the same or even better availability using AWS and other providers that care about critical availability, if architected correctly.
https://docs.aws.amazon.com/wellarchitected/latest/reliabili...
Storage is our business and we target an even worse 99.95% availability[1].
Availability has a cost. That cost is complexity.
We would very much prefer to have boring outages more often than have fascinating outages very rarely.
It seems that Dropbox's system is complex enough but still isn't as designed to be highly available (5x9s) as a well architected system on AWS using S3.
Yet It is surprising a public billion dollar company like Dropbox cannot go beyond this for their mission critical customers (govt agencies, healthcare, military, etc) willing to spend hundreds of millions.
So yes, I expected way better.