And if everyone else on the AWS / CF is down as well, then it is no one's fault. We all keep calm and just wait it out.
And if everyone else on the AWS / CF is down as well, then it is no one's fault. We all keep calm and just wait it out.
It's the same as the old phrase about no one ever getting fired for buying IBM.
At the end, your customer, assuming it is a SaaS, just need an excuse for "their" customers or the End User. And there is nothing better than a big household name's fault so no one get the blame.
So the whole thing is written off, everybody is in the clear, and everyone can get on with their other business. :)
I often feel sad for people that have not developed the critical thinking skills needed to look past marketing to where the realm of logic and reason resides
The simple reality is that if your site is down because AWS is down then it means a lot of other sites are down too. Which means you don’t look anywhere near as bad as you would if you were the only site not working.
Logic and reason is all well and good but human perception is a very real thing that businesses need to keep in mind. It isn’t always entirely logical.
At most one region was impacted, and it is easy to be multi-regional in AWS, in fact that is kinda of the point
Further this comment is in service of the moronic axiom of "No one ever got fired for buying <<insert large company>>" my response to that has always been and will always be "sure they have and they should"
It is simply not true that buying AWS, IBM, Cisco, or any other large vendor is complete insulator from all responsibility to maintain reliable systems nor should it be
Any administrator or developer that is making buying choices based on that is not an person I would like to ever work or do business with
To extrapolate out to the original point: if multiple top ten traffic web sites are having issues (and this has happened once or twice in the last few years due to AWS or other cloud issues) then your site being down is less notable in customers minds.
The grandparent was talking about a SaaS service so then I would expect the customers to be technically minded people.
If you are selling to masses then sure, but if you are selling me a Line of Business SaaS service then no I do not care if facebook and reddit was down, that is not relevant to how the SaaS product should be running
For instance Basecamp could be down due to AWS. And Basecamp ( and in this case Hey as well ), but are SaaS. And their customer may not always be technically minded. Given the usage of these tools, while inside tech circle, are not only used by technical people.
And At the end of the day it is all about trade offs.
I also dont think anyone ever make a purchase decision purely on brand or not fired for X. For example, despite AMD offering lower price and offer more core and performance. Server Vendors hasn't all switched to AMD at once. In fact, Intel Server still has months of backlog order to fill. This isn't simply because Intel is better connected with Vendors, it is the fact most of those End user / customers are still demanding Intel CPU. Because it is well tested, with more specific libraries, tools, guarantees and support. Many of these factors cant be quantitatively measured, and therefore would only be judged when the final price difference are shown. In this case No one gets fire for using X is another phase for if it aren't broke, dont fix it.
Install new host -> Migrate VM;s live to new host -> Shutdown old host
Zero Downtime
This is simply not possible with a Xeon to Epyc Migration which requires downtime, as well as testing of the guest to ensure nothing weird happens
it is a prime example of Vendor Lockin.
IF there was away to live migrate with zero downtime Epyc would own the datacenter today
I’ve had colo equipment that ran with 5-nine uptime for years eventually get unlucky and be down for an hour and it was “all my fault”. Switch the service to Amazon which achieves much worse uptime, but now it’s “well if Amazon is down, what can you do?”
Frankly when something seen as core internet infrastructure goes down, the measly SaaS companies pretty reasonably don’t take any blame.
To be clear, often these are services where there isn’t the engineering budget, nor honestly the need, for multi-cloud and geographically distributed multi-master services.
But the plan B in this case is switching nameservers, which we could certainly have done (and briefly considered), but it could be error prone and would take longer for those changes to propagate than it would for CF to fix the issue, most likely.
The best option was simply inaction, if that makes sense. There are times when not doing something is a better idea than doing something.
Tangential: nobody got fired for buying IBM
Taking accountability and having backup plans are extremely important, but you simply can't remove every last shred of dependence. You eventually have to accept that there are things that are out of your control and may take you by surprise despite best efforts.
Other than that it's your choice whether to make your infrastructure dependent on a bunch of unreliable centralized SPOFs from big corporations or build highly available infrastructure relying on servers from many different providers running your own DNS servers with DNS routing, failover, etc. You will definitely beat Cloudflare's availability this way many times over.
What if a political event impacts you, for instance? A pandemic? A storm taking out a major data center? A weird Linux kernel edge case that only happens beyond a certain point in time? That only sounds ridiculous because it hasn't happened, but weird things like that happen all the time. There are so many unseen possibilities.
I understand that might sound unreasonable or facetious or like I'm expanding the scope.
The point is, the more confident that you've built something that has no SPOF the more exposed your are to the risk of it, because one probably does exist.
I remember when I first deployed DNS routed system it was too reactive, constantly jumping between servers, monitoring was too sensitive, it didn't wait for servers to stabilize to return them into the mix and exponential backoff was taking servers out for far too long. But even given all that it was still able to avoid outages caused by data center failures and connectivity problems.
> If you design for resilience, you get more resilience and you build confidence as you see the evidence how the system works in real world.
You simply can't foresee or eliminate all risk. This is referred to as "the turkey problem." It's not my idea, but one I certainly subscribe to.
https://www.convexresearch.com.br/en/insights/the-turkey-pro...
Speaking of things that don't make sense... if it's unforeseeable, one will have a difficult time adequately preparing for it
Famous to the point of being a cliche, the titanic was thought to be unsinkable, and I would have a similarly hard time convincing the engineers behind the ship's design to believe otherwise.
The level of confidence you're displaying in predicting the unforeseeable is something you may want to take a deeper look at.
I understand that most of those leetcode corporations don't care much about resilience, likely even incapable of producing highly reliable systems, and may give you a false impression that reliability is something of an unachievable fantasy. But it's not, it's something we have enough research done on and can do really well today if needed, we are not in titanic era anymore.
I have high confidence in these things (not in "predicting the unforeseeable"), because I've done them myself. My edge infrastructure had like half an hour of downtime total in many years, almost a decade already.
show me a time where all of AWS was down in every region all at the same time