We are incredibly blind to just trust just 3 cloud providers with the operational success of basically everything we do.
Why hasn't the industry come up with an alternative?
We are incredibly blind to just trust just 3 cloud providers with the operational success of basically everything we do.
Why hasn't the industry come up with an alternative?
Heck, why stop at having servers on-site? Cast your own silicon waffers, after all you don't want spectrum exploits.
Because you are worst at it. If a specialist is this bad, and the market is fully open, then it's because the problem is hard.
AWS has fewer outages in one zone alone than the best self-hosted institutions, your facebooks and petagons. In-house servers would lead to an insane amount of outage.
And guess what? AWS (and all other IAAS providers) will beg you to use multiple region because of this. The team/person that has millions of dollars a day staked on a single AWS region is an idiot and could not be entrusted to order a gaming PC from newegg, let alone run an in-house datacenter.
edit: I will add that AWS specifically is meh and I wouldn't use it myself, there's better IASS. But it's insanity to even imagine self-hosted is more reliable than using even the shittiest of IASS providers.
That might be true, but the effects of any given outage would be felt much less widely. If Disney has an outage, I can just find a movie on Netflix to watch instead. But now if one provider goes down, it can take down everything. To me, the problem isn't the cloud per se, it's one player's dominance in the space. We've taken the inherently distributed structure of the internet and re-centralized it, losing some robustness along the way.
If my system has an hour of downtime every year and the dozen other systems it interacts with and depends on each have an hour of downtime every year, it can be better that those tend to be correlated rather than independent.
If you're a company relying upon AWS for your business, is it okay if you're down for a day, or two while you wait for AWS to resolve it's issue?
Apple designed their own silicon, a third party manufactures and packages it for them.
You mean like that rule, pedant? It's not name calling if it's an accurate representation of one's behavior.
ped·ant /ˈpednt/: noun a person who is excessively concerned with minor details and rules or with displaying academic learning.
Apple designs the M1. But TSMC (and possibly Samsung) actually manufacture the chips.
Was a nightmare recovering data. Even when the service was operational was sub par.
Just saying perhaps the “shittiest” providers may not be more reliable.
It’s had two in 13 months
> Heck, why stop at having servers on-site? Cast your own silicon waffers, after all you don't want spectrum exploits.
That's an overblown argument. Nobody is saying that, but it's clear that businesses that maintain their own infrastructure would've avoided today's AWS' outage. So just avoiding a single level of abstraction would've kept your company running today.
> Because you are worst at it. If a specialist is this bad, and the market is fully open, then it's because the problem is hard.
The problem is hard mostly because of scale. If you're a small business running a few websites with a few million hits per month, it might be cheaper and easier to colocate a few servers and hire a few DevOps or old-school sysadmins to administer the infrastructure. The tooling is there, and is not much more difficult to manage than a hundred different AWS products. I'm actually more worried about the DevOps trend where engineers are trained purely on cloud infrastructure and don't understand low-level tooling these systems are built on.
> AWS has fewer outages in one zone alone than the best self-hosted institutions, your facebooks and petagons. In-house servers would lead to an insane amount of outage.
That's anecdotal and would depend on the capability of your DevOps team and your in-house / colocation facility.
> And guess what? AWS (and all other IAAS providers) will beg you to use multiple region because of this. The team/person that has millions of dollars a day staked on a single AWS region is an idiot and could not be entrusted to order a gaming PC from newegg, let alone run an in-house datacenter.
Oh great, so the solution is to put even more of our eggs in a single provider's basket? The real solution would be having failover to a different cloud provider, and the infrastructure changes needed for that are _far_ from trivial. Even with that, there's only 3 major cloud providers you can pick from. Again, colocation in a trusted datacenter would've avoided all of this.
What are good examples of
>a small business running a few websites with a few million hits per month, it might be cheaper and easier to colocate a few servers and hire a few DevOps or old-school sysadmins to administer the infrastructure.
and how often do they go down?
Heck, I use old T430 for my home server and still it doesn't go down on completely random occasions (but thats very simplified example, I know)
No idea what are the standards for other companies.
When Netflix was running its own datacenters in 2008, they had a 3 day outage from a database corruption and couldn't ship DVDs to customers. That was the disaster that pushed CEO Reed Hastings to get out of managing his own datacenters and migrate to AWS.
The flaw in the reasoning that running your own hardware would avoid today's outage is that it doesn't also consider the extra unplanned outages on other days because your homegrown IT team (especially at non-tech companies) isn't as skilled as the engineers working at AWS/GCP/Azure.
Sure, that's trivially obvious. But how many other outages would they have had instead because they aren't as experienced at running this sort of infrastructure as AWS is?
You seem to be arguing from the a priori assumption that rolling your own is inherently more stable than renting infra from AWS, without actually providing any justification for that assumption.
You also seem to be under the assumption that any amount of downtime is always unnacceptable, and worth spending large amounts of time and effort to avoid. For a lot of businesses systems going down for a few hours every once in a while just isn't a big deal, and is much more preferable than spending thousands more on cloud bills, or hiring more full time staff to ensure X 9s of uptime.
To be fair, I'm not saying never use cloud providers. If your systems require the complexity cloud providers simplify, and you operate at a scale where it would be prohibitively expensive to maintain yourself, by all means go with a cloud provider. But it's clear that not many companies are prepared for this type of failure, and protecting against it is not trivial to accomplish. Not to mention the conceptual overhead and knowledge required with dealing with the provider's specific products, APIs, etc. Whereas maintaining these systems yourself is transferrable across any datacenter.
Ovh, scaleway, online.net, azure, gcp, aws
That's one's I've used in production, I've heard of a dozen more including big names like HP and IBM, I assume they can match aws for the most part.
...
That being said I agree multi tenant is the way to go for reliability. But I was pointing out that in this case even the simple solution of multi region on one provider was not implemented by those affected.
...
As for running your own data center as a small company. I have done it, buying components building servers and all.
Expenses and ISP issues aside, I can't imagine using in house without at least a few outages a year for anywhere near the price of hiring a DevOps person to build a MT solution for you.
If you think you can you've either never tried doing it OR you are being severely underpaid for your job.
Competent teams to build and run reliable in house infrastructure exist, and they can get you SLA similar to multi region AWS or GC (aka 100% over the last 5 years)... But the price tag has 7 to 8 figures in it.
will they? because AWS still puts new stuff in us-east-1 before anywhere else, and there is often a LONG delay before those things go to other regions. there are many other examples of why people use us-east-1 so often, but it all boils down to this: AWS encourage everyone to use us-east-1 and discourage the use of other regions for the same reasons.
if they want to change how and where people deploy, they should change how they encourage it's customers to deploy.
my employer uses multi-region deployments where possible, and we can't do that anywhere nearly as much as we'd like because of limitations that AWS has chosen to have.
so if cloud providers want to encourage multi-region adoption, they need to stop discouraging and outright preventing it, first.
I also think that they've been making a lot of the default region dropdowns and such point to CMH (us-east-2) to get folks to migrate away from IAD. Your contention that they're encouraging people to use that region just don't ring true to me.
Come to think of it (far down the second page of comments): Why east?
Amazon is still mainly in Seattle, right? And Silicon Valley is in California. So one would have thought the high-tech hub both of Amazon and of the USA in general is still in the west, not east. So why us-east-1 before anywhere else, and not us-west-1?
I'm not sure how many aws services are easy to spawn at multiple regions
Doesn't help you if it what goes down is AWS global services on which you directly, or other AWS services, depend (which tend to be tied to US-east-1).
It's not AWS fault here, it's the companies', which assume that it will never be down. In-house servers also have outages, it's a very naive assumption to think that it'd be all better if all of those services were using their own servers.
Facebook doesn't use AWS and they were down for several hours a couple weeks ago, and that's because they have way better engineers than the average company, working on their infrastructure, exclusively.
However if youbare on AWS many of your competitors are down while you are down, so they can't takeover your business.
Medical devices, banks, the military, etc. should generally run on their own hardware. The next photo-sharing app? It's just not worth it until they hit tremendous scale.
On the second though, at some point, infrastructure like AWS are going to be more reliable than what many banks, medical device operators etc can provide themselves. asking them to stay on their own hardware is asking for that industry to remain slow, bespoke and expensive.
It is incredibly difficult for non-tech companies to hire quality software and infrastructure engineers - they usually pay less and the problems aren't as interesting.
If you would like to go back to the days of managing your own machines, be my guest. Remember those machines also live somewhere and were/are subject to the same BGP and routing issues we've seen over the past couple of years.
Personally, I'll deal with outages a few times a year for the peace of mind that there's a group of really talented people looking into for me.
Most financial institutions are implementing their own clouds, I can't think of any major one that is reliant on public cloud to the extent transactions would stop.
>Why hasn't the industry come up with an alternative?
You mean like building datacenters and hosting your own gear?
The agreement is more of a hybrid cloud arrangement with AWS Outposts.
FTA:
>Core to Nasdaq’s move to AWS will be AWS Outposts, which extend AWS infrastructure, services, APIs, and tools to virtually any datacenter, co-location space, or on-premises facility. Nasdaq plans to incorporate AWS Outposts directly into its core network to deliver ultra-low-latency edge compute capabilities from its primary data center in Carteret, NJ.
They are also starting small, with Nasdaq MRX
This is much less about moving NASDAQ (or other exchanges) to be fully owned/maintained by Amazon, and more about wanting to take advantage of development tooling and resources and services AWS provides, but within the confines of an owned/maintained data center. I'm sure as this partnership grows, racks and racks will be in Amazon's data centers too, but this is a hybrid approach.
I would also bet a significant amount of money that when NASDAQ does go full "cloud" (or hybrid, as it were), it won't be in the same US-east region co-mingling with the rest of the consumer web, but with its own redundant services and connections and networking stack.
NASDAQ wants to modernize its infrastructure but it absolutely doesn't want to offload it to a cloud provider. That's why it's a hybrid partnership.
vanishing or delayed six hours? I mean
The cloud is the solution to self managed data centers. Their value proposition is appealing: Focus on your core business and let us handle infrastructure for you.
This fits the needs of most small and medium sized businesses, there's no reason not to use the cloud and spend time and money on building and operating private data centers when the (perceived) chances of outages are so small.
Then, companies grow to a certain size where the benefits of having a self managed data center begins to outweight not having one. But at this point this becomes more of a strategic/political decision than merely a technical one, so it's not an easy shift.
There is a middle ground of bare-metal hosting. Abstract away the hardware and networking. Do the rest yourself.
Multiple-region redundancy costs more both in initial planning/setup as well as monthly fees so a lot of AWS customers choose to just not do it.
If the companies at the scale you are talking about do not have multi-region and multi service (aws to azure for example) failover that's their fault, and nobody else's.
Companies do not change their whole strategy from a capex-driven traditional self-hosting environment to opex-driven cloud hosting because their IT people are lazy; it is typically an exec-level decision.
We used to have that, some companies still have the capability and know-how to build and run infrastructure that is reliable, distributed across many hosting providers before "cloud" became the "norm", but it goes along with "use or lose it".