AWS down again?
aws.amazon.com
aws.amazon.com
Seriously, those commenting “oh boy! Time to rethink this whole cloud thing!” You’re either so new to this stuff to have no experience to remember the days before cloud, you’re trolling because you’re high on nostalgia remembering the good ol’ days (perhaps with some “member berries” https://youtu.be/mPs8-ZBSjok), or you’ve just straight up forgotten what it takes to build and run your own infra from the ground up.
And to those who happen across this comment before posting your own: please contribute something to the discussion beyond the above. Thanks.
Rather than a late-night panicked run to the smoking server, now we can just shrug and wait a few hours and it fixes itself. Most websites aren't that critical, and it's nice not having to lose sleep over devops issues.
Since Amazon failure modes are so varied, it's impossible to make the above tweaks fully automated. There'll always be a new way the cloud can misbehave, and often your service can at least maintain partial service by being nimble enough.
They take care of all the infrastructure for ya, and if something breaks, you just twiddle your thumbs while they twiddle the dials.
Of course this sort of laziness is not a valid approach if you're a huge enterprise or running some live-saving project that needs seventy nines of uptime, but for non-critical websites, it's a great sigh of relief...
It is even more, making downtime's responsibility somebody else's problem. You can now simply point to AWS is down, AWS is slow, AWS is causing error, and there is nothing we can do about it :). And as long as management knows everyone is having the same problem they are perfectly fine with it.
Heck, if that was even possible... "everything is green" dashboards... :-)
"We need this new fancy dashboard to monitor our other dashboards in one place, more accurately!"
If you accept that then it means that we need to investigate and scrutinize the whole supply chain of any product before trusting it and it would also mean that any company switching even the smallest part of their supply chain would require us to reevaluate their product.
When we buy things (either products, services or software) we consider the company selling it to be responsible for the whole and deferring responsibility to a vendor is basically just saying that you were either promising more than you knew could deliver or that you did not do a good job picking vendors.
Some of the reasons are no one wants to be seen using old tech; stuck being the last COBOL dev. So there's a big incentive for developers and their managers to be using the latest tech; it looks good on the devs resume and grants them job security in the field, and it makes their managers look like they're sharp and on top of new technology.
Also, recruiting people familiar with AWS services is often easier than finding someone versed in on-prem tech. And for companies in less than desirable locations, this helps remove a staffing issue since it can all be done remotely.
And when something goes wrong on-prem, the CTO is expected by the rest of the C-suite to fix the issue. When it's a reputable 3rd party, he can absolve both himself and his team of responsibility.
Everything up until this sentence makes sense in a twisted sort of way, but that's the part I don't get. When did we absolve companies from their choice of vendors? If I buy a service and that service does not work then the company selling that service does not get to blame their vendors, that they chose, just like they don't get to blame individual devs that they hired.
>just like they don't get to blame individual devs that they hired.
When the vendors become a symbol, a public utility status or practically the only choice. And this isn't a Tech company thing. It can be applied to all other industry. You can fire and hire another Dev. If you are looking at cloud vendor, you only have AWS, Azure and GCP to choose. For lots of reason the latter two may not be an option due to competition or features. And AWS is the best you can get, are you sure the Dev they have are the best they can hire?
The first, I can agree with, but the latter two are definitely not the case. If they were a public utility they would be regulated as one and if they were they only choice there would not be companies online during these outages. The site you are writing this on is an example, an online service with a pretty large (although not massive) following that is online enough to be the place where all the people who work at companies that depend on cloud services come to discuss their outages.
It is not whether they are one as in business or one as in politics that has to be regulated. They are utility like as people treated it as one and it behaves as one. Many things goes down with it with AWS just like many things goes down with the grid.
HN is comparatively tiny and simple. Doesn't requires any of the features on AWS. And has no reliability requirements like SaaS. HN doesn't even use CDN. Even though I am not a cloud advocate for 99% of cases it makes zero sense to build a SaaS with your own Server.
If a restaurant ran out of some menu item for a few days because a supplier didn't have a good harvest, eh, it happens. If a bookstore doesn't stock a certain book, no big deal. Grocery store self-checkout broken yet again? Oh well. Band got too drunk and can't perform? Reschedule it. No snow at the slopes? Climate change.
In the real world, shit happens, and most of it isn't critical.
If you offer a 99.999% guarantee and your customers pay for it, well, you better deliver -- subcontractors or not. But many businesses don't need to do that and a few hours of downtime a year is just a minor inconvenience. To go from "a few days of downtime" to "a few hours" is really easy; just about any provider can offer that. To go from "a few hours of downtime" to "a few seconds" is a huge investment and not worth it to many folks.
I have built and run my own infrastructure from the ground up, and I've been made to transition to the cloud.
The experience hasn't been great. It may simply be sour 'grapes' because, after all the expertise a whole generation has built up learning UNIX, and all the internet protocols (DNS, ARP, Email, reading RFCs, networking, routing) we get told that all that old stuff is just 'legacy,' and that we should retool around amazon's proprietary services instead.
Some of us old timers argued against this only to be shouted down by people who don't even understand TCP/IP.
The current generation couldn't invent the internet. You know why, cos they would never have the patience to spec it out like the old timers did. Go read a few RFCs and try to imagine a scrum team today putting as much thought into an up front design.
Today we'd just cruft together an MVP, solve only the interesting parts (or more likely the easy parts) and then move on, letting dashboards which lie to cover it up.
Many of us have been taking shots at those 'big boys' since the start of this trend.
Now that because of recent events and we have a chance to be heard you're telling us to be quiet. Why?What are you afraid of?
I haven't found it too helpful when dealing with AWS and the serverless trend, whose popularity is really just based on price and economics, not technical superiority.
Serverless turns every simple system into a distributed system with the number of failure modes now multiplied by ten.
That sounds like fun.
I do know how to use serverless for the record, I just think it's an overhyped, overpriced waste of time.
Not that that is a bad thing, but just something to think about.
Yes, of course. Not really any different from the Internet we have today, though.
There's no need to design security in the system when you literally know everyone who is using it. And everyone who was using it had the same goals in mind.
So, I don't disagree with the sentiment -- people today would probably do it a little bit differently; however, I do disagree with the expression -- people designing these protocols weren't naive. They were trustful because they had to be.
In the early days of building something new, nothing works without trust; not the Internet... not Bitcoin... not a nascent venture... nothing.
Because when TCP and it's predecessors were invented there were only a few computers in entire world. In initial ARPAnet there were only 4 hosts (In September 1973 there were 42 computers connected to 36 nodes)
But each computer had many users. That's why there were so many ports, because the thinking was there will be big computers with many users each running their own internet connected clients and servers.
That was true even in the beginning of 1990's, when I want to high school, I had access to Unix shared between 2000+ people.
Throughout the 80s, 90s, and early-to-mid 2000s, there was a certain level of trust in the claims people made about PBs (Personal Bests) and WRs (World Record/Ranking). There was no practical way to record, host, or especially upload literal hours of footage (VHS footage) of a run you did. Even if you did somehow achieve all of the above, it would be a grainy, low quality video which is hard to see, maybe with a stopwatch nearby so people can verify your claim. People would be watching this through RealPlayer, if they could watch it at all!
So what do you do in such a situation where people have no practical or easy means to verify claims? You build credibility off of how active you are with other members of the community. You post and comment on forums about what strategies you're trying, what difficulties you're dealing with, and what new information you might have uncovered through trial and error. You don't prove your work, you prove your worth. Your standing is evidence of your claim.
To me, this is a great example of "personality-credit" communities that's existed online; Usenet and BBS aside. The mentality has largely faded away with improvements to bandwidth and services like Twitch and YouTube, but considering the technological challenges of what someone in say, 1993 would be dealing with in trying to prove they just set a new record can really give a glimpse into what things used to be like.
The fact that the people who could invent the Internet mostly work for a few giants doesn't mean they no longer exist in the current generation.
Not to mention the amount of garbage in the cloud, the constant learned helplessness that we have to endure even knowing that the situation could have been avoided or even mitigated/solved if the access to the box was possible.
The status-quo of the cloud is uninspiring to say the least...
New cloud products are targeted at new companies and services.
If a company has already invested all of the R&D into building self-hosting and they've got it running properly with well-defined and measurable economics, it doesn't really make sense to upend it all and rebuild in the cloud.
But for new services, embracing hosted platforms is a proven accelerant for development. Skip past the solved problems like hosting, get straight to work on the interesting problems your business is trying to solve.
> The current generation couldn't invent the internet. You know why, cos they would never have the patience to spec it out like the old timers did.
Oh please. This is just "back in my day" ranting about "kids these days" and how you think one generation is superior to the other. Give it up.
Is there anything about this particular incident that is new that contributes to your position?
Maybe we run in different tech circles, but I feel like the "on-prem vs. cloud" has been litigated fairly extensively here and elsewhere. In fact, as you said yourself:
"Many of us have been taking shots at those 'big boys' since the start of this trend."
Your experience with UNIX isn't worthless, but it's worthless to anyone who's working "further up the stack" than you. If you feel like your skills are degrading then you need to find a job somewhere that's actually building infra that shops will build on top of. Your skills aren't day-to-day anymore and that's a good thing. Your generation made all that junk turnkey in the way that you probably think about dealing with Ethernet frames. Taking my first AWS job literally obsoleted all my system's knowledge -- super humbling experience -- it wasn't completely worthless but none of the problems I've spent years working out solutions even existed anymore.
You're confusing people building products with people building infrastructure -- the devops role makes this messy because you're usually "using" infra tools like a dev rather than building them. If you're working on foundational elements then literally nothing has changed. If you're shipping products then absolutely retool around a cloud provider, the infra isn't your secret sauce and if you have to move back on prem because of cost it will be good problems to have.
I mean this completely sincerely, take a job at a company that's providing hosting/services to people. All the old timers with deep deep systems knowledge are gods.
Even though I'm on the side of the "UNIX graybeards" here, this is a super-great point. We do need to recognize that it is a great time to be building networked applications, precisely because the younger ones these days don't need to understand TCP/IP or anything else related to infrastructure.
I confess that I got caught up in the hype as well and built quite a few "multi-AZ" apps that I thought would help me get to five nines that much faster. (for non-cloud folks, that's 99.999% availability, which was something to pursue before the cloud.)
Of course, when those abstractions break, those same younger ones are completely helpless, and my single-server apps have been running non-stop for years at traditional providers except for a few minutes for reboots following updates. I've never had a multi-hour outage, especially one that's completely out of my control where I can only point a finger at AWS and say, "sorry, it's not my fault."
Yes, there was also, at various points and times, specs and big-picture thinking, sure.
Losing skills you worked on for years is just part of this space. We are continually building on new abstractions so that we can focus on building solutions.
This really feels like a rant of kids these days. SCRUM doesn’t mean you can’t do upfront design.
4 Microsoft SQL Servers (new hardware for 2 of them) 11 IIS Web Servers (new hardware) 2 Redis Servers (new hardware) 3 Tag Engine servers (new hardware for 2 of the 3) 3 Elasticsearch servers (same) 4 HAProxy Load Balancers (added 2 to support CloudFlare) 2 Networks (each a Nexus 5596 Core + 2232TM Fabric Extenders, upgraded to 10Gbps everywhere) 2 Fortinet 800C Firewalls (replaced Cisco 5525-X ASAs) 2 Cisco ASR-1001 Routers (replaced Cisco 3945 Routers) 2 Cisco ASR-1001-x Routers (new!)S
As an old fart myself, this is very cheap argument. Old generation wasn’t superior. Old generation couldn’t invent transistor, how useless we are!
You always move up the stack as tech progress and that’s a good thing.
I look at the components of a system and think "what happens if this is turned off // breaks in an unusual way // goes slow", and ensure that the predicted effects are known and acceptable based on the likelihood of failure.
That's the same whether it's an AWS managed DNS service, storage bucket, or a raspberry pi on my desk. As a systems engineer I know what that component does, what happens when it doesn't "do", and ensure the business knows how to work around it when it breaks.
If your business can't cope with an AWS outage (even if it's not as efficient) then you've got problems.
Plan for failure, and it doesn't take you by surprise.
Often at work I see code implementations that "work" in that they usually work but can fail. I'm not a great engineer by any means (I'm actually an economist that stumbled into code), but I believe that one of the reasons I've been able to gain a good reputation at work is specifically because I design with that principle in mind.
This is a staggering bit of revisionism. As someone who was around at the time, I remember that many RFCs were written based on already-working code. They had some advantages: nobody much cared what they were doing, so they didn't have to answer to multiple levels of management, and they clearly gave almost no thought to security from bad actors, but if you think there aren't people today--including at those very cloud providers you disdain--doing work at least as well-thought-out as those early pioneers, you haven't been paying any attention at all.
I'll stay off your lawn, but maybe take off those rose-colored glasses and stop pretending the past was rosy.
> The current generation couldn't invent the internet.
Come on now, this is an overdone, lame argument and I can't believe you're seriously suggesting this. Do you also lament the fact that kids these days can't bind their own books? The point of tools is to be built upon, not to sit around marveling at your own genius. If you build a good tool, the folks that come after you don't have to think about it. That's how you make progress.
Virtualization aside, we haven't abandoned basic infrastructure, but centralized it in the hands of a few huge, expert providers. IMO this is a good thing, and was both necessary and natural as the Web grew to offer more and more opportunities to more new professionals. In detaching HTML from HTTP from ARP, etc. we gave rise to entire new professions like full-time UX (which arguably the Old Guard was never good at beyond a small audience of academics and engineers), or various flavors of front-end developer, or serverless ecosystems.
The Web and associated technologies advanced so quickly it was impractical for a single IT or network department to know all of it anymore, and some of the newer webapps wouldn't have been possible if that same team or company had to also manage all of their own basic infrastructure like it was the 90s still.
Now you can be a front-end only shop, or a UX consultant, or a network engineer who never has to touch HTML, or, or, or... maybe big enterprises always had and could always have all of those in-house, but the division of labor has been a huge boon for small businesses and startups and nonprofits, who just don't have the same resources.
As someone who grew up configuring zmodem and running BBSes and having to (mis)configure NetBIOS all the time, I am so, so glad I never have to worry about OSI layers and such ever again. It's boring to me, and the experts at it are SO much better at it, might as well let them handle it. Especially when the cost of that outsourcing is often like <$100/mo. Well worth both the time and money... and sanity. The division of concerns lets you focus on the things you're either interested in and/or good at.
Our professionals haven't gotten worse. The stack has gotten much deeper.
The problem isn't the cloud, the problem is all of the web properties going down all at once, exacerbated by the fact that AWS charges a premium for any viable cross-region / multi-cloud architecture (given its relatively high egress fees). This is discounting the inter-dependence among AWS services themselves, some of which aren't multi-region (afaik).
Engineers like to build stuff. Historically, running your own infrastructure and configuring servers from the ground up were the building blocks of the internet. These new cloud services took away those building blocks and they had the nerve to charge for it.
In the real world, there are more interesting problems to solve than setting up and maintaining your own servers so it isn't really a complaint. It's actually more fun to get the hosting stuff out of the way and focus on the problem at hand.
HN comments are basically notorious for exaggerating the downsides of hosted services while overplaying the ease and benefits of DIY. A lot of the comments here are similar to the famous Dropbox comment thread where users couldn't understand why anyone would want to use Dropbox when they could simply set up a complicated self-hosted rsync contraption to sync their own files.
This is a community of highly technical users expressing legitimate frustration and concern that the biggest player in cloud hosting is
a) having a service impacting event, and
b) the status page says everything is up
That's why people talk about the beneits of alternate hosting arrangements. If the status page said what was happening, there would be no need to ask "is AWS down?" the status would be clear.> You’re either so new to this stuff to have no experience to remember the days before cloud
I do, and it was much more peaceful.
Everything you can do on "in the cloud" you can do in colocation. Which in the long-run is cheaper, more secure* and it's yours! Including the data.
There are caveats, network, component failure. Investment in to these and you can have a pretty king setup.
The cloud enabled magnitude of email spam, brute-forcing, botnets, security vulnerabilities and much more. Operators are lazy don't want to combat it. Providers are bias, then again you can say that about any business.
People flock to the cloud like it's the greatest thing, when all your buying in to is a expensive price-plan for a company who will happily knock you off their service if you somehow brush up the wrong way and then charge you for a closed account.
I do laugh when something goes wrong for FANG. Partly, I'm cynical and want to see the world burn, but it's also these companies exploit their userbase, their staff and the environments resources. When Facebook locked themselves out of their own offices due to the BGP issue, now that's funny.
My colocation costs are:
$5000 covers for a three-year 1Gbit 2u server racking space in two different DCs. Where I have full-control, as many services as I desire, allowed to host what I desire and where the internet space is actually mine. It may be a small cube of internet but I know it's my network, my traffic.
Cloud is whitewash for me, I won't buy in to it. It has it purposes and if your happy with it, fine. But for me Colo for life.
Also, you should consider renting out space on your cube.
You could even design APIs around the internet protocols so they can create their own DNS records, mailboxes. Wait...
Now, how many people think they can do better than Uber?
1) Like any utility, the cloud is predicated on the assumption that it never goes down. If AWS itself goes down, it causes knock-on effects more most mid-sized (i.e. 1 region, multi-AZ) services on their platform, so we can't have it.
2) If we're going to have to accept downtime (1) then yes, it IS time we rethink the cloud--either how we can achieve no downtime, or how we achieve fault-tolerance in some other way.
Maybe the cloud has a higher uptime than your on-premise infrastructure (see the AWS, Azure outages). Make sure to compare the actual outage time v.s. the stats doctored by various political pressures and weaselly worded SLAs (how do you mean you had an outage? Only 49% of your requests were failing!).
Your customers will be more understanding if your outage is part if a wider outage that makes national news.
Any services you integrate with are likely down too. If two services with 99% uncorrelated uptime together drops to about 98%. It doesn't drop if the downtime is perfectly correlated.
Even if you don't directly integrate, your customer's workflow might. They see many services down and say, "cool, time to get caught up on laundry". If only you're down it's more aggrevating.
Yes they definitely are.
Getting higher uptime is super easy for smallish inhouse deployments.
You just don't install any updates and let the server run, trusting your VPN to shield you from possible security issues.
The maintenance burden is the reason why people often prefer the cloud services, not the uptime. Because maintaining the instance with updates, reading all patch notes and steps for migration, keep every health metric monitored and respond quickly on issues without getting stuck googling for possible reasons is quite a bit of work and quickly forces you to employ n+1 people.
We run some bare metal servers. They just never go down. Solid continuous pings for years as monitored from elsewhere on the Internet. That's because they're just boxes on a rack somewhere running an OS and some steady-state services (ZeroTier roots). Simplicity is more robust than complexity.
SaaS is definitely about the pain of managing and upgrading software, but it's also about OPEX vs CAPEX. Many companies will pay more for things to put them in the OPEX column for various entirely synthetic accounting, investor relations, and tax reasons.
I do wonder if the pendulum there will swing back though since if you price out cloud vs. physical hardware the market has become extremely distorted. Many companies spend enough on AWS to buy an entire rack of hardware at a different data center every month and pay 2-3 employees to manage it. That hardware would be up to 100X as fast and powerful as what they rent at AWS and bandwidth would be almost free. That's a really distorted market. The amortized costs should not be this different.
I think the issue here is that it's not zero-cost to switch. Your processes will adapt to some implicit assumptions that aren't true outside AWS, Azure, or whatever vendor you locked yourself into.
If we somehow managed to have a completely standardized interface here, the market would be more competitive.
They can try to saturate the VPN host maybe, but that's going to be challenging considering that it's going to be limited to connection requests without valid credentials.
and these are likely set to be ignored on multiple failed attempts through fail2ban or similar tooling
1. How easy can I access your physical servers ? 2. What happens if there is a catastrophic failure, for example local power outage or a major flooding 3. How secure is your server? Are you regularly patching your operation systems 4. If I want to run a project that requires double the capacity of your current hardware for a specific project, how long is it going to take to get it spun up?
- Regularly patching is automated and took about 30 seconds to configure, using an automated script
Regarding running your own physical servers, that is a different ballgame, but for all of my projects, if I need to:
- I can pretty easily spin up VPSs / bare metal servers anywhere (netcup, linode, hetzner, etc) and provision there while I wait for new hardware to come in - If you want to double the capacity of your current hardware, you'll have to order it and wait, but it's cheap (vs the major cloud providers) to way over provision if you're running your own physical hardware, so you can pretty easily have 2-4x extra capacity and still come out with extra money in your pocket.
I host in the cloud, but I think people vastly over estimate how much it saves 90% of cloud customers.
There are well-understood answers to all your questions, they are not too difficult, they just cost money - some businesses choose not to spend that money, some weigh cost-benefit and go for AWS, some decide to go for in-house servers, some go for hosted Virtual Servers, some go for serverless.
Why does everything have to be built one way?
I'd also say some choose not to spend the money, but fail to consider the cost of that choice.
For example: Doing old-school manual deployments that require herculean efforts to update at off-hours on the weekends burns people out, and makes it hard to attract new talent. In other words, you've made the decision to spend more money on finding and retaining people. But it's definitely way cheaper to pay for colocating a single Dell server you bought 3 years ago than what you'd spend in the same time on AWS.
And if your hardware never dies, paying for the redundancy might seem silly. A lot like paying for fire insurance despite the fact your house has never even burned down.
Of the options I mentioned, you seem to have picked only self-purchased hardware with no redundancy and no backups (your addition) to compare with, and also throw in manual deployment (why?).
Most businesses are somewhere between a forgotten old dell server in the closet and fully hosted multi-region auto-scaling fully bought in to AWS, and that's OK!
1: Access to the data center requires either an access card plus biometric verification (fingerprint in one DC in my case, retina scan in the other), or an ID-verified appointment. Then, you still need to know where my server is, plus the access code for the rack (or, you need to be an average lockpicker, but beware, there's cameras...).
2: Each data center has dual-provider AC feeds, plus generators, and provides A/B feeds to my rack. I've not had a dual-feed outage in the last 20 years or so.
3: No cloud provider that I'm aware of guarantees server security or does automated patching (for servers, not services). So, keeping your server up-to-date seems equally important for both cloud and non-cloud scenarios?
4: At least two weeks, I guess? I have sufficient VM host capacity to accommodate 30% unplanned growth, but 100% would require new hardware. So: ordering two servers, installing these in two data centers. But if the new project also requires significant bandwidth, getting new Internet connections in might take longer.
Look, I'm definitely not denying that "the cloud" makes it easier to scale fast, but scaling fast is not an overriding concern for most businesses. Cost is, and self-hosting, even with a pretty redundant infrastructure, is still much cheaper than AWS.
Similarly, I've had EC2 instances run for years without ever being rebooted or going down. One of them's still running today after 5 years.
But none of those services were being used 24/7; if the internet went down for an entire weekend, or a hard drive was a little bit corrupt but kept running programs in memory, I probably would never have noticed. I've also had EC2 instances literally just fall off the map and sort of disappear, and had them manually replaced without notice by AWS, had virtual drives fail and corrupt, and had calls to services fail. And I've had my desktop's power supply get fried by a power surge.
Without a lot of experience, running systems seemed easy. But as time went on I learned that it can be easy, and it also can go down if somebody blows on it the wrong way. What we see as being reliable may just be chance. The only way to guarantee reliability is to expect that things are going to go down, and design and build it accordingly.
What AWS makes easy is they give you all the components for reliability, but you have to do the plumbing yourself. I'll bet you the people whose services went down did not properly design for reliability, as they were probably running in one region, in one set of AZs, and relied on distributed system operations that can fail, and didn't properly account for how to deal with those failures. One product I maintain on AWS did not go down, but another did.
Also, the more components a system has, the higher probability there is of failure. Big systems are actually more error prone than small ones.
For any individual company able to offload the blame, that's great. It's not so great if half the countries' doorbells, robot cleaners, various home streaming service setups, the baby camera, the fridge, the smart TV and your phone stop working... All at the same time.
In my opinion, the 'downtime' really should be measured in $NUM_SERVICES_STOPPED X $TIME, instead of just $TIME. And in this case I think any long time Amazon outage is orders of magnitudes worse than your regular old slow IT company outage.
Let no more have your integrations with systems in the cloud frustrate your customers when they are unreachable - let the news explain the downtime. JOIN the Downtime Umbrella NOW and receive 5 downtime lever pulls for FREE! /s
"my outage means the customer is probably also down so they maybe don't care" is not a viable way to run a business in my mind.
A long time ago, I used to circulate snarky little emails at work.
One of them was responding to this very concept.
My managers were throwing out our working and mature UNIX servers (implementing DNS, Mail, and other services) in favour of NT.
The new system crashed a lot, we had some security breaches, but at least it was 'industry standard.'
Managements' response was that with the UNIX stuff we had no one to pin our outages on. Now we could blame Microsoft, and call their support line.
I circulated an email making fun of this justification, which promoted a fictional product called 'Blame Studio' which would help you map out the blame path for any of your products or services.
It would help to make sure that none of the blame ever landed on you, but rather was always redirected onto some other company.
What would be more valuable to you:
1. 99.5% uptime where unscheduled downtime is max 10 minutes vs.
2. 99.5% uptime where unscheduled downtime comes in 2-5 hour chunks.
These outages have almost zero impact on anyone deciding to use a cloud provider.
"Hyperscalers" like Amazon, Microsoft, Facebook and Google build their own hardware and are able to avoid many of these problems. Unfortunately, none of this stuff is available off the shelf to mere mortals.
There’s a startup trying to fix this problem (Oxide) which I think is launching their racks next year. Will be interesting to see what happens.
Not even so sure about that. I've had a ton more downtime ("degraded" in AWS speak) with AWS than any self-hosted systems. And that's with more than half my career on self-hosted.
If a major disaster strikes, like the whole rack catching fire and melting everything, then it's true that AWS could recover quicker than self-hosted. But most problems are not of that sort.
Edit: Seems to be flapping between the 'sorry' error and a blank page. Thoughts and prayers with the SREs, if they call 'em that over there.
There was a whole mature infrastructure built around UNIX and open standards at the time.
We were told to scrap that and replace it with Windows NT and its descendants.
While UNIX wasn't perfect¹, Windows was an order of magnitude worse.
Compared to our UNIX machines, those systems got the internet protocols wrong, were full of security holes and crashed all the time. Objectively worse performance and security was tolerated for a decade or more, just because we all convinced each other that this was the way things were going.
It took Microsoft maybe ten fifteen years to work their problems out. And now there's a Linux env built into Windows.
This cloud thing isn't the end of computing history, just like Windows on the server wasn't the end of history.
Cloud isn't better than on prem. It's just popular right now.
> Désolés!
> Une erreur s'est produite lorsque nous avons tenté de traiter votre requête. Soyez assuré que nous travaillons déjà à la résolution du problème que nous pensons trouver très rapidement.
> 申し訳ありません。
> リクエストの処理中に問題が発生しました。 現在問題を調査しておりますので、解決するまでもう少々お待ちください。
Amazon uses a lot of Java, so the two stories might be actually connected somehow.
Source: me
Google "site: aws.amazon.com" and try any of the links.
Most will accept it as a fact of life and continue to pay for it both directly and indirectly, as long as there's cheap money going around the cloud business can't do wrong.
But it's hilarious to see people indulging in byzantine "World scale" resilient systems that depend on a single vendor.
Have fun with that…
Real world is not so easy though...
For (almost) entirely self-contained systems it can still be useful, of course. But everything wants to be interconnected to everything these days...
There's always a service provider somewhere in the chain who can drop the ball.
AWS works 99% of the time,plus it's someone else's problem
The nice bit though is many services can go down. Yeah it stings a bit (money, reputation, time, etc). But overall it is not that big of deal.
But for the places where you can not go down. Tons of planning and tons of backup plans with backup plans, and a different style of producing code.
You can fully manage your hardware and software fleet while still paying a host to provide data center service, and networking too.
While I considered myself a decent Windows NT admin, 20 years ago, the reason I went all in on Linux and FLOSS software at the turn of the century was because I dreaded the powerlessness that these proprietary solutions gave me, when they failed. You'd call the vendors, pored through logs, finding obscure, undocumented error codes, etc. With FLOSS and self-hosting youve got all the information at your fingertips. And if you encountered bugs and you can dig into the sources, patch them, re-compile and fix things - and share them with others and feel that you're contributing to our profession.
When I got the chance to do cloud projects over the last 5 - 10 years, I always took these opportunities, hoping to ensure that I keep up with the tech. At first I was hopeful to offload the boring ops tasks, take our config management to the next level and automate even more. With every platform I got to work with, AWS, Azure and GCP so far, we kept finding bugs in their APIs, outdated or otherwise incorrect documentation and very unpredictable performance, unless you actually can run stacks at scale to average it out (as in more then just 10 - 20 instances or a larger clustered SaaS of the cloud vendor). Many times we also encountered undocumented limits that required requesting support and waiting for approval by the cloud vendor, to get even their mid-sized resources allocated. It all works very nicely on the free-tier-eligible, smallest instances and services, if as slow and high latency as is to be expected, but as soon as you actually need some decently sized storage, compute or bandwidth, it becomes quickly more expensive than what you can put in two or three datacenters for redundancy yourself, if you look at the yearly costs. So far none of the PoCs I was involved with ever got approved long term. They mostly end up as reference implementations for our customers or show case material for the corporate blog. :-(
No thanks, I don't want to go back to feel that powerless as I did on closed systems ever again. Luckily, although many seem to think that cloud is the only option to run at a global scale, you can still provide lots of valuable services on the internet using robust hardware, housed in well connected datacenters.
I know I'd rather someone else do it, then having to drive 2 hours away to a data center at 4:00 in the morning. But I don't know exactly what you're working on , I definitely can't imagine some use cases where a few hours of down time is just unacceptable. I know I wouldn't want to run a logistics firm with servers that go down all the time.
As far as I can tell, there really hasn't been an AWS outage where you couldn't have avoided issues with multi-region and some careful selection of which products you're using. Which isn't much different from what you would have to do with your own infrastructure.
For example, my company recently purchased $400k worth of hardware that would cost $80k/month if it was all EC2. There's amortized costs of the supporting infra (cooling, power) and ongoing maintenance/depreciation, but paying for the hardware in 5 months can't be beat.
It sounds like you're at sufficient scale that it makes sense. For most people, it doesn't make sense.
> Our slack group for this issue is at 3,400 people, haha. It'd be funny if I wasn't one of them. > Where do y'all work that has 5000 employees on a single issue?? > One that has an arrow under it's name.
https://<region>.console.aws.amazon.com/
I'll go ahead and make what will seem like a weird comparison. I sometimes think about this as I do about Microsoft Office.
What?
Over the years the MS Office suite has reached an asymptotic limit with regards to functionality users would actually be willing to pay for. I'll say that somewhere between 2013 and 2016 the suite had everything most users would need to do the usual stuff one does with Word, Excel, PowerPoint, etc. In other words, it became good enough and maybe even better than good enough.
This is the parallel I see with Linux servers and the Internet infrastructure needed to run a reliable site or service. Not just the OS, but the system as a whole has, over time, approached an asymptotic limit on functionality, reliability, uptime, ease of use, capacity, bandwidth, management, etc. Not sure we are at that limit yet. It sure feels like we must be close. As many have echoed, today one can deploy a range of servers in multiple configurations and achieve extremely good uptimes with as much functionality as needed without necessarily having to go to cloud service providers.
Is it the same? Not sure. I haven't operated at the highest scales in building web services, so I don't know. My gut feeling is that short of things like seriously large DDoS attacks, a self-owned infrastructure could do just as well as something hosted by a cloud provider. With knowledgeable engineers running the show there should not be any issues in setting-up, managing, maintaining and supporting such a structure.
Note that I am not hating on cloud services. No. What I am saying is that because technology tends to get better over time, it is only natural that the alternative approach has gotten better and better to the point that it would be perfectly sensible for a business to consider rolling their own rather than automatically resorting to cloud providers.
it's usually a month or more after a large outage to see the full breakdown on what happened. people expecting to see it the same week are kinda not living in reality.
at first:
We're sorry!
An error occurred when we tried to process your request. Rest assured, we're already working on the problem and expect to resolve it shortly.
and now: This page isn’t working
aws.amazon.com is currently unable to handle this request.
HTTP ERROR 503Of course if you have millions of visitors a day, you'll do better with a more robust setup, but mine costs me €80/month.
That direct URL aws.amazon.com gives me a weird error, but everything else seems to be fine.
Edit: SSH access has been restored if using their web client, everything else still times out
P.S. A comment from an exhausted engineer to another.
P.S.S. Seriously, I need to get out of the house more.
Also, the rather rare and brief maintenance window from the provider is always middle of the night for all my customers.
Are we at the point where the people who maintain the infrastructure are now completely different than the ones who built it, and are struggling to keep it running because they don’t understand it as well?
So I do think tech debt is accurate, in that the average human deludes themselves into thinking they will pay it off at some point.
I'm not sure tech debt can ever be fully paid down. You can pay off a bit, and you can stop the debt from accumulating further, but the only way to realistically unburden yourself of a really horrible, big pile of tech debt is to just rebuild the thing from scratch. Or don't, and just keep making money while you wait for the business to fail. (From an investor's perspective, this is the equivalent of riding the car into the ground.)
Look at some of the code that was open sourced by Yahoo years ago. Other large engineering organizations today have technology which far exceeds the complexity of what was made public by Yahoo a decade ago. Unfortunately Yahoo is pretty much the only example of a large company open sourcing big portions of their platform code.
Twitter used to go down so frequently the "fail whale" was a whole cultural thing.
Big AWS outages are rare, but hardly new.
https://www.theregister.com/2017/03/01/aws_s3_outage/
https://arstechnica.com/information-technology/2012/10/amazo...
Every service you mentioned (maybe sans Gmail, the only additions I remember in the last twelve years are a new UI and Hangouts) has bolted on so many features over the last years: AWS was virtual machines + SDNs in the beginning as Amazon only intended to sell spare capacity on their own servers, now it's a global one-stop-shop for everything that can be done on the Internet. Facebook was a social media feed, now it's event coordination, groups, chat, image and media hosting at global scale.
And apparently, no one at these organizations ever thought about re-working their infrastructure with "lessons learned over the last decade" in mind. Every new feature was simply bolted on, on top of an infrastructure that was hardly even envisioned to ever become the scale they are today. And that sort of refactoring costs serious amounts of money and developer time, not to mention that it doesn't make sense to develop new features on a code base that's going to be shut down in a year, so management doesn't approve it out of a fear they will be "out-featured" by a competitor and cannot react (=copy, like Instagram's Stories that were a clear rip-off from Snapchat) in time or that their own PKI/OKR goals and with it their bonus payments won't get hit.
That mindset/scale issue is also why IBM mainframes are still so common, why travel PIRs seem to be stuck in formats over half a century old or why "put CSV files on an FTP server" is the standard on bank transfers... big corporate/government clients pay a shitload of money for virtualized mainframes on new hardware that still can run the 70s-era code and even more money for people able to speak COBOL, because that is still cheaper than the alternative - reworking everything from scratch on a modern foundation, testing data integrity and edge cases, revise interfaces to hundreds or thousands of clients. Hell, even Internet standards have the same problem... we are still using protocols like BGP that have been around since before I was born, and tacked on security only a few years ago after a couple of fat-finger incidents.
Modernization in such entities only tends to happen when laws or regulatory frameworks change, and then it can become a real shitshow for those at the lowest rungs of the IT ladder that have to implement them - simply take same-sex marriages and try to shoehorn them into a database that was labeled for "husband and wife", or trans/inter people with gender data represented by a boolean field.
Of course you need to make sure the network is actually decentralized (see Solana). Or maybe you can just rely on Amazon and GCP and Azure, because surely they won’t all fail at the same time.
You usually just don’t need 100% uptime. “Sorry were closed, come back later” is fine.