AWS's us-east-1 region is experiencing issues
health.aws.amazon.com
health.aws.amazon.com
Mistakes happen all the time but when all the people who intimately know how these systems work leave for other opportunities, disasters are bound to happen more and more.
Honestly the lore of w40k is quite fun to read, if you’re into dystopian and fantasy sci-fi.
It seems like both sides are fine to have them be reconciled, but it's an important narrative gadget that can be used to get humanity to fight itself in-universe.
Also interesting is that humanity's "lost" technological progress seems to eclipse some of the other races in the W40K narrative, with even the Tau (space dwarves with robots) and Eldar (space elves with crystals) freaking out when humanity brings giant robots, because the sheer physical impracticality of a gigantic human shaped robot is noted, with nobody aware of how they continue to work.
BattleTech had something similar: it was considered cheaper to keep replacing humans than to replace the mechs, because hardly anyone knew how to repair or build mechs.
Scientists only get to talk to the public every 100 years or something, wasn’t it?
IT was never allowed to talk to scientists.
Seemed like a modernist idea, even at the time of publishing.
"What were your duties at your last position?" "Performing the daily ministrations and singing the praise of the machine god."
What's causing people to believe that the latest round of attrition is any different?
I’d definitely agree that it is probably harder to become good in older organizations - the technologies are probably generations behind the current state of the art and the learning curves are high for those older technologies.
Just thinking through keyboard, but it’s probably reaching the point where enterprises need to evaluate aqui-hire or outsourcing entire development departments due to attrition, due to the incentives to leave for regular employees.
The Great Recession
The Great Resignation
The Great Dying (due to COVID-19)
Repeating this wise comment: https://news.ycombinator.com/item?id=23769427
The COVID death counts are hopelessly over-counted. This is why there's a cottage industry of people pointing out things like "COVID deaths" which mysteriously also suffered from being murdered, or drug overdoses, or undiagnosed leukaemia.
Then you get into the problem of care homes being authorised to report COVID deaths without any testing or formally trained opinion at all.
There’s very likely to be severe undercounting of COVID-19 due to the same reasons that crimes of victimization are underreported: shame.
Not to mention the swaths of homeless and disabled people that probably didn’t get counted.
https://www.medicalnewstoday.com/articles/how-are-covid-19-d...
Can you in good conscience say that the typical rate of death remained steady while the following happened:
* a majority of given populations remained at home (lock downs - meaning no travelling to work in multi-tonne death traps)
* practised increased hygiene protocols (masks, more frequent hand washing)
* did not visit elderly homes (at least, less than usual)
* many people were reluctant to get timely health care (due to fear of catching COVID from medical facilities)
* ate worse and excersied less
* infected elderly patients were sent back to their nursing homes (to typically infect the entire facility)
There are so many confounding factors that on the face of it, should result in a radically different death profile... and almost every country faced the above to different extents.
Anyone claiming to be able to work out the actual number of people who died from COVID-19 from excess deaths is being disingenous at best.
I think you have that backwards: disingenuous at worst, a scientist at best.
If your company is hostile to people sticking around for decades, then it makes it that much less likely that you end up stuck with an machine that relies heavily on poorly documented tribal knowledge that's likely to start falling apart as your core people start cashing out.
We're slowly but surely converting the world's institutional technical knowledge into re-usable and automated runbooks.
That's the dream. Obviously there are companies that sink between v1 and v2, but that's life.
Fundamentally I think the cloud business is robust, it's a fundamentally reasonable way of organising things (for enough people), which is why it attracts customers despite being arguably more expensive.
I've been in this situation in much smaller scales, and yes, you'll see massive drop in productivity but that's the cost of going from prototype to product.
I've been an SRE on a tier 1 GCP product for over three years this is not the case. In my experience, our systems have only gotten more reliable and easier to understand since I joined.
It's not like there are only a few key knowledge holders that single handedly prevent outages from happening. In reality, you don't need to know shit about how a system works internally to prevent outages if things are setup correctly.
In theory, my dog should be able to sit on my laptop and submit a prod breaking change without any fear that it will make it to prod and damage the system because automated testing/canary should catch it and, if it does make it to prod, we should be able to detect and mitigate it before the change effects more users using probes or whitebox monitoring.
This is happens for 99.9% of potential issues and is completely invisible to users. However it's what's not caught (the remaining 0.01%) that actually matters.
Google doesn’t have nearly as hard a time retaining good people as Amazon does.
Strict adherence to the "hiring bar" means we fail to bring in good people who aren't desperate enough to act out the cultish LP dance during their interview. Hiring new grads seems to be the only area where growth is not stalling - but that can obviously only help so much.
My team is hiring for 2-3 people and we are being buried alive without that growth happening sooner - but I can't in good conscience recommend this place to anyone I respect or like.
Sounds like regular skeleton-crew enterprise IT.
What is the "cultish LP dance" here that is weeding good people out?
>"My team is hiring for 2-3 people and we are being buried alive without that growth happening sooner - but I can't in good conscience recommend this place to anyone I respect or like."
I appreciate your candor. Are you in a dev role or are you on the SRE side? Is your description true across pretty much all teams/services then?
The "culture fit" interview process focuses on leadership principles, so lots of questions like " tell me about a time when you went above and beyond for a customer". Being yourself will get you nowhere, you need to research the questions and the script that is expected of you.
> What service does your team work on?
I'm a partner-focused SA, so not a developer and not aligned to a particular service.
Now I work at one of the slightly more sane FANMAG (to include $ms) companies.
Pretty sure I dodged a bullet, maybe the engineering manager spared me because he liked me more than I realized.
Ones I'd never consider, even for a cushy VP role- FB, Goggle.
Ymmv.
In my experience, they also give the toughest programming questions. It is a lot of prep overall.
Maybe the service teams are less heavy on LPs due to not being customer-facing?
Is "SA" systems admin here?
Broadly though it is a pre-sales role to help people get started, followed by ongoing guidance as the customer iterates (this is the part which often turns into free support).
Speaking of the research, the recruiters email you the principals and specifically mention to you to review them and to consider them during the interview process. They even send you a document about the STAR method of interviewing to help you have a smoother interview. To me as a guy from Baltimore who doesn't know anything half the people here do. I don't think the interview process could have been smoother.
No offer, recruiter emphasized that they were halving the “cool off” period for me (so I could interview again soon), and maybe they do this for everyone, but it’s clear there was one interview making the difference here. Interesting that this is apparently a common problem.
What gets me is that every recruiter who reaches out (hundreds by now I estimate) wants me to complete their timed coding test to qualify to interview again.
I found Google’s interview process comparatively much more respectful (although more demanding), and have been happy working there instead.
I’m sure it’s helpful to weed out applicants who actually cannot code, but what’s the point in doing it twice?
Thank you!
So, when I used the example of designing and starting to build the replacement e-mail system for ASML for literally zero additional cost [0], I pointed them at the URL for the Invited Talk that I gave based on that work. And when I used the example of when I broke all e-mail across the entire Internet, I pointed them at the URL for the article that was published in The Register. When I talked about what I had learned about Chef and DevOps, I pointed them at the invited talk I gave in Edinburgh entitled "From zero to cloud in 90 days" and the accompanying tutorial I taught called "Just enough Ruby for Chef". And so on.
I really feel that having the URLs to backstop my stories helped me sail through that part of the interview.
In my case, I wasn't being hired as an SDE, so there wasn't much programming tests they wanted me to take -- one of their senior developers did connect me to a shared coding platform, which we really just used as a shared whiteboard. He asked me some questions on how I would solve some problems, and I used my 30+ years of experience with Bourne shell and bash to show him stupid first order solutions to those problems and then we talked about what some of the second and third order solutions might be.
The longer I work at AWS, the more convinced I am that everything depends on the people you're working with. There are good teams and bad teams. There are good managers and bad managers. And if you can find a good team with good managers, then you're golden.
In this respect, I don't think AWS is materially different from any other employer I've known.
[0] They had already bought all the hardware, including some stuff I scavenged from a closet where they had been sitting new-in-the-box for a couple of years; the OS was covered under their site license; all the application software was open source; and my time was free because I was already there under another contract)
The Great Resignation had to have taken a huge toll on regular enterprises. There are probably going to be some unlucky (or lucky, depending on how hardcore they are) people in the position of maintaining aging legacy systems and retrofitting them into the future.
COBOL, for example, is becoming a lucrative language for people in the financial and insurance industries. Legacy Java is all over the place, I’m sure. Legacy .NET is in the middle of a huge industry retrofit, (.NET 5 was the official post-legacy rebrand and they’re on to .NET 6+ now).
The Great Resignation was people leaving jobs they didn't want (front of house/service industry/gig) for jobs they did want (career-track jobs) after those jobs resumed hiring again after the pandemic settled down.
Labor force participation went up not down due to 'the great resignation.'
Of our senior engineers & team leads, 70% have joined in last 6-9 months.
Only 3 full time senior engineers with 2 years or more tenure.
We've grown during COVID but we've also just burned through people.
Turnover has hit the point where we stopped doing going away zoom toasts.. people just sort of disappeared.
I would, and have. Granted not with an entire org, but it’s still good form to take a few moments to say goodbye.
If we can have meeting after meeting for working groups, agile kayfabe, status reports, etc for hours recurring weekly.. we can spend 15min on the phone saying thanks, good luck, and see you again a handful of times per year when a teammate leaves.
It takes a while to find a Vice President, I guess.
Having worked at a few other large tech companies now -- Amazon's incident response process is honestly great. It's one of the things I miss about working there.
(I have a data-driven status page for my personal website. If Oh Dear decides my website is down, the status page gets automatically updated. Obviously nobody is ever going to visit status.jrock.us if they are trying to read an article on my blog and it doesn't load, but hey at least I can say I did it.)
To make a judgement call on whether the issue is severe enough to warrant the legal/financial risk of admitting your service is broken, potentially breaking customer SLAs.
I have two arguments in favor of honest SLAs. One is, if customers expect that something is down, it can give them a piece of data with which to make their mitigation decisions. "A lot of our services are returning errors", check the auto-updated status page, "there may be an issue with network routes between AZs A and C". Now you know to drain out of those zones. If the status page says "there are no problems", now you spend hours debugging the issue, costing yourself far more money in your time than you spend on your cloud infrastructure in the first place. If having an SLA is the cause of that, it would be financially in your favor to not have the SLA at all. The SLA bounds your losses to what you pay for cloud resources, but your losses can actually be much higher; lost revenue, lost time troubleshooting someone else's problem, etc.
The second is, SLA violations are what justify reliability engineering efforts. If you lose $1,000,000 a year to SLA violations, and you hire an SRE for $500,000 a year to reduce SLA violations by 75%, then you just made a profit of $250,000. If you say "nah there were no outages", then you're flushing that $500,000 a year down the toilet and should fire anyone working on reliability. That is obviously not healthy; the financial aspect keeps you honest and accountable.
All of this gets very difficult when you are planning your own SLAs. If everyone is lying to you, you have no choice but to lie to your customers. You can multiply together all the 99.5% SLAs of the services you depend on and give your customers a guarantee of 95%, but if the 99.5% you're quoted is actually 89.4%, then you can't actually meet your guarantees. AWS can afford to lie to their customers (and Congress apparently) without consequences. But you, small startup, can't. Your customers are going to notice, and they were already taking a chance going with you instead of some big company. This is hard cycle to get out of. People don't want to lie, but they become liars because the rest of the industry is lying.
Finally, I just want to say I don't even care about the financial aspect, really. The 5 figures spent on cloud expenses are nothing compared to the nights and weekends your team loses to debugging someone else's problem. They could have spent the time with their families or hobbies if the cloud provider just said "yup, it's all broken, we'll fix it by tomorrow". Instead, they stay up late into the night looking for the root cause, finding it 4 hours later, and still not being able to do anything except open a support ticket answered by someone who has to lie to preserve the SLA. They'll never get those hours back! And they turned out to be a complete waste of time.
It's a disaster, and I don't care if it's a hard problem. We, as an industry, shouldn't stand for it.
You can't go from a metric to diagnosis for a customer - there's just no 1:1 mapping possible, with errors going both ways. AWS sucks with their status delays, but it's better than seeing their internal monitoring.
I remember one night a long time ago while working at Google, I was trying a new approach for loading some data involving an external system. To get the performance I needed, I was going to be making a lot of requests to that service, and I was a little concerned about overloading it. I ran my job and opened up that service's monitoring console, poked around, and determined that I was in fact overloading it. I sent them an email and they added some additional replicas for me.
Now I live in the world of the cloud where I pay $100 per X requests to some service, and they won't even tell me if the errors it's returning are their fault or my fault? Not acceptable.
What happens if the monitoring service goes down or has a bug that causes it to incorrectly report the status as okay? Obviously this won't usually happen at the same time as the actual service goes down, but if it did, would people really believe that?
"If you're having SLA problems I feel bad for you son
I got two 9 problems cuz of us-east-1"
This outage is reportedly impacting 5 services in 1 region.
For those impacted, pretty terrible. But as a heavy user of AWS, I’ve seen these notices posted multiple times on HN and haven’t been impacted by one yet.
us-east-1 is their largest reason, someone told me it's 50% of their revenue
multi-region failover is awfully, awfully expensive.
In my last 7 years I imagine we had ~1 outage a month on average from AWS failures, but who knows if my imagination is accurate.
A lot of institutional knowledge in these massive tech corporations is disappearing and we're starting to reach the tipping point.
Does that salary cap apply to regular enterprise developers working for HR, Accounting, etc., too?
If so, kudos to Amazon.
It’s just a little amazing to imagine that people doing the same work in different places of the country have such huge gaps in salary caps. I think the national average for a high-level software engineer is less than $150k per year.
Also, times are good and rates are crazy. Even at VARs, you can make a lot of cash. I have a buddy who went from $150k to $600k. The guy paid off his mortgage and is at a point where he could burn out and work at Home Depot if he needed to.
Source?
It is public information within America that we are to be at “Shields Up” readiness.
Multi AZ isn't that hard, but generally requires extra costs (one nat gw per az, etc...)
But multi region in AWS is a royal pain in the ass. Many services (like SSO) do not play well with multi region setups, making things really complicated even if you IaCed your whole stack.
(I actually love that we have strategies and infrastructure for multi-region... it just tends to come up at scales and for applications where it is not justified.)
What's the point of cloud if we have to manage robustness of their own infrastructure. I can understand if that's due to natural disasters and earthquakes, but the idea should be that a single AZ should never go down barring extraordinary circumstances. AWS should be auto-balancing, handling downtimes of a single AZ without the customer ever noticing it.
It might not be a good analogy, but if a single Cloudflare edge datacenter goes down, it will automatically route traffic through others. Transparent and painless to the customer. I understand AWS is huge, and different services have different redundancy mechanisms, but just conceptually it feels like they're in a conflict of interest to increase robustness of their data centers - "We told you to have multi-AZ deployment, not our fault".
Another way to put this is make sure as an AWS customer, to 3x multiply all costs + management of multi-AZ deployment into your total costs.
Worth deliberating on. I’m curious as to what the lifetime cost of ownership for an on-prem data center is relative to lifetime cost of operating in the cloud.
Do they acknowledge the problem?
It's been a joke for years how bad us-east-1 is.
It's the only way to be sure
It is a joke.
I would delete my parent comment if I could.
EDIT: Also AWS Lambda seems to be down and AWS EC2 APIs having a very high error rate and machines slow startup times.