Microsoft blames outage on small staff, automation failures
theregister.com
theregister.com
“Staff must park behind the building”.
I'm not sure of the technical terms, but "staff" can both be the total mass of employees, and the individual employees.
"Small staff" means a small number of employees in the same way "small army" means a small number of soldiers.
Something, something countable vs uncountable nouns?
I'll also note that it's totally standard for the theregister.com to have headlines that incorporate puns or colloquial language. In this case the original headlines has "slim staff" which is also awkward but has a different mental image :-)
For example - "original meaning" here is kind of strange. What if the original documentation is wrong or ambiguous? Then we don't want the original meaning, we want the intended meaning. But I'll leave it in here as an example.
> We have temporarily increased the team size from three to seven, until the underlying issues are better understood, and appropriate mitigations can be put in place.
> Microsoft admits slim staff and broken automation contributed to Azure outage
> Just three people were on duty in Australia when 'power sag' struck and software failures left them blind
> Storage hardware damaged by the data hall temperatures "required extensive troubleshooting" but Microsoft's diagnostic tools could not find relevant data because the storage servers were down.
Should've self-hosted it instead of trusting some cloud vendor.
Unfortunately, this will never happen.
Anyway, with "abstractions" I mean all services like AWS App Runner that are build on top of foundational services (ec2, ebs, s3, vpc) that drains resources and money at the expense of the low level stuff.
They only do it because people buy it from them.
I think our costs are around 7x what we had before with no material improvements.
Just a rant ...
Middle managers aren't going to give up their salaries when there are perfectly good underlings to sacrifice first, especially when they can just tell the Chat GPT to do the codes like they read in that ebook they bought last night.
A way I've always explained it to people is that ChatGPT is based off our knowledge and if knowledge is never improved by a human constantly updating ChatGPT it will never improve it will just make up shit to fill in the blanks and it will not be free of error. AI can loose coherency on it's data if the data it is training on is full of errors. Like data death due to compressing what is compressed over and over again.
How did a single-AZ failure cause outages for two dozen services?
Why did a single-AZ failure mean "approximately half of Cosmos DB clusters in the Australia East region were either down or heavily degraded" and require those clusters to do a cross-region failover?
Overwork and tiredness never caused any problems whatsoever, right?
Is this the purpose of incident analyses? Blame?
> Due to the size of the datacenter campus, the staffing of the team at night was insufficient to restart the chillers in a timely manner. We have temporarily increased the team size from three to seven, until the underlying issues are better understood and appropriate mitigations can be put in place.
Status report is at https://azure.status.microsoft/en-us/status/history/, the tracking code is VVTQ-J98