I don't think anyone reasonable was suggesting that the site would just disappear overnight after such a large loss of staff. (Despite even chieftwit claiming this as evidence for business as usual.)
Which is arguable a good thing, as it allows alternatives to grow into the space
But anyone with a modicum of experience would know that makes no sense. Websites of all sizes are designed to largely self operate with no manual intervention for decent periods of time. Otherwise all large websites would be down during holidays, etc.
You also said "Otherwise all large websites would be down during holidays, etc." What makes you think large websites don't go down in some fashion or another during Holidays? Are you thinking that there would be nobody to handle incidents over a holiday?
Because every large website has people on call over holidays, and they handle incidents that would become extended outages if left unattended. Many large companies publish incident reports, and if you care to review them, you may find that indeed, their sites do experience partial outages all the time, with no exceptions for holidays (full outages are rare, of course).
Many companies do have "skeleton staffing" during major holidays, so they try to mitigate the risk of a major outage by instituting "code freezes" and not shipping new features.
That doesn't speak to how resilient those sites are, it speaks to the fact that yes, it's a normal things for sites to go down, and yes, humans are needed to keep them up. This whole thing speaks to the difference between "robust" and "resilient."
Robust: The site doesn't go down.
Resilient: When the site goes down, it comes back up with no or little impact on the user/customer experience.
Robust is an architectural property. Resilient is both an architectural and organizational property. Cutting staff impacts resilience, but not robustness. But for most users, the difference between them is invisible. If you see a site that appears to remain up over a holiday, it may be resilient rather than robust.
My last employer one day had virtually every system fail in the world at the same time due to a hidden dependency on some routing in a single data center; it took 8 hours to restart all the systems since that had not happened before. It was a fun slack to follow. There was a lot of institutional knowledge that was able to get everything working again. Imagine this kind of issue with people having no idea how things work...
I think every company eventually has at least once of these. It is usually a wake up call to the company that what used to work when they had N number of users no longer scales with 10x the traffic…
(It also usually has a lot of “I told you so” moments too)
Resilience has three aspects:
1. Anticipating and preventing problems ahead of time 2. Preventing a problem that has already occurred from getting worse 3. Recovering to a good state after a problem has occurred
First, "Anticipating and preventing problems ahead of time" is of a manifestly different character to either "Preventing a problem that has already occurred from getting worse" (sometimes called stabilization) or "Recovering to a good state after a problem has occurred" (sometimes called remediation).
Another model we use that emphasizes the difference is to divide the work into "proactive" and "reactive" categories. Everything we do that is not during an incident is proactive, everything that happens during an incident is "reactive."
It's a simpler model than the three categories you presented, but it can be helpful to highlight the fact that companies often under-invest in proactive processes and end up paying the "robustness and resilience debt" in reactive activities.
Second, I suggest that proactive work such as "Anticipating and preventing problems ahead of time" applies to both resilience and robustness. If there is an incident that is resolved by rebooting a service that has some kind of memory leak, that's reactive.
If the postmortem identifies a bug in the code causing the leak and a bug fix is deployed to prevent that incident from happening again, the system has become more robust.
Whereas, if no bug fix is made, but the runbook is updated to suggest rebooting the service when it runs out of memory, this is going to increase resilience. Proactive activities drive better outcomes in both cases.
I hope you're right, but honestly there are also legit reasons to think otherwise. It could be a fuck up due to staff shortage and knowledge loss crisis, but it could also be an impromptu, poorly implemented and communicated decision. Both scenarios would not be surprising at this point.
I can only speculate as to whether it’s embarrassing for Musk, but my assumption is that he enjoys flaunting his power to cause wreckage without justification rather than being embarrassed by it.