Unless you have a emergency hotfix (you don’t), go hit the gym or walk outside.
If a work tool going down triggers you enough mentally to start angrily ranting online, it’s a sign you need to chill out and focus more on your health
Unless you have a emergency hotfix (you don’t), go hit the gym or walk outside.
If a work tool going down triggers you enough mentally to start angrily ranting online, it’s a sign you need to chill out and focus more on your health
But in general, it's not feasible to do everything in house.
And I don't think that GitHub is devoid of SLA: https://github.com/customer-terms/github-online-services-sla
The issue is that they're not achieving two nines uptime in practice.
I worked for a few years in an exceedingly well capitalised place which ran everything in their own data centers, money no object, with a truck parked somewhere, ready to go, with a smaller version of our critical infra. We had a serious business-stopping outage once every 18 months or so, every time for fringe reasons one only learns about when trying to run a large data center. Its convenient to blame the cloud and pretend that self-hosting in private sector was so, so great with six nines.
It's not all that different from, say, an AWS region having a service impact. People would rather complain about AWS than prepare and utilize a well-tested recovery plan to shift to a standby region. Oftentimes there's no fallback plan because the business already considered it and decided it was too costly relative to the benefit, but when the incident happens, they still can't help but complain. Humans being humans.
However, I'm also of the boomer opinion that you should get what you pay for. "Ranting online" about a service (you pay for) being unavailable is a reasonable reaction. It's not like they have a call center you can dial into for support ...
Github does have an SLA: https://github.com/customer-terms/github-online-services-sla. But like virtually all SLAs, it's really just a token gesture. I have never seen an SLA that pays the losses you suffer due to the outage.
not to mention that any business which could potentially lose enough money that they would need to let go of developers from a github outage should probably already have some business continuity plans in place.
My significant other was just let go from their job as a scapegoat for an organizational error: 3 layers of failure - IC, manager, director, and the IC was let go. The error caused a 7 figure loss for the company that has 10 figures of revenue per year. The manager and director may not see any consequences, though the director will probably be forced out by end-of-year due to incompetence. The new executive has taken to firing employees much more eagerly than their predecessor, like some sort of Jack Welch acolyte.
Their firing has put a lot of things into perspective for me. Mostly, fuck "at-will" employment and its negative effect on the American social contract.
But also this "angry ranting" online that the original poster was referencing. Not everyone has the privilege to calmly respond to things that directly impact their livelihood.
- Have never seen it mentioned on this forum, in any context.
- We're potentially pursuing legal action.50% of social media traffic is bots and i want to believe
My SO was fired - this isn't a Google review where we were treated poorly at a restaurant.
we do not bill.
the outages they promised their customers are now different.
this has costs for everyone downstream.
It all ties back to the OP, where the issue you've might not be as bad as you think. I have been in situations where we have dropped all procedures to push a hot fix because we were actively bleeding money, and in situations where you know there is an issue, and you let it be.
Maybe I pay for a service and I want that service to work consistently during core business hours.
With every single of these enterprise 'cloud' offerings you are giving (almost) complete power over your business/project to somebody else who couldn't care less about your success or failure, you are simply irrelevant for them. I see it at work too, every time critical external systems go down whole bank stops still, just because few bucks were saved yearly on some on-prem servers.
Look at it this way, you are learning some important lesson today and finding great area of improvement for resiliency from now on.
Please read https://berthub.eu/articles/posts/cyber-security-pre-war-rea...
I sympathized with them when they said this a handful of months ago, but then I saw this [0] page that shows how it's been shot for years prior (which tracks with my memory).
I feel for the GH engineers that have to deal with this, especially the SREs. I also don't hate the downtime right now, as I'll make a cup of coffee and do something else. I will say though, I did have a hotfix a week or so ago during the Actions outage, which really was a pain.
You're right that getting angry and ranting isn't the right reaction here, but I do no give them the LLM load excuse. I don't give them an out for having awful uptime during work hours for a product we pay quite a bit for as an org.
This is the outcome of violating single responsibility principle in business.
Or, maybe your #1 IT priority was moving everything to Azure instead ;)
Largest shareholder (27%), not majority.
Edit: it seems GHE is also affected by the outage and isn't much better.
- we are getting no benefit beyond getting excuses replies to our emails - right or wrong, but we depend heavily on GitHub actions so this has a very real impact on us
There is no separate enterprise platform unless you buy their self hosted server product.
Which means they have a pricing problem. Rate limit non-paying users.
It's absurd that I have dozens of repos, many with GH Actions that run CI, test and then package and push to prod/package managers, and haven't paid GH anything.
The idea that we should be fine with this unreliability is just amazing. It's not a mental health issue to have problems when important infra fails.
the comment doesnt say you should be "fine" with the unreliability.
they are saying people shouldnt get so emotionally worked up over it. which, while i wouldn't phrase it in the way the parent did, i agree with the direction of their point.
"Github’s servers are constantly on fire as their usage increased something like 50x due to LLM sloppers pushing large amounts of trash code"
"Unless you have a emergency hotfix (you don’t), go hit the gym or walk outside." with assumptions like above.
If it was, "hey they f*ed up but there's no point having an overly emotional reaction" that would be fine. But he seems to be justifying this. Especially with Github's record up to now of unreliability I think it's completely fair to be annoyed.
No qualifiers necessary. But try arguing with a lawnmower...
I have little time in my day to spend with my daughter. But I also have responsibilities. And one client has decided to bet on github. It's the only client I've ever worked with who has their code on github. EVERY other client has hosted their own gitforge or used bitbucket.
So now I am not angry because some critical piece of US-american infrastructure is down all the time.
I am angry because instead of spending time with my daughter I have to work on this later, because there are due dates and "Well, fucking GitHub was down" ain't gonna cut it.
Which they encouraged by pushing Copilot down everyone's throat
Is this rage bait? Isn't Google, Amazon, and all the other services in the world similarly impacted by LLM's? Isn't Claude, OpenAI, etc.? Why can they handle the load but not Github?
They could easily tighten things up in that regard, and make a choice that is right for their main users at the expense of The MS corpo mandate/mission. It is a choice to do otherwise.
However, your "(you don't)" comment is not going to calm down all the people, who, you know, DO have a hotfix to push now, and DO have an angry customer that could not care less for which part of our infrastructure is breaking _their_ workflow.
The only things would calm everyone down is guarantee that Microsoft would be paying for _our_ SLA breach compensation. But they don't. And I don't think anyone is paying their GH bill with a prorata of the number of time the platform was actually available.
(I'm also aware that the wording of the contract probably clearly says that you should not use GitHub for anything critical, that Microsoft is only a small startup in their garage, that you can't credibly expect 90% uptime anyway, and that it's all the fault of LLM slop ! Bad LLM slop ! Also, please buy our LLMs to generate more slop, please.)
If they are facing an influx of commits from bots maybe they need to start metering commits from those accounts and charge them money.
This is a self imposed problem.
They get paid millions of dollars by my company. They need to get it together.
Github and their parent company are active participants in pushing LLM-driven coding.
If people and companies are paying for critical infra they expect it to be available - it's not MY problem that their servers are on fire - last time I checked, having too much demand was a good problem to have and they've had over a year to fix it and come up with a plan.
I'm so glad I migrated to GitLab months ago.
I’m not a Meta fan but it’s interesting that they manage to keep their systems up with an order of magnitude more traffic. GitHub’s uptime is inexcusable.
I find it hard to believe that github APIs don't have rate limits to handle traffic spikes.
This should not be an excuse since they probably can use AI to fix it /partially sarcasm
hey bud your bias is showing
github enterprise is fully operational at the moment
https://eu.githubstatus.com/posts/dashboard
We cannot pull any actions images (or how it is called) for example.
> Unless you have a emergency hotfix (you don’t)
Oh, we do. Given the sheer number of users, it's almost guaranteed someone is on fire ever time GitHub is down. Statistics is a very charming branch of reality.
My mind immediately went to this classic xkcd: https://xkcd.com/303/
Secondly, the scale is WAY off and makes it look far worse than it is. It makes it look like 99.5% availability is practically zero availability.
Finally, GitHub’s availability is not a binary all-or-nothing proposition. They report incidents on a granular level, for example, webhook firing can be impaired while Git hosting may be working fine.
So with the "elapsed time" out of the way (now over an hour, incidentally), how about "how long will it be down?"