Normally I defend GH in the comments of these incidents but it’s been an impressively bad month by their standards, even when you filter for critical components filter out sev-2’s and 3’s.
Normally I defend GH in the comments of these incidents but it’s been an impressively bad month by their standards, even when you filter for critical components filter out sev-2’s and 3’s.
They should install OpenClaw for that as well.
Not at all, you merely move the goal post of at what layer the "root cause" actually could come from! At that speed, it's always something short and sweet, while when you actually want to long-term address things, you have to have time to even investigate organizational issues or whatever the actual problems stem from.
But you have half a day? "Post-mortem: Push X wasn't properly analyzed before deployment, in future more testing" and call it a day.
https://damrnelson.github.io/github-historical-uptime/
Unfortunately, it doesn't look like it's being updated with new data. But it wouldn't look any better for GH if it was.
https://github.com/DaMrNelson/github-historical-uptime/issue...
GH going down used to be quite rare. If it failed to load I'd spend a bunch of time trying to figure out what was wrong with my internet connection, just to read on HN that it was down for everyone.
This week GH failed to load and I automatically assumed it was a GH issue - just for it to be followed up a few minutes later by a marketing coworker complaining about internet connectivity. Turns out the office internet connection was dropping about 50% of all packets.
It is bad enough that business-side managers are noticing that GH issues are slowing work down. That would've been unimaginable a few years ago.
Is it possible that there has been a change in the way the data are collected/recorded that even partially accounts for this sudden onset?
IMO as a github-watcher, I think they changed their definition of what constitutes a sev-0 between sev-1 for the better. In particular, they had a few "sev-1"'s around the turn of the year that would be classified as sev-0's if they happened today.
Pre-4/22 GitHub sev-1 was a normal SaaS company's sev-0, imo. So I think their new system is more reflective of reality. My guess is that a few of their big customers bullied them to have more accurate SEV categorization.
To be clear, your observation that "they changed their definition of what constitutes a sev-0" is based just on your external observation of incidents and their designations, correct? I.e. they haven't officially released a statement saying they have changed their standards
I don't envy their position of having to scale that fast on something that has to be instant and real-time. As far as I know, you can't do CDN/edge caching shenanigans with a remote git repository like Google can with a YouTube video. It's gotta always be reading/writing to the latest, single source of truth.
The easy solutions like caching and read replicas don't work and you're forced to go the route of sharding or similar techniques that have much more painful tradeoffs.
I'm not sure if that's why everything keeps breaking but at that scale write-heavy workloads are never going to be easy
However, they have reported numbers along rather inconsistent dimensions. Like, historically they've focused on number of repos and users and later PR's and issues, and often catch-all terms like "contributions" which includes all of those + comments etc... but the number of commits alone (which apparently is the main culprit now?) has been mentioned very sporadically. This has made it hard to get a consistent sense of historical growth.
Without any other information, however, it is reasonable to assume that a 14x in commits is the prime candidate for instability. Especially since commits are write traffic, which is much harder to scale than read traffic. Plus every 3 - 5x increase in scale can reveal bottlenecks in your distributed systems that you never knew existed, so they probably have like 2 - 3 "generations" of bottlenecks to figure out!
Think about countless actions that have to run almost at every push and PR push! Also, remember that we were used to use external services for "actions", and they basically killed the competition by offering their own CI actions at no cost to most users.
Also, they did a lot of reworks in the last years, not necessarily for the best like the PR diff page, and probably not in the most efficient way.
At this point you would get better uptime by just self-hosting your own GitLab, Forgejo or Codeberg instance instead of dealing with Github's unreliablity.
There is no defending them with their clear neglet and carelessness of the platform.
I’m working on an open source Forgejo browser called Joui. It’s coming along nicely, and is so much snappier than GitHub in every single way.
People may have had complaints about functionality, features, commercial issues, but the thing used to at least have a decent uptime until recently.
Now it's a unit in their AI hype machine.
It's quite disappointing objectively, but I expected worse from MSFT.
The more surpassing part is that Microsoft hasn't figured out a way to manage/contain the AI-sourced traffic better so it doesn't create all this noisy neighbor problems for non-AI usage/users.
bucket1 = clients that were working just fine before (users and whatever automation they had in place) bucket2 = ai clients that contributed to, if not flat out caused, the scale problems
then slowing down/limiting the bucket2 clients while keeping the bucket1 clients rolling as-is, is both doable and keeps existing customers happy while the underlying infra gets scale/perf improvements needed to support ai clients at scale.
But a couple of years ago they were crowing about how much work they were doing to prepare for “a billion developers”. If they had actually done that then the actual load from agents should have been no problem.
I'm not sure how reliable the data is, but average uptime seems to have dipped measurably starting within a year of the aquisition, according to https://damrnelson.github.io/github-historical-uptime/
Personally, once a game I own is janked from my hands because of organizational decisions, that's the time I'll stop consider the game "in good shape", but I'm sure the people who had to buy the same game a second time still enjoy it.
This was happening to some degree pre-acquisition, but since the acquisition it's been this non-stop.
Some of it's good. The Nether and the oceans were really boring before their respective updates.
They should have called Minecraft "done" around the acquisition time and started on Minecraft 2.
The user profile / contributions and PR UX is pretty much the entire "hub" product since git is a fully separate offline app.
Is it? Seems a text description of "Make a website outlining 'How cooked GitHub' is with a modern style" to basically any LLM would produce exactly that UI and design, literally nothing of that design a human had any influence on, besides the ones selecting what training data the used LLMs was trained with.
I think most of us who've tried using LLMs for web-design can recognize that style and design at this point, regardless of model actually used.
I'd actually say what really makes an excellent engineer stick out among many great engineers, is their ability to communicate clearly and knowing what needs to be communicated vs not, basically being way better at language and communication in general, and they also understand the important of it.
Same goes with the "LLM does web design" example from before, a web designer with great communication skills in web design, will (naturally) have a better prompt for something that'll potentially could look good, compared to a web designer that isn't at good at communicating what they actually want.
3D type stuff too, it's useless outside boilerplate.
Very little spatial reasoning training, no end-user subjective reasoning inference (Google is starting to though even in unrelated chats), so it's no surprise the LLM doesn't know what you want.
Since I don't even know what I want half the time until I saw it, the subjective reasoning piece is key - that is, being able to predict what I'll want to pretty good accuracy. Then you have your agents etc.
Well, we're at least two people who care, since we were conversing about how good/bad the webdesign is, then you jumped in here :) If you don't care, why bother to reply to people who seemingly do care? What kind of conversation are you expecting here, "Yeah, do tooo"? :|
In this case, it appears more effort was put into content than presentation, which is a possibility and the creator is in the comments saying as such, but humans operate on heuristics by default. The majority of sites that look like this have been a complete waste of my time and I usually just click away at this point.
Is it though? If the page is near unreadable?
* Almost pure-black background rendering every not-pure-white colour barely readable
* Dark-grey and low saturation colours used almost everywhere, for both fonts and other coloured elements (the orange cells in the calendar are the most readable thing)
* Thin fonts - coupled with the dark grey colours this just adds to the readability issues
* Yet another incredibly long info-dump of a page
And then as far as actual information:
* Vanity metrics as the main information, that is a lot of things with no context or historical information
* A lot of aggregates and rollups that aren't that useful
No, I haven't tried Reader Mode.
It's a good demo for UI state syncing though, I'll give it that.
I originally made the core data functionality of this site for myself because I was curious what the uptime stats for each service were (I build something that heavily depends on GitHub), and to viz the distribution/severity of those incidents, again per-service, over time.
It involved a lot of back-and-forth, and is not a one-shotter; maybe closer to 40-50 shots over maybe ~10 hours of human time. A couple memorable things that made it complicated, irrespective of the UI: sneaky bugs around double-counting time for overlapping incidents, no GitHub API for incidents so you need to puppeteer-scrape the backlog of incidents to get historical data. Although, you all are right to call out that the CSS was three shots, though, and it shows :) I thought it looked so cool in ~January 2026 and now it gives me the ick, too!
For people who are curious about how much direction went into the information architecture/presentation, it was fairly substantial. I wanted a contribution graph style viz and it took many turns to get it working the way I wanted. The swimlane viz for selected-day-incident visualization was also me, because I love swimlane graphs.
I ended up sharing it with some folks and they wanted to reference it, so I put it on a website. So it's jokey for sure, but I take my jokes seriously! I'm grateful that people have feedback on how it can better functionally and visually :)
Totally, my comment was all about the styling and design literally, and is in no way a comment about the data or actual contents of the website, hope you didn't take it that way as well, as it does seem proper in that regard!
Thank you for sharing it, and even greater thank you for sharing the process behind building it, for me that's more interesting almost :)
Most part screen is taken by picture. Contrast ratio is really low. Hard to read Should they remove that useless banner, current status which is the most interesting part coud've been made visible right away.
I would call this whole thing highly un-ergonomic
Moving everything from GitHub to Forgejo and Tangled for now. These outages haven’t effected me for the past month because of this.
I plan on focusing primarily on these areas:
- mobile experience is first class, even on old/slow devices
- diff viewer is fast even on extremely large pull requests
- stacked pull request support
- user interface is modern, accessible, and theme-able with a light touch of whimsy
- search is accessible from anywhere
- opinionated keyboard shortcuts and commandk palette from day one
Many other longer term goals that I’m not mentioning here for now while the roadmap is forming.