GitHub Actions and Pages are experiencing degraded availability
githubstatus.com
githubstatus.com
We're a good year+ into the use LLMs for all major bits of software that we all rely upon and GitHub here is down to one 9 of uptime. I've been using GitHub for a _long_ time, my first commits there go back to August 2009!, and I honestly don't recall GitHub going down as much as it has in the last year.
I'm sure there's other things happening in the background, but I can not help but believe that this is directly correlated with the increase of LLM usage.
Though I would love to hear someone else's pet theory how a rock of the internet went from four+ nines of uptime to maybe one.
I have an Ops background and I strongly suspect they were given a stupid timeline for the Azure migration.
I've got to believe Microsoft have decent Ops people but the management wanted to move faster than was reasonable and screwed it up. Move one thing at a time and double check it all works and you can do a migration like this.
I always figured these problems were directly related to the migration.
* There's a huge disconnect between the actual product and what was sold to the buyer. It could be that the product was a borderline fraud, or it could be that the product was actually much better quality than the expectation on the buyer's side, but the buyer isn't interested in most of the product.
* Fear spreads in the acquired company that their product will be discontinued or reshaped into something else.
* The pay is good, probably much better than before the company was bought.
* Many internal teams end up lacking real purpose and try to insert themselves into every new internal project only to obfuscate their irrelevance.
This leads to some pathological developments, where the teams previously working on acquired products start doing a lot of useless, for-show work. Internal initiatives sprout like mushrooms after a summer rain, but they are all plagued by very broad (and mostly irrelevant) team involvement, duplication of existing products / services, and fear of being discovered. And while there are plenty of such initiatives, their role is to be a superficial distraction. In reality, everyone is afraid to touch the old code or do any sensible integration because it could lead to blanket firing of a lot of people. The middle-management behavior becomes a sort of exchange of favors, where everyone is afraid that the other can blackmail them into losing their job, and so everyone is trying to be extra nice by offering a slice of a pie to another manager.
What this leads to, in reality, is insane inertia (often despite highly shortened software release cycles), astronomic amounts of unmitigated tech. debt, opacity in communication with management, persecution of those who genuinely want to improve the system.
Based on this, my prediction is that Github will not survive. Just like Skype didn't. Somehow or other, Microsoft will find a way to replace the product with... MS Outlook with a new skin.
Seems pretty conclusive. Very similar story when they bought skype.
It got worse around April/May this year, while that graph stops at February.
I'm sorry but this made me laugh out loud. That isn't true at all, this has been going on for years. This conversation[0] from six years ago has discussion about the outages starting to become much more frequent in December 2019. It has never gotten better in that time, it's just continually degraded.
Most people work around it by self-hosting github (which has other problems but uptime aint one).
Its common knowledge that Win11 needs 32gb RAM, but im glad to hear that Microsoft is aware of it.
Looking forward to this 8gb variant :)
So they will make Windows run on 8GB because people can't purchase cheap RAM anymore, and they want to stop people migrating to Linux.
This is not directly caused by AI slop, only indirectly through RAM prices.
If you want to have a horrible time, try reading some of the runner code. It's early 2010s-style Windows-First, MSFT C# crud that has trouble not racing several threads to inconclusive status codes.
If your platform can't handle the use patterns of AI, then perhaps don't go around telling everyone to use AI for everything? It's a self-inflicted wound, you could also just not do this. Too bad Microsoft has bet its future on AI not being a giant bubble, huh?
That plus migrating clouds is insanely difficult to manage. They're almost certainly drowning in traffic and trying to keep up.
Though I'm sure some of the blame can go to internal slop code.
As someone who has spent many years working in high load environments, this is not an uncommon pattern.
You design a system and it works great. It can handle failures, load spikes, it is horizontally scalable, things are great. You think you figured it out.
And then load keeps increasing and you suddenly hit a tipping point where everything keeps failing, and you cant keep up. The things that you thought were perfectly horizontally scalable turn out to have a bottleneck you didn’t even think about until you got to a truly massive scale. Your systems suddenly don’t have the excess capacity to handle load spikes or catchup work, so suddenly any failure cascades and recovery is more and more difficult. You can’t solve the problem with additional hardware, and your perfect scalable design actually can’t scale any more.
This doesn’t have to be about GitHub using LLMs in their code to still be related to LLMs. GitHub gets a lot more commits now because of LLMs and probably get a lot more reads because of LLMs as well.
It could be that the extra usage just pushed them past one of those capacity thresholds.
If I had to guess it's because Github is sitting on top on infrastructure held up by toothpicks and duct tape
GitHub Will Prioritize Migrating to Azure Over Feature Development - https://news.ycombinator.com/item?id=45517173 - October 2025 (63 comments)
Hotmail used to run on bsd, and they tried to move it to WinNT, back in the day. Same thing. Disaster. And it was held up as a reason to never use winnt for serious work, for years.
Microsoft just always shoots themselves in the foot.
The nobody-got-fired-for-choosing-Microsoft crowd. And they don’t look few and far between.
Maybe this really depends on region, but I can't say this is true. I've experienced tons of hardware failures that have caused instances to throw weird errors, instances to randomly stop, instances to randomly disappear (along with their corresponding EBS volume).
Standard EBS volumes are only 99.8% durable. If you run a lot of instances on EBS, some of them will disappear.
GCP meanwhile has balanced zonal PDs (not even regional, zonal) at >99.999% durability.
I've never lost a disk on GCP, I've never even had an instance stop once without me telling it to stop.
Azure is absolutely a hot mess though. Had a VM a client was paying for backups on. Couldn't restore the backup because they changed generations of VM platforms too many times, literally no way to restore it. Had to get their support to eventually give me a .vhd that magically appeared in a OneDrive share a few days after asking about it.
Probably both. But it still throws a thorn into the theory that LLMs are about to replace software engineers any day now.
You'd think they could LLM-code their way out of this situation easily if LLMs were the software engineer replacement they are being marketed as.
Web search, steaming, high frequency trading, and many other systems are resource demanding, and keep scaling. Whether to handle the surge due to bots or wider adoption. But GitHub can't scale git? It isn't even git failing.
Incompetence in leadership is what makes a tech business technically unable to meet growing demand.
I don't think this is true. Take LLMs for example, you need GPUs to serve them and these are in short supply at the moment making it hard to meet demand. I don't think this is necessarily the fault of incompetent leadership.
Is there a single Google service that has ever had reliability this bad? Facebook? Instagram? TikTok? Apple? Those companies all run hugely scaled write-heavy platforms. Why is GitHub so uniquely unreliable?
Me and 1 billion of my best friends can upload 5GB 4K videos to iCloud Photos and live stream to Facebook all day long but GitHub who is owned by Microsoft the second largest cloud computing provider on the planet can’t handle 180 million users pushing code and running pipelines on VMs?
I'll give you an example: I pointed Codex at one of my repositories asking it to implement some features. Very quickly, it ballooned a 3 minute CI workflow to 30 minutes per commit, and the number of commits it started making increased 10 fold. So we're looking at a 100x increase in usage just from myself.
You factor that into the fact that all public repo get access to this same compute, basically unlimited minutes, up to 20 parallel tasks and 6 hours per request. Now because of AI, every small little weekend throw away side project is running a full dev ops shop. And the pressure never goes down because the AI just start a cron job to run these systems every single day, whether anyone is even looking at these repos including the authors, who may themselves just be bots.
And the systems are sophisticated: parallel builds across windows, mac, multiple linux distributions, multiple architectures, embedded targets, web targets, full feature matrix, the works. The default MO of these agents is to do as much testing as possible without any regard for resource constraints, and GitHub offers them unlimited resources, so it's like an addict meeting a dealer.
So in a sense, maybe being attached to Microsoft is the problem. Or at least their unlimited wallet is -- Microsoft is the enabler in all of this.
It still just sounds like a really vanilla "horizontally scaling VMs" problem, especially for yesterday's incident that was focused on pipelines.
If this is a "GitHub is giving away more capacity than it has" problem, that's easily solved by rate limiting and queuing.
This is GitHub's explanation:
> During the incident, some Actions Runner Controller (ARC) runner pods became stuck in an idle state. Affected users can delete those pods using kubectl or redeploy their Actions Runner Controller application. ARC will automatically create replacement runners.
> The next releases of Actions Runner and Actions Runner Controller will include an automatic recovery mechanism, preventing the need for these manual steps in the future.
That sounds a lot more like a major architectural flaw and a really obvious oversight.
https://www.reddit.com/r/kubernetes/comments/f2bcsu/parodyfu...
With GitHub, we threw that away. It's centralized on steroids. Now we're all depending on one platform that has to be massively scaled to deal with the massive load of serving almost every major software project on the planet. And when it is down, we all notice.
I see this repeated a lot and IMO it's simply not true.
The main decentralized advantage of Git is that you can continue to do VCS operations without network access or access to the remote host.
Most of what Github does is managing collaboration. The only thing we've "thrown away" by using Github is the email based workflows or directly pushing git branches to various hosts. But unless you're gonna have team members SSH into each other's machines you'd still need a central repo somewhere.
My only guess is there's some circular dependency they want to avoid by keeping Okta.
Jokes aside, yes, quite correct. Even frontier models like Fable have very obvious limits.
I'm really kind of surprised they let us do that - like, why didn't they just have you upload the binaries after building on your local machine?
You can do that already with GH Releases. Actions is if you want CI/CD managed by GitHub. And you can also use your own machines via self-hosted runners.
So, from the sales point of view, Github Actions was, at the minimum, an alright idea. Not brilliant, but quite obvious and expected. From the engineering standpoint, however, this is a disaster on many levels. But, that never stopped Microsoft before. They don't try to win the market by making an objectively better product, their tactics are and always have been to make a product that can claim (with an asterisk) to be able to do a lot of things the customer wanted only to discover afterwards that those promises were phony.
With what to show for it? If GH did 10x in volume/git commits, it's all LLM sloppy-pasta. Where's the 10x productivity? Where are the amazing apps?
I think it is a natural progression in technology. Failure rates were high for initial aircraft designs and safety improved over time, for instance.
> Yup, platform activity is surging. There were 1 billion commits in 2025. Now, it's 275 million per week, on pace for 14 billion this year if growth remains linear (spoiler: it won't.) GitHub Actions has grown from 500M minutes/week in 2023 to 1B minutes/week in 2025, and now 2.1B minutes so far this week. So we're pushing incredibly hard on more CPUs, scaling services, and strengthening GitHub’s core features. And as a fine purveyor of hand-crafted shit code for many years, I'm not gonna weigh in on that.
x.com/kdaigle/status/2040164759836778878
These are folks that routinely make it a point to press on system design and scalability during interviews.
Now they suddenly can’t scale or design systems but we should accept that?
For like 10 years the only feature development they did was by stealing ideas from GitLab. I wouldn't be shocked if little if none engineering discipline has took place at all during this time if it's this brittle to frequent change.
Guessing it was mostly held together with duct tape and poorly written tests/monitoring systems if any other corporate driven software.
We're really going to find out over the next few years which businesses have good practices or not.
Even my own systems at home are burgeoning under the load of the my more ambitious hobby projects I'm doing for fun. I don't envy "we had to build five more datacenters to keep up with demand" class problems just from a how many people have to sign off on them perspective let alone the technical difficulty of doing so
Why do we keep giving excuses to poor engineering disciplines + poor management? This problem is entirely GitHub's making and acting like it's some unplanned natural disaster is low key pathetic.
Next: proficient engineering managers, not just people who immediately bend a knee to the middle manager.
Oh wait, I went into sci-fi territory again. My bad.
In any case, the number I'm focusing on here is the 300k cores part (x2 if you're counting in vcores). It does not seem like too much to ask for (significantly) more than that at github scale. It doesn't feel like a hardware issue, is what I'm getting at.
1MW ~ 6700 NYC citizens' residential usage, apparently, lol. I don't know exactly why, but that citizen number seemed a surprisingly large (well, probably because I've played with approximately-MW lasers).
No-one's stopping you hosting your own projects.
selfhosted runners do not go down when our CI is down, its operates separately.
Their MO was to court an executive and sell second-rate tools to them before the people who had to use them had a chance to say anything. It doesn't matter how much evidence you can provide to the contrary, once the million dollar deal is signed, you are going to be tasked with finding reasons to say that your executive was shrewd for buying this pile of junk and unfulfilled promises, and not an insane idiot sucking away your job satisfaction as fast as they can.
They did a lot of deals based on how their products would have features their competitors already have 'soon' when they haven't even started them, and a long track record of taking 3 major releases to get from something to good, and then breaking everything again by doing a 4th major release that re-imagined everything and made it horrible again.
I'm not going to claim that Apple was or is a panacea. Apple doesn't use vaporware which is big, and their Cycle of Awful is 2 releases instead of 3. You could afford to skip 1 waiting for the next even-numbered version, instead of being 2 versions behind and getting pressed to upgrade.
Windows NT was one of them. Up to Windows XP the products were pretty solid and each had visible improvement against the previous one.
Their language products were/are still solid IMO. Maybe Visual Studio is sluggish, but we can still use an older version if we want. Plus they put a lot of effort optimizing VSCode, too.
Even back in the MS-DOS/16-bit Windows days, when things broke down quite easily, I think they still provide the best bang for individual users and developers. There was no competitors who could provide so much value back then.
Disagree. I was admining NT4.0 boxes back then and it was
1. very slow
2. constantly leaking that required weekly scheduled reboots
3. security wise it was nightmare even by that times standard
For the last point what was the golden standard back in the mid-late 90s, if we don't include mainframe/minicomputers? Was it Solaris or BSD?
Now that I think about it, I feel sad that I do not have the technical prowess to compare operating systems :/
My WSL instances end up getting overwritten when I reinstall Windows so I thought that was just a general OS thing...
Every bare-metal Linux system I have in my home is upgraded in-place. My oldest is roughly twenty years old. My newest is around four years old.
Back then the other options for home computers are just too technical for normies. I tried to install Linux back in early 2000 and I couldn't even get through the install process. Skill issue for sure, but if I couldn't do it back then I'm sure none of my class mates could.
Windows 95 was a brilliant hack that allowed 32-bit GUI programs to run on hardware with 2Mb of RAM (Windows had 4Mb of RAM as the official minimum, but you could run it on 2Mb) while preserving compatibility with the majority of DOS software.
If the scheduling was self hosted it would be inexcusable but you can always just connect whatever you want to webhooks.
They have a strong motivation (self preservation) to continue to misunderstand the problem. If they did what is best for us, then we could avoid a substantial fraction of all GH subscriptions by using a FOSS tool to hit the Pareto frontier by replicating just enough GH services to watch commits and PRs.
I don't disagree that it's obvious they've got problems but I'm just saying it's obvious to me the part that falls over (the scheduling of jobs) and why that would impact self hosted runners, which do no scheduling but depend on it to function.
As for 'just a message queue with some database updates and sharding that's easy to reason about'... Here's a job scheduling problem as an example: imagine you schedule a job, and there's no runner available. How do you disambiguate between no runners available because you've reached capacity, runners not being available because they're on a real network with faulty connections, and runners not being available because of a faulty rollout of internal updates?
A simple message queue for job scheduling is fine if you own everything and can deal with the operational overhead of identifying those cases by hand, but Github can't do that.
Sure, but the part that actually schedules where a 'job' gets run is based on a relatively simplistic tag system. Reading the yaml and plopping some job metadata into a queue-like system isn't where I would expect their issues to be, but at their scale I'm sure everything becomes fragile and inscrutable.
> imagine you schedule a job, and there's no runner available. How do you disambiguate between no runners available because you've reached capacity, runners not being available because they're on a real network with faulty connections, and runners not being available because of a faulty rollout of internal updates?
You don't need to. GitHub Actions runners, and most CI runners that I've interacted with appear to have a pull-based model where they ask for work that matches their declared tags/shape (usually platform/runtime/OS/etc.). This probably amounts to a database query, but who knows.
> A simple message queue for job scheduling is fine if you own everything and can deal with the operational overhead of identifying those cases by hand, but Github can't do that.
I highly doubt it's a simple message queue. My issue is git repos and their CI infrastructure have very low coupling to other repos or entities in most circumstances, at least conceptually, so parts of the system (ie. regions, shards, etc.) should be able to function even when others are down (ie. it shouldn't break for everyone). There's clearly centralization and coupling that isn't obvious from an outside perspective, which sorta tells me it's incidental, but that's a guess.
This is no surprise given standard Microsoft operating procedure - https://news.ycombinator.com/item?id=47616242
I'm certainly not surprised that they tried to sneak in billing self-hosted runner minutes at some ratio to potentially recoup the costs here.
0: https://docs.github.com/en/enterprise-server@3.21/admin/mana...
Edit: to be clear, I mean the part where they maintain state on their end. I have no idea if they do that with another runner instance, that seems unlikely.
Sitting on Github these days is the same as sticking to twitter a decade ago, expect next mecha hitler, I suppose.
9% availability would be an uptime of ~33 days a year, I think at that point, we're pushing the semantics of "available" if the service is down the entire year except one month on average.
I thought they should rename to Anglia Busways and have bus replacement trains instead.
I have sympathy for the on-call team trying to resolve it, most of us have been there done that.
But seems something is systematically going wrong at GH
Outages happen, but this many outages so close together, and so many of them so major/long lasting, something is systematically wrong for sure. It's been seriously hamstringing our ability to ship code at my company.
What is happening at GH?
Rate of change trying to keep up with new challengers? Over-reliance on AI? Engineers trying to debug slop?
They're at 93.91% uptime over the past 90 days, according to https://mrshu.github.io/github-statuses/ , and that doesn't even include today's outage yet.
A glorious one nine of reliability.
I don’t love GitHub, but that number is a little misleading… That’s the intersection uptime of all GitHub services, most of which I (and most users) do not care about; code spaces, copilot, packages etc…
When you take out those uptimes, it becomes a lot higher. I’ll admit there seems to be a lot more incidents than usual though…
In large part, the move from AWS to Azure. Azure's just bad.
Azure remains too bad to move Github to, but the relevant executives have pretty clearly decided that that's no longer a good reason to delay the move.
It can be both things. I've worked professionally with AWS, Azure, and GCP, and Azure is just really flaky and unreliable. It's easily the worst of the three.
Yes, we call it: Microslop.
We have many agents per employee working in parallel pushing way more commits than was humanly possible before AI, triggering GitHub actions a lot more than the workflows were built for, causing Actions costs to escalate (they really aren't cheap if you compare to hosting it yourself), meanwhile working with YAML workflows is just a pain, and just writing code would be so much more fun and AI compatible[1].
At the same time, GitHub has about ~3 different PR review UIs? And they're all half-bad? Any decently sized PR triggers their "optimized for large PRs" UI which jumps around randomly in my experience. If you don't get that UI and keep the scrolling one (there's an old and a new one btw) then god forbid you click a line number because at some point your browser will randomly scroll back to that line and it won't unstick. Now Linear[2] (and others) is replacing the PR review experience for the agentic era.
I'd love to see a solid AI first Git + CI + reviews.
[1] Cloudflare CI https://blog.cloudflare.com/ci-workflows/
[2] Linear PR reviews https://linear.app/changelog/2025-01-23-pull-request-reviews
Source: me
Now I'm stuck twiddling my thumbs with PR checks stuck/failing...
This is annoying and I'm here because it's down. But it would have to be far worse to come close to actually being worth changing.
If I could right now:
1. go sign-up elsewhere 2. Log into GitHub and point Actions to that new host 3. All my actions files immediately worked without question
I'd probably give it a spin and make a wiki page explaining how to swap back and forth. No meetings. No design issues. No scheduling. Just a flip switch on who to pay for computers.
I think people were so excited to move away from jenkins to something 'managed' just because of how much a dinosaur jenkins is and how much a pain in the ass it is to upgrade it... but now we are seeing how managed can bite you in the ass if the manager is incompetent.
This is multiple times this month that this has been a problem.
Has GitHub completed it's internal migration to Azure yet? Or is it still ongoing? None of our devs want to switch away from GH, but we will have to at this point.
Step one - migrate my build workflows to Docker.
My Github actions are now basically: "checkout / set env vars from secrets / docker-compose builder run make".
I used large machine runners to run full the Docker (Podman actually) on Github first to avoid dealing with docker-in-docker complications. This step also provided some very nice robustness advantages, as I can now trigger deployments from my laptop if needed.
Step two:
Migrate to self-hosted runners. I used my former homelab server to set up a build machine. It has 16Tb of fast NVMe SSDs and thanks to Podman container layer caching, my entire lint workflow now takes 30 seconds. Faster than just one "npm install" on Github before.
And Github's self-hosted runners are actually surprisingly easy to set up and use. They are also somewhat more robust.
Step three:
Swap Github for something else.
But for people who either don't pay anything at all or phenomenal amount one 9 of up time is all you need.
If you're actually trying to run a business I guess you can call and gitlab and get an Enterprise contract
Nobody cares about ATProto or whether your commits are a damn NFT or some bs just literally improve upon the experience.
That’s it.
It’s as if no company is focusing on the product experience or anybody’s experience anymore. It’s all ooo look what I got I got this I can do that too me me me but nobody will ever buy that.
Say what you want about huge companies like Microsoft or Walmart but they spend a lot of energy understanding the human experience to sell products and less on their own perceived self-aggrandizement.
GitHub is the best version control online and it’s not even close.
Github is just the laziest default. It's not _terrible_, but it's also not great.
Also GitHub just sucks.
This is a man who's spent a significant portion of every day for the last 15 years on GitHub.
>My personal projects and other work will remain on GitHub for now.
People always talk about "leaving GitHub" but it's all talk.
I'd love to know what the most common root causes for these outages are.
I'm not sure why this particular industry is so abysmal at making things even semi-reliable after decades of research, educated workforces, and loads of cash.
As a side note, I started getting these symptoms (actions staying queued forever or not running at all) at 5pm PDT yesterday, intermittently. And definitely full outage by 8:30pm PDT. So for sure full outage for 7 hours and counting, even if you use other runners, and I strongly believe partial outage for ~15.5+ hrs before that, even if it hasn't been acknowledged by GitHub yet.
You should check out Avrea — they’re building around exactly this problem, with a lot more control and reliability than relying entirely on GitHub Actions. New player, but insanely fast. You can even move your existing runners over with one line, so pretty easy to try.
If you implemented a tool and it worked one time for your presentation to management that's all that mattered.
The actual company employees using it downstream in prod basically had to constantly QA the alpha software they were forced to use and the authors of the tool were hard to track down if they even still worked there. And if you did find the author or the team they would be very resistant to admitting there was an issue because it LOOKED bad.
So many tools I used were fragile and buggy, it was clear the authors just presented the happy path to management to get the note added to their promotion packet and the rest of the company just had to deal with the fallout.
My team implemented this product that the entire company used that was broken and buggy as hell but they kept presenting the product to management as this amazing product and nothing was ever done about how broken it was. One of my team members came from Apple and said Apple's tool to do the same thing was much better. The tool my team worked on was a well known pain point amongst the rank and file but management was very detached from the rank and file, which I guess ultimately was the primary problem.
If Github is having the same issues I feel for them.
Aug 06, 2026 - 16:27 UTC - Update - Pages is experiencing degraded performance. We are continuing to investigate.
Aug 06, 2026 - 16:19 UTC - Update - Pages is operating normally.
Like, everyone talks about how you can just prompt a new CRM instead of paying for Salesforce.
But in this case, I raced with Fable and Sol and had time to build a full featured, fully functional CI/CD workflow system that has most of the features of GitHub actions (not the ecosystem obviously) and costs less than the GitHub runners per minute, while being able to scale to millions of workflows... And I had it working end to end before the outage was resolved and was already running my apps deploys through it.
It would take a little longer to do in a company with red tape and more investments in the GitHub ecosystem, and of course it only replaced actions. But it works.
There is no moat.
I should eat lunch.
If the frontier AI used GitLab as a default I'm sure they'd be the ones suffering now.
Especially troublesome in the middle of trying to fix a high score security vulnerability when the release vehicle is Github.
Also, doesn't even have RAG offering.
Seems like the only reliable way to run GHA jobs is to not use their runners. Hope they at least didn’t break self-hosted runners operations
> Customers using self-hosted runners may see errors or rate limiting when runners register.
Maybe they're hosting in us-east-1 though :)
Wildly exaggerated. I use github almost daily and can't remember when the last downtime was.
Self-hosting a service like GitHub that operates at GitHub scale is difficult.
Self-hosting a service like GitHub that operates at the typical small/medium company's scale is trivial.
A single machine (with separate runners for CI) will cover many companies' needs. It being a single machine eliminates a lot of the complexity and failure modes associated with a distributed system and makes backups/restores/maintenance easy.
Even self-hosted runners are impacted.... How can that be?
The cost of this globally has got to be in the hundreds of millions to companies that use CI/CD through GitHub Actions. What if prod is broken and GitHub actions is stalling the deployment of your hotfix? What if this makes your organization miss and SLA and diminish user trust? What if this makes you miss a release that you were contractually obligated to meet? This is happening during peak dev hours on a Thursday (not that it would be acceptable at any other time).
I don't understand how a service this critical to the global technical infrastructure can fail like this at all, let alone for more than a few hours. Like where's the backup generator for crises like these? You can't even use self-hosted runners? WTF? Like how can you not bring your own backup in a crisis event like this?
Not that Microsoft has a good reputation, but holy moly, you'd think they would prepare from something inevitable like this.
It's hard to draw a direct analogy there, but I feel like it echoes the same sentiment.
A soapbox I have is that GHA workflows are scripts that could run on your machine without any of the YAML stuff. Who gives a flying about the DAG or the logs? Which, by the way, if you're willing to walk to the milk store to buy your milk, could be recreated in a much more testable and maintainable way without any of the YAML bs that GHA prescribes....
But DAGs are pretty, and logstreams showing up in a browser application instill trust (for reasons that fly far above the head of yours truly). So people go for that. Pretty DAG, nice logstream; therefore, deliver my milk. All of a sudden.... The CI/CD platform is having its merry way with your SLAs, contract abidements, and hotfix deployments.
What a time to be alive.
Not to rail on the South Park thing, but the blast radius of this issue also reminds me of the episode where the internet dried up.
If this bs with GitHub continues, Parker/Stone will have to make a GitHub episode. How seen would we all feel if that happened?
The fix is merged, but won't deploy... it's been hours
Thankfully it's a batch job, and isn't interrupting production ATM
There's always the escape hatch of running you GHA workflows locally, but unfortunately, despite the existence of packages like `act`, there is no way to fully recreate the GHA runtime locally. Tons of the special YAML syntax just can't (more accurately, "just doesn't") get interpreted by those local actions runners.
We never went this route, but at my old org, I always advocated for considering GHA to be wrapper around a single bash script (or whatever script you want to run), as a means of completely breaking out of the GHA hellscape that is programming in YAML, who's turing-completeness is pretty dubious.
Unless you have things set up this way, you (the client of GitHub) would have to completely redesign your CI on the fly, run it locally, and then figure out how to get the D compliment of the I to work in a way that is auditable. Fat chance for most teams I bet.
Thank god you're dealing with a batch scenario. Silver lining for sure. Still, embrace the anger.
What makes my blood boil is that there's millions of DEVs literally crying at the moment worrying about how GitHub's failure to be responsible will put their jobs in jeopardy.
And fingers crossed for you my friend. We're at 5+ hours at the time of this writing.... You're batch job may still have a chance!!!
I was complaining about how I should have used AI instead of manual brainwork for this instead but turns out that might have been the problem.
Would you know it, they're actually UP to the ... One-nines. 93.40%
latest server i set up is simply a bare repo + hooks to make a local deploy after running tests and shit
super easy to set up having ai do it, zero dependencies, deploy is still 'push it to the main'
i have several remotes for backups and stuff
I think we need to ask for some money back, this is BONKERS.
Any more of this, and we'll go full Gitlab, like we should have years ago!
Maybe people are actually pretty tolerant of downtime if other things about the product or service are done right
It was before it became a news and a trend in X.
Really frustrating experience.
Now that the calculus has changed, I think we should start asking ourselves whether anyone has the right to unlimited free public repositories.
We can all agree that trying to attribute any single metric to repo quality is subject to Goodhart's Law[1] (e.g. what happened to GitHub stars) yet we can also agree that the Linux Project is a much more vital repository to keep public infrastructure availble for than say, someone's vibe-coded to-do app. Is it impossible then to make any quantitative distinction between these two? We can't have both an uncontrollable firehose of AI slop and unlimited free public storage/compute at the ready for it. I say we actually need to decide what software is worthy of using up these resources for free. We already have companies like JetBrains[2] using dynamic pricing for their products. It's time Github do the same. If you want a hosted VCS for vibe coded projects that aren't being used or depended upon widely, then host the repo yourself.
There is no better time to self-host.
Or has Microsoft made sure (through Windows licensing terms and pricing) that it's not possible to compete with their own CI offering?
Edit: CircleCI seems to offer 750 minutes/month (whereas GitHub offers 1000 minutes/month).
Honestly I dont think I've seen a tool I used regularly with such a moat lose it simply because they cant keep the service up. Also its hard to not see a pretty strong correlation between a bunch of these big companies doubling down on AI and just having their service go to s%#$. We've had similar issues with Digital Ocean recently, which is pretty lined up with them adding a bunch of "inference" services and rebranding the site to add all the "agent" marketing slop. Looking to move to Hertzner when the time allows.
Is the AI slop that bad? Culture change?
From the GitHub COO on April 3rd:
Platform activity is surging. There were 1 billion commits in 2025.
Now, it's 275 million per week, on pace for 14 billion this year if
growth remains linear (spoiler: it won't.)
GitHub Actions has grown from 500M minutes/week in 2023 to 1B minutes/week
in 2025, and now 2.1B minutes so far this week.
So we're pushing incredibly hard on more CPUs, scaling services, and
strengthening GitHub’s core features.After 6 years of this nonsense of "centralizing everything on GitHub", it is not a good idea at all.
You might as well self host like I said before [0].
Once this goes in, I'd expect to see 89%, which is zero nines. (I'd like to say, "a new low!", but sadly we've had this before)
It seems to manifest in many places: AI slop, Github outages, recently Miscoroft sent me an email demanding that I subscribe to 365, in order for MS Office (which I already paid for) to continue to work. I simply moved on to the other, free, provider.
Perhaps, it's time to evaluate the use of the Microsoft products?
point clanker to forge.smol. ai/llms.txt
for now its just a fast agent native git remote and u can check docs for the extras.
Among with assorted thoughts on AI or new model point release announcements.
I'm genuinely curious what changed, what their processes are and how they internally think about their reputation being in the gutter with all these incidents.
but i always guessed that maybe half of us seeing how many tokens we can spend for a year straight maybe stressed tested their platform for a year straight