GitHub was down
githubstatus.com
githubstatus.com
At my job, if something we go wrong... management just tells us roll it back. That always fixes the problem, right? :P
If rolling back works for a simple system, why wouldn't it work for a complex one?
Because one cannot step into the same river twice. Heraclitus would probably make a good engineering manager.
edit - from wikipedia: In the field of psychology, the Dunning–Kruger effect is a cognitive bias in which people assess their cognitive ability as greater than it is.
I interpreted GP as saying the aforementioned management has just enough cursory knowledge to want to apply the same hammer that worked on a simple system to that on a complex system, but not enough knowledge to realize the unknown-unknowns that they aren't even aware of.
For context, there were four (public) incidents this week:
• Incident on Feb 27, 14:31 until 18:54 UTC - https://www.githubstatus.com/incidents/q07bfjh7jf1t
• Incident on Feb 25, 16:36 until 18:48 UTC - https://www.githubstatus.com/incidents/xp2qc958g4wt
• Incident on Feb 20, 21:31 until 22:16 UTC - https://www.githubstatus.com/incidents/bd29l6zgr43g
• Incident on Feb 19, 15:17 until 16:09 UTC - https://www.githubstatus.com/incidents/fxbbtd7mhz1c
I'm sure the pagination might better for performance, but it's terrible UI.
Ctrl+F on a list of 10000 entries is far easier than clicking through 400 ajaxy pages and trying to figure out some custom and buggy filtering system that probably doesn't allow regex.
Past 10000 records most sites probably ought to just let you export in something bigquery compatible anyway - Regular Joe isn't going to have more than 10000 of anything, and anyone who does can learn how to use proper data tools.
Did you miss the part where I noted Github’s lists fail to load (let alone render) long before that point?
Edit:
Realized *.github.com takes you to your .github.io sites.
Wait, what? On iOS Safari I can only see repos in desktop mode now (except the issue tracker which is responsive anyway). Which is a good thing. Not sure why you have the exact opposite experience?
(I do vaguely recall being asked if I would prefer desktop mode on my phone a while back, and I said yes.)
I've been doing on-going client work for someone using BitBucket and for weeks it feels like every other day has an outage related to their pipelines (CI) feature (the thing I happen to be working on).
It was constant banners about service disruption. There's a lot of UI outage related issues too, like the pipelines page starting to show a new build but never updating any of the progress until you reload the page -- which sounds like some type of API outage somewhere. I'm not sure if that gets reported as an outage but it makes using the platform not fun.
I imagine that there are stability issues that any provider will have to deal with as they scale to account for the masses.
Plus of course the service was so damn slow that using it was a daily pain.
Enterprise = 99.95% (quarterly)
https://help.github.com/en/github/site-policy/github-enterpr...
They're having a bad February but January was good. We will see what March has in store
> Our Uptime calculation is based on the percentage of successful requests we serve through our web, API, and Git client interfaces.
Just curious, how do they measure this? What is the actual calculation?
They obviously don't have beacons on the client side, I wonder if it's based on statistics, at this time of the day on a Tuesday we should be getting x requests but are getting only x/n.
Says who?
Not answering this directly, but the paper Meaningful Availability [0] released recently really changed my opinion on how to calculate and visualize availability. There's a discussion on HN as well [1].
[0]: https://www.usenix.org/system/files/nsdi20spring_hauer_prepu... [1]: https://news.ycombinator.com/item?id=22424173
[Edit: To the point: a high rate of randomly-timed failures is a kind of degraded experience, but not as critical as blocky patches of downtime. A 1% rate of randomly-timed failures is much much much preferred than having the service go out three straight days every February.]
Also: uptime is not the same as "customer delight". It's all about time.
Do you think that's an accurate generalization for all software and business contexts? I think a novel insight about the paper is that windowed user uptime is able to visualize the differences. (See Figure 20 from the paper.)
Whereas of some GitHub request fails and I retry it's a minor annoyance, but in most cases I won't even know whether that was GitHub's flaw, my local system or some networking in between.
So when a customer finds a broken service it is in their financial best interest to repeatedly hammer the broken service and drive down the uptime calculation to trigger their rebate.
Just an observation, not a suggestion. I’d fire any customer I found doing this.
-- Chris Pinkham
-- Jesse Pinkman
-- Walter White
The issues, comments, PRs, wikis etc... that we all came to depend on aren't.
That said, I think it's a bit weird that they don't store the data of the services around the code itself in git, like they do with e.g. sites. That way you'd have an `issues` branch that you could still access if github is down.
But that would probably pave the way for easy migrations away from Github.
bingo
Also you have to deal with sysadmin error, i know us sysadmins are practically perfect in every way, but occasionally we make mistakes....big mistakes. ;)
So redirecting might not always be possible.....
People's comments are meaningless, you can look at historical GitHub up-time and see that it hasn't changed meaningfully.
"And it hasn't happened yet"
Ah yes, now you have to backtrack from: it happened! to... no wait I promise it will happen! Based on... what? The fact that some acquisitions don't go well?
This is all pure speculation with no substantiation.
I recommend learning about confirmation bias.
I went back from the time of Microsoft's acquisition, and that status seems heavily underreported. At least when I checked now, it was all green, green. That does not reflect my experience.
Danger was already in freefall by the time of the acquisition. You can't blame Microsoft for that.
Why do you say that? I thought it was just renamed App-V and it's still going strong to this day. https://docs.microsoft.com/en-us/windows/application-managem...
1. Nothing is going to change. We bought this company because we love it
2. We need to show a higher profit for this quarter, cut all expenses for every subsidiary by 15% by Friday
3. Cut back on training, R&D, and support teams. they are a huge cost center
4. Bunch of employees leave after retention bonuses, replaced with MUCH cheaper labor
5. Need to show better on our next quarterly filing, slightly increase prices
6. Through attrition, replace more good people with cheap drones, until nobody knows WHY things are the way they are.
7. More increased prices, and way increased support contracts
8. Wonder why we have lost all this marketshare. Look at Company X, they are doing great, lets buy them.Nokia... I don’t know about that one. Change to smartphones hit every “old” phone brand... Ericsson did not survive, Siemens did not survive, Alcatel is just a brand now, even Sony has a hard time... Nokia would probably die no matter what. All the Maemo/linux based OSes (that kept changing names all the time) were nice, but so was Palm’s WebOS...
$ git push
Enumerating objects: 26, done.
Counting objects: 100% (26/26), done.
Delta compression using up to 8 threads
Compressing objects: 100% (15/15), done.
Writing objects: 100% (15/15), 1.49 KiB | 1.49 MiB/s, done.
Total 15 (delta 12), reused 0 (delta 0)
remote: Resolving deltas: 100% (12/12), completed with 10 local objects.
remote: Internal Server Error
To git+ssh://github.com/<redacted>/<redacted>
! [remote failure] wip -> wip (remote failed to report status)
error: failed to push some refs to 'git+ssh://git@github.com/<redacted>/<redacted>'There really is no good reason bee so dependent on GitHub.
Could be the IO? I remember colleagues working on getting stuff running on azure and they experienced horrible IO latency, as well as very low throughput for lots of small IO (aka unix-style software).
Was a few years back so it might have improved since, but if those things are still non-optimal and github is built with a unix-style vision of tons of small IO access…
This specific outage taught me that github apparently stores git repos on-disk which I was not expecting though (because the API access complained it could not delete repos until they'd fully backed to disk or something).
I know I would if we had a month with only 2 9s of uptime.
https://i.imgur.com/vy7onDT.png https://www.githubstatus.com/uptime
didn't notice there was a selector for which type of downtime on that page that didn't default to all.
I'm not sure how that results in 99.98% uptime on the other tab.
Number of outages and duration total are much better metrics.
"Waa, that was so archaic the way we used to do things!"
/s
Appreciate the work.
There are several ways to work around this, but none are really satisfying.
Translation: our shitty software update practices are now affecting Github, not just Windows!
If anyone from Microsoft is reading this, why is your company so incompetent at software updates in the past few years?
I'd wage that Microsoft practices had nothing to do with this, but I'll wait for RCA.