GitHub Outage
github.com
github.com
bower jquery#1.11.3 not-cached git://github.com/jquery/jquery-dist.git#1.11.3
bower jquery#1.11.3 resolve git://github.com/jquery/jquery-dist.git#1.11.3
bower foundation#~5.5.2 cached git://github.com/zurb/bower-foundation.git#5.5.3
bower foundation#~5.5.2 validate 5.5.3 against git://github.com/zurb/bower-foundation.git#~5.5.2
bower ember#^2.3.0 ECMDERR Failed to execute "git ls-remote --tags --heads git://github.com/components/ember.git", exit code of #128
fatal: remote error:
Mid bower install. Rly srry guys!!! =[Shit happens.
I was able to painstakingly rebuild the server after 9 hours without anyone noticing. To this day one of my biggest fuck ups and prouder accomplishments.
Recovered all the data by using open file handles in /proc/.
Not a fun two hours.
Shit happens. Live and learn.
The URL "git://github.com/components/ember.git" [1] suggests this is an internal GitHub Bower build log, but your post history doesn't mention anything about GitHub (let alone whether you work there), so I'm not 100.00% sure.
[1] https://webcache.googleusercontent.com/search?q=cache:3e00jl...
Assuming this is, in fact, a GH Bower log, the first thing that came to mind was that this architecture isn't (and possibly should be) using a dual-silo approach: when you upgrade, the upgrade gets loaded into a new blank namespace/environment, tested, and if it worked (passes CI test coverage or something like that), the main entry point is switched to the new environment (maybe with a web server restart or config rehash) and the old environment gets purged (possibly after a trial period). The current stack looks quite akin to "click this button to flash the new firmware and DO NOT UNPLUG your device or you'll brick it."
But then I realized... wait. You guys have like... isn't it like, a few dozen RoR worker boxes? Was this crash on the inbound router or something? xD
It's all good though; consider this "curious criticism" - like constructive criticism, but with extra sympathy. And hey, GH's never broken on my watch before (not that I need it atm... hopefully); this is interesting =P
Source: I work at GitHub.
So, you're saying you broke it.
Everybody, over here, gang up on this guy.
> I hope github does make public a official reason tomorrow
I do too, but, mostly I find post-mortems quite interesting to read.
If your services are at some 95% uptime or lower, you're doing something (serveral things) very wrong and its probably not that interesting to me.
Getting from 95% -> 99% you probably did some interesting things there.
Going into multiple .9s beyond that, you're likely doing quite a bit of interesting stuff but what I find more interesting is where you went wrong. Figuring out not just where you went wrong, but, what your incorrect assumptions were and WHY they were wrong. "We believed X could never fail because of Y, and even if X did fail, it would not cause production impact because of Z!"
But now my "technical breakdown info" box has no tidbits in it. :P
I'm glad you're back up now (sortakinda - what sort of traffic are you sustaining right now? :D), but a rough idea of what asploded would be really cool to know about.
Speaking of which, I'd like to take a moment to make a strong point about the fact that disaster-recovery situations don't get blogged about enough. Vague "we fixed it" datapoints get buried in status update logs like it's something to hide and hope nobody brings up.
In situations like these, the only constructive perspective is for everyone to accept that something went horribly wrong and not make a fuss about it, and if such a mentality can be established, this creates an environment within which we can share technical breakdowns of "we found ourselves in XYZ position and then we did these thirty highly specific things in heroically record time to be up and running again", and I think sharing this type of info would potentially be more educational than setup tutorials or the "we switched to X and it improved Y by 1400%" type things the Net's full of. Sure, you'd have to generalize and probably give a lot of backstory about infrastructure, but it's becoming trendy (in a sense) for companies to describe their operations in precisely this way, so it's not completely nonviable.
(Note my use of the angle of "we found ourselves in XYZ position" - maybe a small highlight of what led up to the disaster would be included (worth considering if the information would be educational), maybe not. In a blog context, moderating comments to keep the discussion on-track and constructive may be necessary, but IMO would be worth it.)
The repo is at ... https://github.com/cjb/GitTorrent, so just clone that and ... oh ...
First we connect to GitHub to find out what the latest
revision for this repository is, so that we know what we
want to get. GitHub tells us it’s 5fbfea8de... Then we
go out to the GitTorrent network.
So yeah, wouldn't have saved you.http://gittorrent.org/ and `git clone git://gittorrent.org/gittorrent`
If you are desperate, you can also call another developer and add another remote pointing somewhere else. So Git is a win anyway.
Try that with a "SubversionHub"...
So I had to install Gems from RubyGems, not that big of a deal, and I had to look them up on RubyGems since Google gives me GitHub first, but that's OK too, but then all the documentation seems to be on.... Github. Except not, it's also on RDoc. (Though RDoc kinda sucks, compared to GitHub...)
Pretty big win imho, Rails got a few more points with that. :D
Now I'm wondering how NodeJS is faring at this...
Yeah, but they'd be wrong about that purpose.
Distributed systems used by people (eg. email, BitTorrent) always have their major hubs (Gmail, The Pirate Bay). That's understandable: no product reaches critical mass without a main stream. The strength of a distributed system isn't that it has no points of failure: it's that, in the event of a significant failure in an established node (eg. TPB's downtime at the end of 2014), the community can retarget around a new solution (eg. KickassTorrents) at the point in which the inconvenience of the downtime outweighs the inconvenience of switching habits, without a significant dip in service associated with the switch.
In contrast, a truly centralized system like BitKeeper would just outright block progress if the central node were to deny access (which was, of course, the purpose that led to devs like Linus Torvalds changing their focus for a few months, so they they could scale back to full kernel production as they constructed a workable alternative, in Git).
Bower certainly isn't the only thing that does it. It's a trend that's becoming more and more common.
And sitting in a ruby talk when the rubygems compromise happened just after everyone had been told to go upgrade a gem because of flaws in a specific gem.
It's not unique, but it is terrible.
Oh wait, we have.
It is obviously gonna be hosted on GitHub. Oh, the irony...
It looks like webhooks are wedged.. no?
$ git push origin master
Counting objects: 5, done.
Delta compression using up to 8 threads.
Compressing objects: 100% (5/5), done.
Writing objects: 100% (5/5), 433 bytes | 0 bytes/s, done.
Total 5 (delta 4), reused 0 (delta 0)
remote: Unexpected system error after push was received.
remote: These changes may not be reflected on github.com!
remote: Your unique error code: 4fce1b2367b5304dd3761538b8fd0c23
To git@github.com:myrepo/myrepo.git
a62b7f1..e88431a master -> master
$ git push origin master
Everything up-to-date
Note: Values are fake, but message is real. $ git clone git@github.com:influxdata/influxdb-ios.git
Cloning into 'influxdb-ios'...
remote: Counting objects: 10, done.
remote: Compressing objects: 100% (8/8), done.
remote: Total 10 (delta 1), reused 10 (delta 1), pack-reused 0
Receiving objects: 100% (10/10), done.
Resolving deltas: 100% (1/1), done.
Checking connectivity... done.
Even though this doesn't: $ wget https://github.com/influxdata/influxdb-ios
--2016-01-27 20:03:39-- https://github.com/influxdata/influxdb-ios
Resolving github.com... 192.30.252.131
Connecting to github.com|192.30.252.131|:443... connected.
HTTP request sent, awaiting response... 503 Service Unavailable
2016-01-27 20:03:39 ERROR 503: Service Unavailable. $ git clone ssh://git@github.com/influxdata/influxdb-ios
Cloning into 'influxdb-ios'...
remote: Counting objects: 10, done.
remote: Compressing objects: 100% (8/8), done.
remote: Total 10 (delta 1), reused 10 (delta 1), pack-reused 0
Receiving objects: 100% (10/10), done.
Resolving deltas: 100% (1/1), done.
Checking connectivity... done.git clone git://github.com/MunGell/Codeigniter-TwitterOAuth.git .
where git clone https://... failed. YMMV, of course.
jim% git clone git@github.com:pfsense/pfsense.git
Cloning into 'pfsense'...
fatal: remote error:
GitHub is offline for maintenance. See
http://status.github.com for more info.Linus Torvalds: [facepalm]
Additionally, it should be easy to learn and have a tiny footprint with lightning fast performance. It should outclass SCM tools like Subversion, CVS, Perforce, and ClearCase with features like cheap local branching, convenient staging areas, and multiple workflows.
That would be a significant upgrade.
# pre-push (there is not post-push)
# Update in several places
git push bitbucket master
git push gitlab master
# ...[1] https://ipfs.io
In an ideal world, a link to a source code repository (or a link to anything for that matter...) would never fail because there would be automatic mirrors to at least provide read-only access to it.
It sort of forces one to ask whether this is the result of fundamental mistakes within HTTP / DNS itself. It's not realistic for a web designer to put time into making external links fault tolerant by running some query to switch the routing link to an available node.
https://ipfs.io/ipfs/QmNhFJjGcMPqpuYfxL62VVB9528NXqDNMFXiqN5...
:-) IPFS includes an optional http gateway to give the traditional web access to IPFS before browsers support the IPFS protocol natively.
npm i semantic-ui
npm ERR! fetch failed https://github.com/derekslife/wrench-js/tarball/156eaceed68ed31ffe2a3ecfbcb2be6ed1417fb2
npm WARN retry will retry, error on last attempt: Error: fetch failed with status code 503We would get these detailed as fuck responses on why one project (BitBucket) is worth more to the developer than GitHub instead of these little jabs at why it is better. I want thoughtful responses.
If there was a duplicate response, the employee that responds to various threads online can link them to their answers.
This isn't even directed at you @ntaylor. I think it's an exploitable marketing strategy. A clear example of this is Katie from PornHub. /u/Katie_PornHub (or whatever the user name is) posts on reddit, which in tern gets more interest in PornHub. PornHub is basically using reddit for free advertising.
Basically. I want to know WHY something is better than something else, and that WHY is with as much detail as possible with a lot of thought put into it. Give me a pros and const list between the two. Anything but two good things about it.
- - -
Holy fuck, this would just add a human element to advertising. You hire humans to serve ads to people. Whenever "GitHub" is mentioned, and you have a competing better service, their only job is to advertise your service to anyone have problems with the other service.
In cases like these, it always seems to happen naturally. Humans are willingly advertising for free, when they could be paid for it.
The downfall of the internet:
> We make up the ad network
> It carries us far;
> however---giants must fall.
Me.This would only get you so far. What if a user has a question? How will you help them out if it's a bit. The solution?
Write API's for every site that you want to use. You can either scrape every site for responses that fit your criteria, or let google do the digging for you with the "as it happens" notifications which would indicate that you have to scrape that site for the information.
If the information you get from the site is something you want to respond to, you can.
- - -
Holy fucking shit.
I just described a system that would allow for you to use all websites as inboxes for chat messages.
- - -
Refine:
With this early model, we would have one program on our machine that would understand X websites. It can parse user comments directed at us from a website (A). It can send responses back to A, and so on.
Now, users on Program X can chat with theoretically any person on any website (A, B, ...), but users of website A can only chat with other people on website A OR people using Program X.
Why doesn't everyone use program X? Because it does not exist yet.
We use Bitbucket in our ~30 member academic robotics research lab, it's wonderful because they are nice enough to supply us with as much as we need (private repos, teams, etc.) for free!
Guess they're too busy working on the issue to notice the misuse of "effecting" instead of "affecting."
1. Commit fails
2. Try again, see if it fails twice
3. Check internet connectivity
4. Try github in browser
5. Try github in a different browser
6. Go to HN to see if it's down for everyone
7. Write snarky comment
8. Go try again...
git remote add backup <bitbucket or gitlab url>
git push backup
while true; do curl -s https://status.github.com/api/status.json | grep good && tput bel || echo -n .; sleep 1; doneruns script
while true; do curl -s https://status.github.com/api/status.json | egrep 'good|minor' && say -r 160 "github is up again" || echo -n .; sleep 60; done
Edit: accept a status of minor as an indication of up-ness. # watch for a website to come back online
# example: github down? do `mashf5 github.com`
function mashf5() {
watch -d -n 5 "curl --head --silent --location $1 | grep '^HTTP/'"
}Hubot has assumed control...the singularity is upon us!!!
No word on twitter yet: https://twitter.com/githubstatus
git:(master) ok this is actually really annoying now
zsh: command not found: ok
git:(master) fuck
look this is actually really annoying now [enter/↑/↓/ctrl+c]
look: is: No such file or directory
git:(master) fuck
> No fucks given
That's kind of perfect.
The more things change, the more they stay the same.
I would like to have one bash script that I could start in such worst case scenarios that tests everything and tells me exactly what the problem is.
Just imagine everything is down and you can not trust your dashboards.
What would you put into such a script? What would you test for?
and also why I do include my dependencies in my projects
having a build tool fetching dependencies dynamically on github or whatever is not a convenience, it's a PITA
just sayin' even extremly reliable hosted services can fail, do own your repositories, going back to write some code ;)
edit: Now says "19:32 Eastern Standard TimeWe're investigating connectivity problems on github.com."
edit2: "19:47 Eastern Standard TimeWe're investigating a significant network disruption affecting all github.com services."
https://www.dropbox.com/s/ap1nqogmqza2e12/Screenshot%202016-...
And:
cd /[mason@IT-PC-MACPRO ~]$ cd Code/rollerball
[mason@IT-PC-MACPRO rollerball (master)]$ git pull
remote: Internal Server Error.
remote:
fatal: unable to access 'https://masonmark@github.com/RobertReidInc/rollerball.git/': The requested URL returned error: 500
[mason@IT-PC-MACPRO rollerball (master)]$Nothing for this yet, though.
EDIT: There it is: https://twitter.com/githubstatus/status/692505376554618883
A side project of mine called StatusGator [1] monitors status pages and alerts you when something is posted on them. I built it to handle the use case that you're trying to diagnose an issue only to remember to check the status pages of your dependencies and notice that they are already working on it.
It's pretty useful for those services that keep their pages up to date. But not very useful in cases like this when it's not updated.
January 28, 2016
00:00 EST The status is still red at the beginning of the day
January 27, 2016
20:02 EST We're continuing to investigate a significant network disruption affecting all github.com services.seems pushes are going into oblivion, but the client side gets a return code that the push was successful.
$ traceroute github.com
traceroute to github.com (192.30.252.130), 30 hops max, 60 byte packets
...
9 level3-pni.iad1.us.voxel.net (4.53.116.2) 17.609 ms 15.057 ms 10.113 ms
10 unknown.prolexic.com (209.200.144.192) 9.186 ms 9.462 ms 9.315 ms
11 unknown.prolexic.com (209.200.144.197) 17.753 ms 17.767 ms 18.851 ms
12 unknown.prolexic.com (209.200.169.98) 9.922 ms 9.542 ms unknown.prolexic.com (209.200.169.96) 11.471 ms
13 192.30.252.215 (192.30.252.215) 13.569 ms 192.30.252.207 (192.30.252.207) 9.660 ms 192.30.252.215 (192.30.252.215) 13.150 ms
14 github.com (192.30.252.130) 9.051 ms 8.833 ms * circuitry git:(feature/middleware) git push origin feature/middleware
Counting objects: 18, done.
Delta compression using up to 8 threads.
Compressing objects: 100% (18/18), done.
Writing objects: 100% (18/18), 5.01 KiB | 0 bytes/s, done.
Total 18 (delta 9), reused 0 (delta 0)
To git@github.com:kapost/circuitry
* [new branch] feature/middleware -> feature/middleware
This branch is not appearing in the repo: https://github.com/kapost/circuitry/branches Downloading: https://github.com/driftyco/ionic-app-base/archive/master.zip
x Invalid response status: https://github.com/driftyco/ionic-app-base/archive/master.zip (503)
Error Initializing app: [object Object]
errorHandler had an error [TypeError: Cannot read property 'error' of undefined]
TypeError: Cannot read property 'error' of undefined
..."Damn! I'm tired of fighting installing cordova/ionic. WTF is happening now? Oh Github is down... okay that's new"
(they're using a different but similar image now, but they used to use that exact same logo)
Between 2016-01-28 00:39:47+00 and 2016-01-28 00:43:26+00 there was a flurry of BGP updates that caused that.
I'm not sure on the exact timing of the outage, this could either be a symptom or a cause.
But my website loads the content from .MD files via raw.githubusercontent, and those appear to be down.
(Positive in this sense has a similar meaning in behavioral psychology's terminology for operant conditioning, adding punishment.. which both come off as odd with our common use of positive as good and beneficial :/ https://en.wikipedia.org/wiki/Operant_conditioning)
...Carrying on with your analysis, since uninterrupted positive feedback loops result in explosion, I guess that would mean either (for this outage) github's fallback status page would go down or the parts of the network carrying that traffic, which could not handle their increased load, would go down.
Probably better we would rework these control mechanisms and reroute \ rewire \ entirely change these processes.
[Oh My Zsh] Would you like to check for updates? [Y/n]: y
Updating Oh My Zsh
Username for 'https://github.com':
> 00:00 EST The status is still red at the beginning of the day
They seem to have some time issues as well since it's not yet January 28th in EST.
> We're investigating a significant network disruption affecting all github.com services.
$ brew reinstall go --head
==> Reinstalling go
==> Cloning https://github.com/golang/go.git
Updating /Library/Caches/Homebrew/go--git
Username for 'https://github.com': ^C
No way would I do self-hosted version control. I have better things to do than babysit servers for commodity services.
I can learn from my mistakes and improve progressively.
If I rely on another instead, I cannot, and I cannot (necessarily and confidently) see to it that they learn from their mistakes and improve.
By relying on another (at least an unreliable \ uncommunicative \ uncooperative one), I cannot improve my chances of them not making those same mistakes again, taking me down with them.