GitHub was down
status.github.com
status.github.com
Making your organisation too dependent on a remote service can indeed be a scary prospect and I'm not sure what GitHub offers to mitigate this.
But hey, I guess there's the comfort in know that when it goes down it's our fault!
If you know you have a major deliverable coming up, you can choose not to touch your git server until it's at a somewhat safe moment to do so.
If you have "danger times", then any unplanned disruption in that time will hurt so much more.
The odds of GitHub going down and your local copy/instance simultaneously is very low so availability is high.
Reliability is how often GitHub or your local instance goes down.
Not that I'd suggest using it for every scenario, but being aware of alternatives is always a good thing.
[0] http://fossil-scm.org/index.html/doc/trunk/www/index.wiki
At every company I've worked for, internal services have been less reliable than github. Certainly way less reliable than gmail.
I get that it's scary, in that it feels like you're giving up control over something important to your business. But I'll posit that you never actually had control, only the illusion of control.
That's a weird way of phrasing it. When you run your own services and they break you have total control and do have the power to fix it.
When you buy SaaS you are relying on someone else. You may very well have more reliability and uptime but you are nonetheless giving up control.
Sure, but most of the time you just reboot the server or some other temporary hack and kick the problem down the road.
The power and skill to fix isn't worth much if you don't have the time.
Some folks just hate the feeling of not knowing what's happening and how long its going to take to fix, vs having direct access to work on problems themselves even though its not necessarily "better" in any sense of the word ... hence the illusion
Software teams are often pathologically unable to create a deliverable that is scrutable and manageable by an operations audience. Building a product that can stand on it's own two feat without constant minding is work.
For all the testing dogma that files around, constant delivery has resulted in systems that are more brittle than ever. If someone with all the institutional knowledge is always there to catch the system when it fails, why bother making it easy to investigate failures? Why bother with analysis and building in fault tolerance when someone can worry about a long term solution for that failure mode when it causes them to get paged at 3 am?
So it becomes easy to say minders are necessary, and they must be developers. Business incentives mean that answer isn't always scrutinized as heavily as it should be, because it means spreading maintenance costs over time instead of an upfront investment in resilience and maintainability. Refusing to provide a toolset and manual for maintenance means some nobody with access to google can't fix 99% of the problems that could occur. That way the magic black box creators can ensure they're the ones getting paid to do it.
We don't have or need Windows engineers, kernel engineers, or Cisco engineers on call. Nor do we have nginx, postgres, cpython, exim, apache, php, mysql, Active Directory, Exchange or Office engineers on call. We use weird enterprise software from companies that have gone out of business, we don't have high-priority support contracts with many of the rest.
Software that needs minding from the developers is just bad software. That's technical debt they took on to get the product out the door.
If I had total control of a service, I would make it have 100% uptime. Wouldn't you? The fact that there exists no service with 100% uptime indicates to me that nobody has total control of their services.
Similarly wrt having the power to fix my services. If I had unrestricted power to fix things when they broke, then I would use that power to fix all breakages immediately. Since nobody seems to be able to do that, I conclude that nobody actually has unrestricted power to fix breakages in their systems.
Or do you mean to say, I have some limited ability to control and fix the systems I run? I would agree with that. That's my whole point.
edit: We don't like the recent changes made in githost. ( Replaced the $35 option with $149 ) We're a team of 15 people.
edit2: should have clarified; I'm hooked to GitLab, just looking for another service that does managed hosting. Not looking for GitLab alternatives.
If you're really concerned, you run it yourself. GitLab insider a docker container isn't hard to manage. You can even 1-click Gitlab onto a DO Droplet in just a few minutes: https://www.digitalocean.com/community/tutorials/how-to-use-...
[0] https://gitea.io
thank goodness I'm on bitbucket :) One of the things they offer is your own local server synched with their cloud servers for redundancy.
then architect your build system be able to switch to another service with the flip of a switch, if its that important to you.
Unfortunately people only care about API's now instead of standards.
Obviously, this doesn't trivially scale to many repos/users/etc, which is why Github exists in the first place.
However, I'd say one of the biggest practical problems with most of the popular DVCS tools is still that they don't also have a simple, out-of-the-box way to set up a centralised repo. As you say, using Git+SSH works up to a point, but it's still irritating and somewhat time-consuming to set up if you work with a lot of projects that might each have different contributors. Last time I checked, none of the main alternatives was any better.
This seems to be a fundamental weakness in whatever deployment system is used, then. Relying on a single point of failure outside of your organisation's own control for anything critical is always risky, and in this case it seems to be entirely avoidable.
What I personally do in AWS is bake my artifacts or other git-sourced data into AMI's. If you want a middle ground, you can instead push your artifacts to an s3 target -- s3 has better reliability / track record than github for this purpose.
If Github being down means you're not fixing your site today, I'd call it a runtime dependency.
(Hint: sans hug of death, because a runtime dependency brought it down.)
Deploying/scaling is somewhere in the middle.
The beauty of Git's design, though, is that you've still got everything to use it distributed even if you, under normal circumstances, just use it centralized like a fancy SVN.
There are probably a lot of other tempting features to put on your deployment critical path but I'm not in an environment that uses github so I've forgotten.
It's not a runtime dependency, it's a stupidity to rely on random repositories from the internets when building and deploying (remember the left-pad farca that halted work for half of the web programmers). You should always have your local copy of necessary sources or use repositories that can trivially be swapped with some other mirror.
We used to have the infrastructure for those things (in the form of binary and source package repositories), but it was mainly used by sysadmins, and programmers usually have an alergy for such infrastructure.
Really he is speaking to the problem with SAAS. Sure it's cheaper to rely on someone else to do the heavy lifting for you and they can do this because they 'make it up in volume'. But the other side of that coin is no one really knows how to do that anymore. If you wanted to roll your own it's going to be very hard for your company to do the basics because you've become dependent upon the cloud, and someone else paying employees to do the basics for you... The basics have never been hard.
Opening a port to your local instance of git implies more thinking and security consideration. Of course you could host a mirror on a cloud instance, but then it's saas anyway and you have even more work to do before ever starting coding.
So yeah github/gitlab does the heavy lifting for you, but for small organisations that might be cheaper overall than to pay somebody mastering all required stack to implement and maintain your own instance(s). For big company no question that an in house git team is probably more efficient.
This is our tactic. I tried to do baked AMIs at one point, but the 10-15 minute turnaround in registering them meant that we couldn't use them for staging or testing (too long to iterate changes).
Previously we were capistrano-deploying with git from bitbucket - every server had to individually git pull from servers on the other side of the continent, over the public internet. Susceptible to all sorts of problems.
Github is an off-site backup, and a nice interface for pull requests and code browsing.
There's just no reason to make github a critical part of your infrastructure.
GitHub isn't a deployment tool. It's a source control tool. Keep your deployment artifacts somewhere else (e.g. S3/GCS) more durable so you can still deploy/rollback during a GH outage.
Our stats show a 1:239K request error rate for S3, and a near 100% availability. There has only been one outage of note in years. GH on the other hand is down with shocking frequency like today and yesterday, though not as much as BitBucket.
Some companies I know write to GCS and S3 in parallel, which is easy given many tools and API's work directly with both (e.g. AWSCLI/gsutil).
At my gig, we use GH Enterprise (on-prem) as well as some smaller plain-vanilla git repos (mainly for CM and other critical services). Any third-party code and all other dependencies live in an artifact manager or private mirror repos.
Personally, I don't use any cloud provider to store my data. Cliche, but true: if you don't control the hardware, you don't control the data on it.
If you _need_ Github to be up, mirror it. If your build or deployment tools _need_ Github to be up, and you _need_ them to be always available make simple dumb mirrors.
Also make sure your tools don't lock on only one item.
http://fortune.com/2016/09/02/us-government-embraces-cloud/
Looks like even the CIA is using some AWS.
"U.S. government agencies are moving to cloud computing and away from their own data centers faster than private corporations"
They could also have a base version of the site that's read-only for source or binaries to allow customers access to the data to fall back to whenever problems happen in more complex functions. Keep it running in some form while they fix things. They might also similarly sell Git appliances w/ high-availability that themselves plug into multiple colo's or cloud providers mirroring changes in the repo to the Github site. Just a wild idea as I have no idea if that's marketable.
When you build a docker image, you copy all your runtime dependencies (`node_modules/`) and build artifacts into the image. When you push a deployment you are deploying that static image with the dependencies baked in, instead of trying to install the dependencies at startup.
Isn't every git client a server? Since git is peer-peer, when GitHub is inaccessible we just merge back and forth (or just commit locally -- it's never been down for days).
Ignore this if you are on a big team of course.
Tried to prototype on XMPP at one point in college but never got very far in the prototype.
(That was back when I was a heavy darcs proponent too and had other crazy things like automated consensus branching as a great distributed way to coordinate efforts on such a distributed project, which would work with the darcs push/pull/merge model but not so well in git's.)
Created by Richard Hipp, creator of SQLIte, it's open source and free, has zero dependencies and works across all operating systems.
And even if you don't want to host it yourself, there's always chiselapp.com [2]
[1] http://fossil-scm.org/index.html/doc/trunk/www/index.wiki [2] http://chiselapp.com/
1. No matter what 99.99999999% availability a service provides, its utterly useless if the time to get back up is unacceptable.
2. Do not have remote services, which you cannot fully control to be a part of your run time deployments.
I hardly think it's reasonable to consider services like GitHub or infrastructure like AWS as being irregularly updated...
A long down time might be worse for you than a shorter down time more frequently.
99.99999999% uptime is 3 milliseconds of downtime a year. That's not even a TCP retransmission.
If this service is up for five years, then it can go down for 15ms and still claim all ten over that period. Which is almost nothing still, but if you needed ten to begin with then maybe it's quite a lot for you.
All I'm saying is - and I think this is what the commenter originally twice above me meant - is that the rate doesn't give you all the information, you at least need a period as well.
The idea that you push the source code then build it on the server makes no sense to me.
If I was github I would forbid (and punish) this kind of behavior.
> 18:19 CDT Major service outage.
$ gl
fatal: remote error:
GitHub is offline for maintenance. See http://status.github.com for more info.Engineers stop the bleeding by 503'ing requests at the perimeter or putting up a static maintenance page. This allows things like caches or DBs or app servers to cool off while a rollback or a revert goes out. Then, when the system is stable again, let requests flow through again (slowly, of course).
edit: confirming, push doesn't work either
18:41 CDT Everything operating normally.
How many servers did the average app depend on in 1998? 1? Get dual HD and you were in decent shape.
Compare to a modern microservice app, that maybe depends on 100 internal services and 4-5 external services. A lot of things need to go right or mostly right for things to function.
It takes an amazing team to cover all the bases.
The fact that VAXClusters were already going out of fashion by that time, and the fact of IBM's parallel sysplex existing by that time already negate your point.
previous $employer had built a globally distributed multisite c&c processing system twice by then and were in the process of revamping it for a 3rd.. 1st on mainframes & remote serial compute devices in the 80s, then on unix workstations in the early-mid 90s.
They were by far not alone in dealing with this level of complexity for the time.
5 racks in 32 cities as POP. Each POP had 10 Sun Netra T1 1RU boxes behind SLB. All those were a service cache to offload the main DC which had...hum... ~100 or so Sun E450 (4x480Mhz Usparc3, 4GB RAM, 20x4GB or 9GB UltraSCSI3 drives, 2x1G NICs).
There was a cluster filesystem also. We had huge databases, etc.
This is just not that hard today.
There's a big difference between what a high-capacity/availability site had to handle in 1998 and 20 years later.
There are also lots of easy criticisms one can level at github, given their uptime and what they've published about their architecture. 'What they're doing is fairly simple' is probably not among them.
My bet is a human.
Hotmail, Altavista, Yahoo Mail, eBay were all a thing then.. and it wasn't like noone was using them at the time either..
which isn't to say everything hasn't advanced, but in my opinion this isn't really a 'difference in kind'..
I think doing this at scale is simpler today because the networking is simpler, i.e., IP fabrics and VMs and containers. If you are building something today and want to use Layer2 I question your santiy.
"As of 1998, it used 20 multi-processor machines using DEC's 64-bit Alpha processor. Together, the back-end machines had 130 GB of RAM and 500 GB of hard disk drive space, and received 13 million queries every day"
There are probably more read-only queries to gists per day than that.
mmhmm
Someone is doing all of the kinds of work necessary to verify that the servers are up, being useful, and nominally working the way the previous deployment did, but not everyone has the tools yet to do that (especially the last one).
We need a compete and complementary tool chain that everyone can use, and so far a lot of these are still business differentiators (i.e., proprietary)
Once is an accident. Twice is a coincidence. Three times is an enemy action.