Incident Report: Inadvertent Private Repository Disclosure
github.com
github.com
One of our engineers came up with a useful script to grab all unique lines from the history of the repository and sort them according to entropy. This helps to lift any access keys or passwords which may have been committed at any point to the top.
I think this is a great example to illustrate the tough edges of security to less experienced engineers. Github will most likely never let something like this happen to you, but on the off-chance that they do it's great to be prepared. Additionally, the response from Github was very well received. No excuses, just a thorough explanation of what happened.
I also can't help but mention that we're hiring, if you'd like to work at an organization that values security and data privacy very highly. :) usebutton.com/join-us
It might be a fun thing to open source as part of a "I've inherited a project, what now?" toolkit that helps you decide what to fix.
https://gist.github.com/jasonmoo/06691c8fea09b62aa35235fc93e...
The email was kind of funny though, part of it was effectively "if you have this data pretty please delete it without looking at it". I'm sure that's the best they can do, but it still made me chuckle.
require 'facets'
lines = Dir['**/*.rb', '**/*.py', '**/*.cpp'].map { |f| File.read(f).lines.map(&:chars) }.inject(&:+)
puts lines.sort_by(&:entropy).map(&:join).last(10).reverse
Using: http://www.rubydoc.info/github/rubyworks/facets/Array%3Aentr...For example, could you pipe the output of `git log -p --all` through this and filter out all the commit hashes somehow?
lines = `git log -p --all`.lines.map(&:chars)
So I found that `git grep /.+/ $(git rev-list --all)` is a better way to get the content of all the files: https://gist.github.com/Dorian/e1514535c3c5036cf327ce61eb34a...But actually an hex number regexp might me far more accurate than the entropy (e.g.: secrets are often long hex numbers).
I tried it and it yields interesting results: https://gist.github.com/28110f0b8105db11e8973d1d0be85259
It's obviously unfortunate in this case, since even a relatively small and quickly fixed bug affecting a tiny proportion of requests still had serious consequences.
However, it's a remarkable achievement (if also a little terrifying for the software development industry from a single-point-of-failure perspective).
Self-hosted git (through the many installable git servers or raw git) running on a correctly sized box is almost certainly the way to go
It's a difference of degree: compromising my self-hosted or on-premise server means that somebody already in my employ has more access than they should. If I'm a small organization, that probably doesn't matter. If I'm a big organization, I probably have an IT staff to deal with this and the people involved are still "nominally" under my control.
The github mistake means that people completely unrelated to my repository can get access.
You don't hear about intrusions into self-hosted source repositories. Not because there are fewer, but because they likely don't have the security infrastructure in place to know that they ever happened.
[Edit: Multiple downvotes within moments of each other do not make calling out the above speculation any less justified.]
Given that git provides many transports and ways to push commits around, I agree. If you have to be safe, there's no reason to use github or a self-hosted git collaboration service (gitlab, etc) on a publicly accessible server, regardless of access control measures. If you really need to have sources on a remote machine, you can limit the potential damage by only sharing archives of a certain revision, without history.
I know many will dismiss it, but if you're serious about the repositories being private, then the most you make them accessible is via an on-premises hosted gitlab instance, which is local to the company network, not accessible via the public network, and only allowed to, if you want, by dialing into VPN first. Then, to be safe, you null-route anything but the VPN traffic on the connected off-premises developer machine.
Access keys get stolen, just like SSH keys are, so you need to use a VPN service that requires additional security like the use of OTP key generators or similar measures.
This probably sounds like a hassle in the day and age of people just going for the comfort of private github or gitlab repositories, but it's what companies have been doing for almost 20 years as standard practice.
You cannot consider any git repository, even on your own root server, safe to keep private code on. CIOs would argue against that practice for good reason. The same CIOs require work laptops to encrypt all data.
If you don't need to be that serious, then an incident like this should be planned and accounted for as part of using such hosting, and shouldn't be a big deal.
The deleted code is very specific-looking. Nobody writes that just casually or out of ignorance. Also it is what was at use in production.
It's very naive to just go and replace that with nice-looking, shorter code.
Key lessons:
- Understand what you are deleting
- Treat production code as sacred
- Add reasonably extensive comments for delicate code (as the original one). Git commit messages aren't enough.
- Try out infrastructure changes in production-like staging servers. I really doubt they properly did, as they say the "majority" of 17M requests failed.
> The impact of this bug for most queries was a malformed response, which errored and caused a near immediate rollback.
Surely they do some end-to-end testing?
I don't know of anyone that would recommend creating tests, even integration tests, that hammers a service to check to see if something like one hundredths of one percent of requests returns invalid data. If anything, the fact that a script is hammering a service that probably (in a Dev or QA environment) has much less data in it's database and file stores, and much less protection (like load balancing and caching) than it would in production would generate more false positives than it would generate in substantial data disclosure regression defects.
To me it looks like poor design , I would expect private repos to be hosted completely independently and in isolation with more secure and throughly audited code with longer release cycle (LTS ?) after the code has been well tested in the public free repos.
It is not excessive if you consider the potential value of the private repos that github has control over. They already do something similar for enterprise edition. It leaves bad taste that smaller customers are not treated with similar caution
I don't know that this would make you necessarily want to make both the change to self-hosting, and the change of platform.