One in every 600 websites has .git exposed
jamiembrown.com
jamiembrown.com
Keeping your entire server-stack up-to-date, making sure you have SSL, using strong encryption for logging-in, hashing the passwords, making sure your server can only be reached via SSH, adding firewalls, filters, etc. etc.
Then some hacker in Eastern Europe comes along (or some beginner at the NSA/GCHQ) and finds out that your .git is exposed and somehow gains all vital user-data and admin data.
Being bashed with a boulder repeatedly would probably be less painful than the torture of knowing "I did it all, but they got me with an HTTP request... because nobody thought of double-checking what our VCS is doing".
How many other glaringly obvious mistakes might be out there right now? I can only imagine.
Don't put secret keys in your repository.
Someone getting a copy of your code should be a big annoyance at worst.
DB_USER=scott DB_PASSWORD=b3withm3pl3aze /usr/bin/python webapp.py
That could be troublesome if an attacker figured out a way to run remote commands on your server even as an unprivileged user.
pikachu@POKEMONGYM ~ $ sleep 99 &
[1] 21340
pikachu@POKEMONGYM ~ $ cat /proc/21340/environ
XDG_SESSION_ID=5COMP_WORDBREAKS=
"'><;|&(:TERM=screenSHELL=/bin/bashXDG_SESSION_COOKIE=8571b679eed8952dd96ad28a54...<etc>(I actually gave the wrong example in my previous comment. While it is true that giving the ENV on cmdline will show up in ps eaux, the more appropriate example is what I just explained in this comment.)
Now, if an attacker manages to get root access then it's game over[1]. That just shouldn't happen. But nobody should be running their webserver as root. So, whatever that user is should be low-powered with only enough privileges to start the webserver & bind port 8080 (and use iptables or whatever to reroute connections to port 80 --> 8080) and the whole setup should be designed that this account won't be able to escalate things further if someone got a bash shell to it.
______
1. You should at least have some way of detecting that it happened and consider all data & files compromised and just wipe the whole machine & start over. Or take that machine offline for investigation into what happened and put a fresh new one in its place.
It would be kind of weird to include that considering how argv works.
What are your other 11 objections?
/srv/www/domain.com/public_html/
|--------->/private/
|--------->/logs/
|--------->/tmp/
Anything stored in /private/ is not publicly accessible by the web server process, but can be read or written by anything running under the user's username.It's specifically for storing things like configuration files.
I think this should be standard practice.
Disclaimer: I worked on it.
It's relatively common to provision secrets with configuration management software like Chef/puppet/ansible/etc using, e.g. Chef's encrypted data bags.
Another slightly heavier-weight solution with some nice properties is to use a credential broker such as Vault: https://www.vaultproject.io/
I typically handle this by versioning a `config.example` file, which includes all the necessary config keys an application expects. The example file defaults these attrs to various strings meant to show they are examples only. I include instructions to copy the `config.example` to a `config.yml` (or some other appropriate extension), and replace the values as necessary. The `config.yml` file is specifically excluded in the `.gitignore` file. The application will only load the `config.yml` file when started, so I also ensure to raise a descriptive error informing team members when they are missing a local `config.yml`.
This allows the `config.example` to also serve as a self-documenting config for the application, as comments can be included that identify and explain each of the config keys and their purposes.
Usually the base config is in VCS but without user/password/db strings. We then manually configure the file with the encrypted strings on the server (usually with the machine name in the filename so that we can use hostname in code to find it and makes it clear the file is machine specific). Not all tools make this easy though and only works if you can add your own code in between. Also prefer files to environment as the files can be locked down easier in my opinion and more obvious what is going on.
I like some of the other solutions that are using encrypted strings but with a keystore server and may consider for the future if they support both windows and linux.
There's a variety of mechanisms for loading this into your environment.
1. Create a JSON file containing encrypted secrets (DB pass, etc.)
2. Upload the file to a secure S3 bucket with fine-tuned permissions and server-side encryption
3. For the instance that is launching, include the permission in the IAM role that allows it to "S3:GetObject" on the specific JSON file you uploaded.
4. Deliver decryption keys to the app in some manner (Chef, Ansible, etc.)
5. When the app starts, it downloads the JSON file and loads the environment variables.
I wrote a blog post[1] about this with more detailed info and an NPM module if you use Node.js.
[1] http://blog.matthewdfuller.com/2015/01/using-iam-roles-and-s...
I have a question about redundancy or "What happens if your gatekeper EC2 instance goes down"? If you have multiple gatekeepers could they be set up this way:
- let's say you have five different web apps using a gatekeeper to hold their secrets
- let's say you have n gatekeepers (let's say 3) and each of the apps knows the address of all three gatekeepers.
- If the primary gatekeeper is unreachable, all five apps would try to contact the secondary gatekeeper, but that gatekeeper would only (ever) respond in the event that the secondary gatekeeper also found the primary gatekeeper unreachable.
It's like a sleeper cell - at any given moment you have multiple replacement gatekeepers ready and waiting to serve, but each of them is unable to respond unless the one above it in the list stops responding. In this way you could lose gatekeepers (even permanently) and build a little bit of resilience into the apps depending on it while you're able to sort out what happened and restore normal behaviour.
Is this a good idea?
Another way is a second file that overrides settings as needed. Although I have found that to be less maintainable if the configuration file changes. That file should be somewhere entirely out of the VC tree.
Either way, the file must be placed in a directory that is not served by the web server.
/include
and
/public
are traditional. Only /public is exposed by the web server.
use git crypt[1][2]
use git-dir and work-tree options/env vars
[1] https://www.agwa.name/projects/git-crypt/[2] though you have to remember to git crypt lock
The right lesson is: Know where your secret keys are and take the appropriate steps to secure them. Whether that's in the codebase, a properties/ini/conf/whatever file, environment variables, whatever - know where they are and make sure you understand possible threats against them.
This story could just as easily have been written about how easy it is to download ALL_THE_SECRETS.txt. Don't feel smugly secure just because you don't store passwords in git.
1. Make a list of everything that absolutely positively cannot live without this data/access/permissions/etc.
2. Put the data somewhere where absolutely nothing whatsoever can ever read it (except root).
3. Figure out what one single change will resolve #2 so that the things in #1 can happen without any other things gaining access.
If you don't do #1, you don't understand your requirements/applications. If you don't do #2, then your data is probably vulnerable through some other mechanism. If you can't do #3 then you probably need to change something else (e.g. stop running all processes as the same user, stop running all services on the same box, stop trusting users, set up more granular sudoers rules, etc).
What I find is that when you come up with an idea for #3, and then come up with a list of side effects, you can actually find a lot of the kinds of issues I mentioned above, for example where the public website CMS (as 'daemon') and the accounting backend (as 'daemon') both have access to the same resources, and thus someone gaining access to the CMS can get the accounting DB user/pass and get access to your transaction records, user database, etc.
It's important to know where your keys are, but it's also important to not store your keys in certain ways that are easily overlooked.
A lesson of "don't put secret keys inside the web root" is also useful.
But a lesson of "know where your keys are and secure them" is a bit too short-sighted. You don't just want them to be secure right now, you want the mechanisms keeping them secure to be mistake-resistant.
Don't put them in the code, even if you promise to be super careful.
It's really bad out in AppSec land
The OWASP Top 10 is deadly.
You only need to make a single mistake and you are hosed. Your attacker can fail an arbitrary number of times and only needs to succeed once.
If you are 99.9% likely to make the right call on anything that could have a security impact then you only need to make 1000 decisions before you probably screwed one up and have a hole.
Some would say this means true security is impossible.
Sometimes our mistakes are not recognizing the right thing to be done, or not recognizing anything at all.
I think the view that true security is impossible is true if security is up to one person. What about a system with multiple layers of sign-off, or an automated system that can help test the security of what you're doing and alert/prevent dangerous behaviour without specific reasons.
Don't use the root of your app as document root!
It's really as simple as that. Almost all modern apps have a subdirectory "public/" or similar. That one is meant to be used as document root. You only have to ensure there are no sensitive files in there.
If you fail to introduce such a directory, you'll have a game of cat-and-mouse, where you have to add extra webserver rules for each sensitive file: VCS, crypto secrets, private keys, and so on. In that setup it's easy and very likely to forget one. Of course, this then creates the feeling of "How can anybody keep track of this never ending list of security details?"
In the end, this is a blacklist vs. whitelist thing. Like with your firewall, you want one rule that blocks everything and allows only specific stuff. The alternative is to allow for everything, have rules to deny all sensitive stuff, and finally get in trouble for having forgotten one rule (e.g. probably because an additional service was introduced after the firewall rules have been written.)
So the connection to PHP is apparent till today, although it may have to do more with crappy super-cheap hosting providers than with the language itself.
.git
web/index.php
src/*.php
Make the document root /full/path/to/web/index.php.Edit: Also use a deployment tool like Capistrano which removes the .git directory as well.
* Wordpress
* Tiki
* ... and so on.
I don't think this is by accident: This technique is too old and well-known to be ignored by large projects. Rather, they conciously don't do that to cause less hassle for the occasional admin: Those can simply dump it into some directory and don't have to setup a proper docroot or anything. And extra work for those who want a more secure setup.This is clearly a usability versus security issue, resolved in the unfortunate, usual way.
It's possible to set it up properly on most of these hosts, but it's much more difficult, and if you ever have issues the support team says you're using a "non-supported configuration." At least WordPress lets you move your config file one level above the webroot and will find it automatically.
Look at any larger Python (Django/Flask/...) or Ruby (Rails/...) project, and you will always find a proper "public" directory, although it may be named differently.
Not sure about Perl though ...
<Directorymatch "^/.*/\.git+/">
Order deny,allow
Deny from all
</Directorymatch>
<Files ~ "^\.git">
Order allow,deny
Deny from all
</Files>
https://serverfault.com/questions/128069/how-do-i-prevent-ap...See also the nginx question: https://stackoverflow.com/questions/2999353/how-do-you-hide-...
Note, if you actually did take it from the stackoverflow, you just infringed on someone's copyright; SO's user content is 'creative commons, attribution required'.
Edit: Thanks for adding the attribution, all clear with copyright now :)
edit: sounded wayyy too snarky lol.
Does a simple access rule, which I can't see there being many sane ways to express, meet the minimum level of creativity/originality to be eligible for copyright...?
I.e.
* No secrets in your repo
* Only copy to the server what you need (and automate this)
* Add conditions to your web server to not serve up .git, in-case the previous two checks failed.
When working in teams I think having additional checks and balances and not one 'perfect solution' is vital.
$ rm -rf .git/
It's way safer to delete the repo history from the production server than to rely on Apache rules copied from a forum.Unfortunately I can't seem to convince anyone that this is good practice. :-(
rsync --checksum --ignore="..."
unless you have a damn good reason not to (--checksum is essential to prevent corruption/malicious modification, without it you are implicitly assuming the version on the remote machine is exactly how you left it: that assumption is why Linus built git around shasum in the first place).rsync is even easier than SSHing to git pull, or opening up a pushable repo on a server. For once the simple approach is clearly better!
That's not true (though not completely wrong).
rsync is stateless. It does not assume the version on the remote machine is "exactly how you left it"; Rather, it compares file size and file modification time; if either changed, it will do a transfer -- efficient delta transfer, usually - which might be as little as 6 bytes if the contents is exactly the same.
--checksum makes it ignore file size or modification time, and compare the file checksum in order to decide if it's time for a transfer (delta or not).
A malicious actor, or bad memory chips, might change your file's contents, but keep the file size and time/date the same. In that case, --checksum will overwrite that file with your source version, and a --no-checksum wouldn't. So it's not bad advice. Whether the cost in disk activity is worth it depends on your thread model, data size, and disk activity costs. (Though, if corruption is due to bad memory, this is the least of your problems)
However, a corruption because of a program error / incompetent edit to the file is very unlikely to leave both the size and modification date intact - and a standard rsync will figure that out as well.
If the comparison is with using Git, then it is clear you're not so resource constrained that you can't countenance running MD5s, since Git would run shasum.
I think in our current laissez-faire climate w.r.t. security, I think recommending leaving security on the basis of saving a few cycles isn't very wise.
> It does not assume the version on the remote machine is "exactly how you left it"
I was ambiguous and sloppy, sorry. It doesn't not check for changes in a secure way, but assumes that, as long as the meta-data for the file matches, the content is as you left it.
When Linus built git, he specifically did so around sha1 to ensure that you ensure the data you think you have in the file, you do in fact have. rsync --checksum is thus a reasonable replacement for git deployment, but rsync --no-checksum isn't, imho.
Sorry if I was vague or misleading, thanks for the clarification.
Remember, git providers go down (i.e. DDOS to GitHub or internal fail at BitBucket). Don't depend on git being up to deploy your code or you'll look like a fool next time a DDOS at GH coincides with a deployment.
If you have the proper secret segregation now, but you're deploying by doing a git pull, now you run the risk of not really having segregated secrets all over again.
The tech industry is shaped like a funnel, with lots of raw, bad ideas at the top and a few smash mega-hits at the bottom. 99% of the ideas at the top are bad; investing more time than is necessary to prove them out is a mistake. 100% of the ideas that make it to the bottom wish that they'd spent more time designing things at the top. But y'know, if they'd actually done that, they wouldn't have made it to the bottom, they'd be outcompeted by the guy who got a quick and dirty prototype up, made his users happy first, and then closed the gaping security holes (hopefully!) before anyone noticed.
The balance is generally far more on the quick 'n dirty side than most engineers (myself included) would prefer, but we could look at this as a cognitive bias of engineers rather than a failing of nature.
I'm sure you've daydreamed of the facial expression of the scriptkiddie the moment he stumbles across your fake keys illicitly, only to be disappointed hours of unfruitful hacking later :)
Sadly, in my experience hardcoding secrets such as (database) passwords and encryption private keys is not uncommon at all in web applications. I don’t like criticising other developers, but sometimes the people who get to make these decisions don’t necessarily have the perspective or experience to make the rights calls.
It really annoys me randomly hearing about critical security issues through tech news websites - there should be a more systematic way for "non-security professionals" to ensure their sites are protected to best practice levels.
When googleing for "inurl:.git", it returns no results. And on top of that, I need to enter a captcha first?
But zero results is blatantly wrong.
For example, check the URLs of these search results: https://www.google.com/search?q=inurl:%22.hello%22
I suddenly started getting a captcha for searches right after using that search. o_O
Deploy the secrets separately: they don't belong in your site's codebase.
Except in the rare cases where it is intentional e.g. an open source repo and you happen to want people to download it from the same domain not github or git.domain.com.
Presumably, for most commercial entities, the parts of the site that are valuable are the assets, which are served from the site as part of it doing the thing it's meant for. For a large percentage, they're running a CMS like Wordpress or Drupal or whatever, where the codebase is public anyways. And for even more, we're talking about a directory of hand-crafted HTML files, where the version control is the HTML of the site plus some "damn, I always forget to close my tags" commit messages.
Because there is no need to expose it. If you expose it then you need to vet not only it's head but it's entire history, and for what purpose?
> For a large percentage, they're running a CMS like Wordpress or Drupal or whatever
I doubt someone running Wordpress would be using version control anyway. Most of the time it is WP + some standard plugins and no custom coding.
> And for even more, we're talking about a directory of hand-crafted HTML files
And anything else you happened to have in your directory. Your passwords file? Even if you commit a delete there is still history!
Because for a non-trivial number of websites their codebase is their IP and product and not something they want to be public.
I'm all for open sourcing as much as possible, but Google doesn't publish their search algorithms for a reason.
location ~ /\. {
deny all;
access_log off;
log_not_found off;
} location ~ /\. { deny all; return 404; }Git is popular, but I find it hard to believe that 1/600 of all websites on the Internet use it.
So in general - it's better not to have it in the first place, because it's unlikely that the person doing the commits knows the whole deployment strategy.
Still - having a real deployment process is much better. It could be as easy as extracting the contents of a tarball generated by git archive.
Of course, doing everything is much better.
Nowadays, lots of open-source projects encourage ordinary webmasters to clone a Github repo and run `git pull` to update.
So I suspect that public .svn folders will be less common.
`git archive` requires you to have a clone of the repo from which you can create an archive. That means people are more likely to just do a local checkout than play with archive on top of it.
Deploy with svn export:
svn export url.of.repo destination/path
Deploy with git archive: git clone url.of.repo
cd repo_name
git archive --format=tar some_commit_or_branch | (cd destination/path && tar -xf -)For example:
git clone --separate-git-dir=<repo dir> <remote url> <working copy>
where <repo dir> is outside of any directory served by the web server and <working copy> is the htdocs root.This option makes <working copy>/.git a file whose content is:
gitdir: <repo dir>
The advantage is that all git commands work as usual, without the need to set git-dir and work-tree, and that there's nothing special to add to the web server configuration.It may be possible that gitdir is still accessible through a misconfiguration or security issue (and you're giving them exactly where to look)
Production servers have no business having the .git directory anywhere.
I agree with you in principle, but in practice this is not always possible, there are situations when having a git checkout in production is better than nothing.
I've seen WordPress sites where a semi-technical administrator updates plugins and themes directly in production: with that git checkout I would at least be able to track the changes and pull them in a dev or staging environment.
This could be the first step to a saner deployment workflow for those sites, where production gets changes that have been tested and validated elsewhere.
If you're using a modern framework with url routing, you don't need to worry about hiding .git or .hg in your webserver config file.
$page = $_GET['page'];
include ($page.".php");
Whilst allow_url_include (http://php.net/manual/en/filesystem.configuration.php#ini.al...) was set to false, I could still craft a URL like:http://example.com/?page=admin/index
which expanded to http://example.com/index.php?page=admin/index where the real admin was at http://example.com/admin/index.php and offered complete access to the backend without authentication or authorization - let alone other files in the file system.
In another project, I found that the server had register globals turned on, and therefore could craft a URL like:
http://example.com/admin?valid_user=1, where valid_user was a PHP variable set to true iff their session cookie could be authenticated in the database.
I think it's terrifying that these things still make it through to production websites
location ~ /\. { return 403; }
My question - do I need to put this once at the top of my configuration file and all it good or does it need to go into multiple places in the nginx config?
It would be great if there was a simple, universal way to say to nginx "don't serve hidden files from anywhere under any circumstances".
Other solutions just seem hackish to me, but every project is different I suppose.
I don't get it either.
So far a few people have requested my .git/ directory but none have attempted to plunder the riches they think they'll find within.
(Is JSON an abstraction? Not really, as it's a concrete spec, but its general kind of serialization format is an (incomplete) abstraction.)
apt-get install zutils
zgrep -r "\.git" /var/log/nginx/Wouldn't be terribly difficult to write a script to crawl it. Git's format is well known and documented.
I wish servers would be configured so they don't server ^\..+$ files by default. I wish servers would behave as secure as possible then it's up to the developer to whitelist features rather than the other way around.
What are the best practices here? What should operators that currently deploy this way do instead?
This simplest solution is to add something like ' location ~ /\.git { deny all; }' to your nginx config (or just hide all paths starting with '.'; you rarely want to serve hidden files).
The most important practice for remediating this is to have a very clear disconnect between deployment artifacts and development repositories. You want to have your developers write code, and then some preprocessing step to distill that repository into the minimal deployment artifact. This artifact then goes through your integ tests on a gamma stage and onto prod. Whatever deployment system you use should know what git commit this artifact originated from, but the artifact should not have that information on its own.
If you want the really poor man's version of this, just have a build system in your repository which populates an "out" directory with everything you actually want to deploy, and then rsync that sucker around.
If you want a hip answer to this, Docker has the "Dockerfile" which will specify all artifacts that should be added into it, and should serve as a distillation of what's needed to run your application.
I'll point out that this issue won't affect many types of sites in the first place. For example, if you use any sane rails setup, you already have the "public" directory which you use as your static files root, and then you proxy to the unicorn/rack/whatever process which will never serve such files.
I would wager that the majority of these sites are poorly configured apache2 + php messes since php made the massive mistake of having the filesystem act as routing; any web framework with good routing (e.g. rails, django, even sinatra, web.py) will not suffer from this without going out of your way a bit.
It depends on what sort of platform and server you're running on but just do what you have to do so you're not serving your .git.
Like this:
git --git-dir=/foo/bar.git --work-tree=/foo/bar.work init
git --git-dir=/foo/bar.git config receive.denyCurrentBranch ignore
Then create (and chmod +x): /foo/bar.git/hooks/post-update #!/bin/sh
work_tree=/foo/bar.work
GIT_WORK_TREE=$work_tree git checkout -f
Then you just create a symlink to the work tree for your website root, or put the work tree there, or whatever, depending on preference. Can't say this is perfect, but it works pretty well for smaller projects.1. Configure your webserver to hide common directories and files.
2. Don't store passwords / credentials in version control. 2a. If you use a test suite, add a test to verify this.
3. Have a make step that builds the deployment content into a separate directory. (e.g. a gulp deploy task)
4. Rsync the deploy content to your destination server.
Instead of 4, more complex systems (with multiple servers, libraries, etc) use Docker to build an image from your deloy content.
That is:
- completely non-obvious and unexpected
- a terrible idea
Why, instead of trying to figure out how to avoid handing this magic file out to everybody, are we not trying to fix it so that no such magic file need exist?The error here is uploading sensitive or hidden information to a web host into a public directory, not how it is stored locally.
If you use the root of your app, including source code and hidden files, as the public directory of your website, one permissions error means all sorts of things might be exposed, e.g. other dotfiles would also be exposed, and potentially all of your source code too, because you're relying on the web server to hide it somehow in every instance. That's the problem that needs fixed here (exposing the wrong files to public root), not that one particular hidden folder exists.
The root of your project should contain nothing more than build documentation/scripts and other developer/user notes & scripts (which for a web application could be as simple as "expose the subdirectory call "public" via your web server" but could be much more for more complex applications that have larger build requirements).
Ignoring this long held recommended practise (which less experienced developers might not be aware of): git originally came from an environment where you couldn't simple expose your repository as your application. The Linux kernel and other projects needed building from source before they could be put into production. So this is in part due to people using a tool in a new context without sufficiently thinking about the possible implications (that the tool designer, thinking about other environments, might not have considered). Security requires a lot of "due diligence" like this unfortunately: you can't expect the tool designer to be aware of all the potential security considerations in your environment, you have to deduce and mitigate them yourself.
It's non-obvious and unexpected if you have never used a VCS and didn't read a single page describing git. In which case almost everything in git will be unexpected and non-obvious (so will be any programming language or technology).
But I agree that the problem shouldn't be trying to avoid handling out the .git dir to everybody. The problem is - git clone in document root of the webserver for website deployment is a terrible idea and was never a supported usecase.
edit: so I came the third! Should try typing quicker.
But let's say we agree it's an issue and should be fixed. What alternative solution would you propose? You can't just get rid of the "magic file" because it actually contains your version history, the thing you wanted to track in the first place.
Also ... webservers should make it very hard to offer up .dotfiles as webcontent.
The model of exporting a Unix directory as the structure of a website is barely good enough for static sites (you'll get into all kinds of problems with URL management), and is completely unsuited for applications.
Now, of course, PHP was created as a tool to add a visitor counter at the bottom to your pages. With a bit of caution, it's indeed secure enough for that. Nowadays people create huge applications using the same security model, and PHP developers don't even think about changing it.