Don't publicly expose .git (2015)
en.internetwache.org
en.internetwache.org
A much better approach (or at least, what I use) would be to set up the repo somewhere private with --bare and set a receive hook to checkout HEAD to the htdocs folder, this way the htdocs only has the content and you get the extra feature that you can sneak extra commands on the checked out source (such as building/minifying) without changing the original source
> Another approach is to use git’s --git-dir and --work-tree switches to move the git repository out of the document root.
Yet another option is to make the htdocs directory a worktree of the git repository, which doesn't require passing flags around, or setting environment variables. Technically, this still leaves a .git, but it's only a file containing the actual location to the real .git directory.
If your history has sensitive info, see about rewriting the history. If that's not possible, maybe fork the repo, remove the sensitive info, and get the team to switch to the fork. If that's not possible either, make the sensitive info meaningless (reset your DB passwork, revoke the API tokens, etc).
(When building systems for clients, this is something I stress. "We should be operating as if it is assumed that an attacker has a VPN into your network space and has nmapped all your stuff.")
a) how do you search a repository's history, along all branches (including undeployed development branches, which may nonetheless contain production secrets) to find secrets mistakenly stored in version control? b) what happens when somebody stores a secret which isn't easily rotated (because legacy systems have hard-coded the secrets etc.)? How do you deal with trying to rotate secrets which aren't well-managed, because the real problem isn't that your teams are storing secrets in source control but that you don't have proper secret management set up across your organization?
The simpler (and more correct!) way to deal with this is to stop using version control for deployments and to start using proper package management and deployment tooling. Version control is not designed as a deployment tool; it is a poor replacement for proper deployment tooling; and teams which think they only have hammers to hit what are not nails but screws, need to learn that sometimes they'll have to go out and get a set of screwdrivers too.
And don't just stop with .git: Delete any folder/file that's not required to operate the app in production.
We do something similar. "Clean and minimal" is a good strategy when you're deploying assets to publicly accessible systems.
Step one: stop randomly smearing crap around. Prod should only have files that came from a .deb or .rpm signed by the legit build process, because that's how you know your system is reproducible and has everything it should and nothing else.
It's also solved the exact same ways, by scripting your stuff on one level or another.
If the deploy process is "git pull", you're praying that Joe McGee didn't push some untested crap to the relevant branch 5 seconds before you deployed.
I'll argue with "that's how you know your system is reproducible and has everything it should and nothing else", though. To get there, you need to look at immutable infrastructure, where you're building a new container or VM image for every deploy. Otherwise you might have libfoo installed on the app server, despite your app dropping support for foo 2 years ago.
You can use git for deployment but indeed you must have a clear commit policy to do that (and it's a good idea anyway to push only major curated versions on master).
You can actually use git in production, properly.
I personally like to mount gitRepo volumes on Kubernetes, and have my CI pipeline automatically update the revision of the deployment whenever it validates tests. Then I have kubernetes roll out the update automatically.
shameless plug: I've developed a service that you run to check against vulnerabilities in your apps/servers and it has a free plan (https://my.gauntlet.io/registration.html) in case you're interested (https://gauntlet.io).
For example, static websites for open source projects, et al.
> It seemed like an accessible git repository was intended on some websites - mostly open source projects where the website’s sourcecode is available online.
> On the other side, we had to hold our breath when we noticed that more than 100 projects used HTTP-Authentication for server-client communication. That means, that the protocol://user:password@host/repository combination is saved in the .git/config file, giving attackers access to the users (companies) GitLab-instance or GitHub/BitBucket account. With a bit of luck an attacker gets access to the CI-Server and then runs malicious code to further compromise your infrastructure.
I just rechecked with [1] and you seem to be right.
I'll update the blogpost in a second. Thanks for the hint!
Not that you should have .git exposed on your public webserver anyway. I do remember participating in a CTF that had a problem like this a few years ago, it's possible that it was the same one the author mentioned.
Hasn't dvcs-ripper [1] been around for longer? It supports other VCSes as well.
Also, the article fails to mention that a simple `git clone` would usually work as well, although that tends to be blocked in similar CTF challenges.
I knew about dvcs-ripper, but thought that implementing another variant might be fun and let me learn about git internals.
Does a simple `git clone` really work? I just tested it and it failed:
``` $> git clone http://x.domain.tld/ fatal: repository 'http://x.domain.tld/' not found
$> git clone http://x.domain.tld/.git/ fatal: repository 'http://x.domain.tld/.git/' not found ```
And yes, the post's background is a CTF challenge that blocked a simple `git clone`.
Parts are even in metasploit.
here's one of the blogpost's authors. Although it has been a while since we published the blogpost, I'll try to answer any questions or listen to any suggestions.
So if I post my website's sourcecode on github, I'm equally vulnerable? I could see problems if said checkout contained a credential cache, but that doesn't seem to be mentioned.
But entries recorded in my reflog are not pushed to or fetched from any remote I interact with. Instead, their reflog is updated when I push (to record that my commits were accepted & their branches/tags changed), and my reflog is updated when I fetch (to record that their commits were accepted & my branches/tags changed).
Or what about helpful developer that checked in some secrets which are visible in the repository history but not the current checkout?
Or that stupid PHP thing where 'config.php' and its MySQL passwords are world-readable, but rely on the web server interpreting it as a PHP script due to its file extension to prevent secret leaks.. not so valid when a copy of the script if available as ".git/objects/00/cf74f2066b0c72a4c4b2a24ef116f1fd23df42".
But of course, even if these weren't problems, the original point still stands: there is no guarantee .git doesn't contain secret data (such as username:password) either now, or into the future, so exposing it is a bad idea.
> there is no guarantee .git doesn't contain secret data (such as username:password) either now, or into the future, so exposing it is a bad idea.
The same can be said of the HTML and images, so I don't find it a useful heuristic. Note that I was disputing your claim that a username+password used to fetch a repo over http would leak into the remote's reflog.
$ git clone https://....:x-oauth-basic@github.com/dw/csvmonkey.git
Cloning into 'csvmonkey'...
remote: Counting objects: 340, done.
remote: Compressing objects: 100% (27/27), done.
remote: Total 340 (delta 19), reused 27 (delta 10), pack-reused 303
Receiving objects: 100% (340/340), 138.93 KiB | 0 bytes/s, done.
Resolving deltas: 100% (212/212), done.
$ cat csvmonkey/.git/logs/HEAD
0000000000000000000000000000000000000000 c9d566bf167dcf3556008df58be37c4a27ff5062 David Wilson <dw@botanicus.net> 1497289486 +0100 clone: from https://....:x-oauth-basic@github.com/dw/csvmonkey.git
If you perform a Git checkout on a web server e.g. as part of an Ansible script, and you embedded secrets in the repo URL (common enough, believe me), then that secret is readable per above.FWIW this isn't some unbelievable theory or hypothetical scenario, I've seen plenty of Ansible setups like this and found domains with this exact problem in the process of writing http://pythonsweetness.tumblr.com/post/52587443706/devs-plea... a few years back
You might not remember me. I'm the poster you're responding to. How have you been? Me, I'm all right.
I was just thinking of when we first spoke… it seems like so long ago! I remember it as clearly as yesterday: you had made a partially-conherent argument that the auth creds for a git URL could leak into a remote deployment's reflog! Oh, how we laughed, and our amusement doubled in size as you fancied a implausible situation where the read-only deployment credentials could be recovered from the very same repo they allowed access to!
It was much later when we crossed paths again, but your talent for sharing inventive tales had not waned in the slightest. For this next performance, you regaled us with the simple truth that no person can be certain that their commit history will not reveal their darkest secrets, and thus should strictly eschew sharing it in a public place; but that the contents of their index was above suspicion, and could be shouted to the world without a moment's thought! Many of us stumbled to determine what byzantine process made the working directory automatically scrub itself of secrets, before finally the jape dawned on them.
I eagerly anticipate our next encounter; what fresh new hilarity will you share with us?
I hope my restatement of my understanding of your position helps make my position clear,
--falsedan
> git clone https://username:password@github.com/...
This command produces a new git repository by cloning the supplied URL
> will end up in git reflog
The newly generated repository's reflog will contain the credentials passed on the command line.
> so yes, it's a problem
Assuming the newly generated repository also happens to be a static HTTP server root, which is the subject of the thread in which you've been posting
I mean, if you're tracking config.php in git, you're already doing it wrong.
So if your deploy process is an anonymous git checkout of Drupal, there's no credential leaking to worry about.
> no guarantee a future feature of Git that you've never heard of doesn't make the problem worse
Okay.
/mysite/.git
/mysite/mysiteroot/index.html
/mysite/.gitignore
/mysite/lib/common # deny access to HG and Git repositories
location ~ /\.(hg|git)/ {
deny all;
}These type of snarky responses discourage newcomers to participate in discussions. I have seen this happen to many people, so please dial back the snark.
Hey! That's not very kind to disparage everyone using a text file in a git repo to manage their passwords/keys.
Putting keys into a git repo is fine! But be careful when publishing that repo, as something you thought was private could suddenly become public.
If your code is public then what does it matter? If it's not then you should be protecting it like any other sensitive information.
You can unintentionally expose your repository by deploying .git by mistake.