A Git Horror Story: Repository Integrity With Signed Commits
mikegerwitz.com
mikegerwitz.com
It's not trying to describe a "ground-shaking vulnerability". There is a current, known issue with display of git signature verifications, and he's including it in his article. That's not the point of the article.
The point is an exploration of options for securing your repository history. That's it.
But then again, I'm just an old grumpy man who is much faster to criticize and be negative than he should be, so who am I to cast judgement? Do I ameliorate or exacerbate the situation with posts like my previous one or this one? I don't know, maybe I'm part of the problem in that way. Either way, I much preferred (so are my motivations selfish? Is that good or bad? I think good, but how do I know?) the old HN when submissions like this one would be labeled for what they were.
Please notice that this article is segregated from the rest of my blog with no glaring/obvious links to it. This was intentional. I posted this article in the hope that it would be useful to others, not because I would gain any benefit from doing so.
I will take your comment into consideration for future articles.
Linus himself is against GPG signing commits.
http://git.661346.n2.nabble.com/GPG-signing-for-git-commit-t...
Ultimately a human being has to be responsible for reviewing code. Once it's reviewed, you can tag and sign it. No GPG signature is going to make up for not reviewing your code.
Various strategies are used to specifically protect all kinds of sensitive systems far beyond simple network security. This article is a detailed exploration of a/some strategy/ies for protecting your git repository against malicious modification -- a scenario that is rare but far from unprecedented.
It's not as easy as using tool x and generate key y... its more than using software, if you want security, you have to think about the whole stack, down to the access rights that each part of your hardware has.
There are good and best practices, but with IT Security, using the best practices in one place and just plain forgetting to unlock you PC when you go to the toilet at the office is more than enough...
People are allowed to be afraid when asked security questions.
You're not even trying. Physical access, game over, the lock is irrelevant.
You fear what you don't understand, and you don't understand because you're afraid to learn. You have simply crafted a vicious cycle for yourself, and are making the world that much more unsafe by spreading your fear and acting as if everyone should share in your fear.
Yes, once someone has hardware access, its devilishly hard to secure an environment, but just stating that its impossible is as wrong as security systems that ask their users too many questions to which they most likely have no good answer/no doubtfree answer.
There are systems which work pretty securely, i have worked with several educated security experts to create security infrastructures in companies.
I'm not saying that i don't understand the problems or don't want to learn solutions to them, i'm trying to defend the position of the user that can't be asked to learn about all the caveats of it security, since it is one of the most complex problems in computer science.
No, not every front-end developer that has to check-in his html und js files can be asked to understand all principles of it security.
In fact what we do in the kernel community is most patches are in fact reviewed via e-mail. This increases the likelihood more people will look at the patch and notice any potential problem, and more importantly, the e-mail archive introduces an additional external record of the proposed change. After each patch in a patch series is reviewed, in isolation, with each change small enough that it can be easily reviewed in a mail reader, and with especial attention made to maintaining bisectability, the maintainer who will ultimately by signing the git pull request will apply it into his or her tree.
For larger changes and for larger subsystems where we have submaintainers, it is the maintainer's responsibility to make sure the submaintainer sends a signed pull request. At that point the submaintainer is assuming full responsibility for the set of patches which he or she is sending to the upper-level maintainer, just as the maintainer assumes responsibility for the patches he or she sends to Linus.
In the case where the submaintainer does not send a signed pull, or the maintainer does not have full trust in the submaintainer (either for social or technical reasons), it is the responsibility for the maintainer to check each and every commit for correctness. If he or she can't do that, he shouldn't accept the pull request; instead, he might request that the patch series be sent to the mailing list so everyone can review the patch series.
One of the things which follows from this is that for me as a kernel or e2fsprogs maintainer, github pull requests are not particularly useful. Most of the time I have no idea who the person is that initiated the pull request, because the primary community is on the mailing list. Worse, the pull request only gets sent to maintainer, which makes the maintainer the review bottleneck. If the patch series is sent to the mailing list, as individual commits, with no requirement to jump to a browser, it allows the review workload to be distributed to other developers on the mailing list.
As far as I'm concerned, a github pull request is the same thing as a patch sent via private, personal e-mail. I'm going to review it very carefully, and only apply it if I know that it is clean. Typically the review involves extracting the commit as a diff, and as has been discussed in other places, github doesn't encourage properly formatted commit descriptions, so I end up needing to extract the diff, and then reapply it with an appropriate commit description. So for me github pull requests generally represent negative, not positive value, because of the quality control and quality assurance requirements of the code bases in which I work.
Thank you for the detailed explanation; it did help to clarify details I assumed, but was uncertain of.
The person requesting the pull signs the pull request by signing the git tag. The GPG signature on the git tag covers the complete history of up to that point in the tag, so there's no need to sign each and every git commit, as you propose. When the upstream maintainer (or Linus, at the top level) accepts the pull request, git will automatically create a merge request that contains a copy of the git tag's GPG signature in the merge request. (After this point, the git tag can be discarded). When Linus releases something, he signs a git tag for the overall git repository, and everyone can check that.
So for example, when I am satisfied that everything in the ext4 tree that I want to push to Linus is OK, I create a signed tag named "ext4_for_linus", and send him a git pull request, which is created using the "git request-pull" command. This creates an e-mail message contains the URL of my git repository and the name of the signed tag. When Linus accepts my pull request, the resulting merge commit will contain a copy of my GPG signature from the git tag. This signature signs my name to the commit id of the git tag, and the chain of SHA-1 hashes inherent in git guarantees that the entire commits back to the beginning of time came from me. Of course, my changes will branch of something like v3.3-rc2, which is in Linus's tree, so what is being signed that is of interest is my new changes between v3.3-rc2 and the tip of the ext4 tree.
Once Linus has accepted my pull, I'll reuse the ext4_for_linus name for the next GPG-signed tag for the next 2-3 month development cycle, since the important bit (the GPG signature) is preserved for long-term posterity in the merge commit. But what's important to note here is that the signature in the merge commit is signed by me, since it contains the contents of what I pushed to Linus, even though the merge commit was created by Linus when he accepted the pull request by merging in my changes into the repository. What Linus signs is the GPG tag for the actual released kernel (i.e., v3.3).
See? So there's no need to refer to the mailing list to validate the pull request or to establish after-the-fact accountability. The mailing list is just used for review purposes. But of course the review is a critical part of the acceptance process.
Ah, I see. Thank you for the clarification.
It was not overall integrity of the trees that I was bringing into question. While the code review process may be a problem for certain organizations, I am not doubting that process for the kernel. With such a process in place, signed tags are likely to be sufficient, yes.
Note, however, that one of the points I was trying to make with the article was proof of authorship (to prevent someone from committing as you, much the same way as you would want to prevent someone from fraudulently sending an e-mail or message through any other means). It is for this reason that the workflow you described is not a complete alternative, but would be an excellent means of augmenting the review process for any project.
That said, author identity is not a concern for most projects, in which case, based on the quoted portion of your comment above, my arguments do not necessarily apply (assuming that all commits that are not yet ancestors of a tag are ancestors of a signed merge, as mentioned in your comment).
Might I ask how the maintainers deal with issues of identity? If they encounter a commit in your pull request that appears to be from a well-known contributor, do they verify his/her identity before merging, or do they not care?
But that's only important from the perspective of knowing that a particular public key belongs to a person, which in turn is only really important when a maintainer wants to know whether or not the person who signed the git tag really is the submaintainer who is claiming to have signed the git tag.
Most of the time, we don't really worry about identity; what we worry about is the correctness of the commit. Especially for patches being sent via e-mail, I really couldn't care less whether the identity of the person sending the e-mail was Christoph Hellwig, or his evil twin. I'm still going to review the patch very carefully, because we all make mistakes. That's my job as the maintainer; to ultimately be a crap filter, and to reject crap, no matter whether it's written by someone famous or some nobody.
So when I send a pull request to my upstream maintainer (which in the case of ext4 is Linus), he really doesn't care about the identities of the authors in my pull request, because it is my reputation which is on the line; I'm the one who signed the git tag pointed to by the pull request. Just if it gets by me, and then it gets by Linus, it's ultimately Linus's reputation which is on the line.
So that's why your question is a bit ill-formed. Someone who is above me in the maintainership hierarchy isn't really going to care much about the identities of those who are below me in the maintainership hierarchy. The trust relationship is fundamentally between me and Linus, and then I have trust relationships between me and those developers below me. So Linus doesn't care about the identities of those which are two levels below him in the maintainership hierarchy, so long as he trusts me to have done my job properly. Of course, if Linus discovered that I had been blindly accepting pull requests from my downstream without doing proper quality control, then he'll yell at me and he might refuse to take future pull request from me (or at least threaten to do so). But normally, there's no need for him to verify the identity of the authors of commits in my pull request; he only needs to verify me, and to verify that he trusts my technical judgement and quality control checks and processes. If decides he doesn't trust me or my technical judgement then he won't accept pull requests from me; and in that case he would be absolutely right to do so.
Security is only as good as the weakest link.
Also, don't make it look like it's more secure when it may not be
Wouldn't the signing of all commits as they are committed solve this problem? (ie. rather then trusting Author information from the commit, trust the signed-by information to give author information?)
The GPG signature cannot be forged (access to the private key is needed).
git log --pretty="format:%H$t%aN$t%s$t%G?" "${chkafter:-HEAD}" --first-parent \ | grep -v "${t}G$"
What of the case where some smartass decides to
$ git commit -am 'Yet another foo G'
So long as it doesn't have a bad signature, it will give an empty string, right? So how do I know where the G came from?
That said, it could be terribly confusing when viewing the output. For that reason, perhaps another character after the tab (before the %G?) would be a good addition.
Foo\tG
This would result in the following output: ... Foo\tG\tWould you not end up with
foo\tG
\t
And break the regex? Or is there absolutely no way to smuggle a \n into the message?Otherwise, you would be correct --- that would have rendered the check useless.
Actually these issues are not even specific to VCS, but security of most digital information.