Github could have a setting like "flag unsigned commits from me after $date" that allows this kind of policy to be communicated.
It's like someone is telling you that they always schnarkel their schnops with shaush, now that the cost of schnacking a schnill is a mere $75k.
After however many days, I pretty much understand all of what joeyh is saying, except maybe for how the social engineering part would work. Thanks Hacker News.
This means that a SHA1 collision on the tree object, or any blob object, still results in a valid GPG signature.
So git GPG doesn't provide any additional security over SHA1; it only provides proof that a certain person signed a SHA1 and the commit message.
"tree 33191145e9a307b7ca5c6e4aa6e74a4904426cba\nparent a2c6664cf085720cba3f92e4a7348c46369a869e\nauthor Greg Morenz <morenzg@gmail.com> 1459911016 -0400\ncommitter Greg Morenz <morenzg@gmail.com> 1459911016 -0400\n\nthe commit message\n"
git ls-tree -r HEADLinus Torvalds http://git.661346.n2.nabble.com/GPG-signing-for-git-commit-t...
NO ONE i know does this. Conversely, everyone expects to be held, and hold others, accountable for the contents of their commits.
We don't have second preimage attacks on sha1 yet, but thinking of keeping signatures separates only buys you flexibility (which is not used by most projects currently on github) and maybe a bit of performance, for a process that is much more brittle.
Also, with things like this:
> It also doesn't add any real value, since the way the git DAG-chain of SHA1's work, you only ever need _one_ signature to make all the commits reachable from that one be effectively covered by that one.
He completely misses the point: who's going to check after pulling that the whole history of dozen/hundreds of commits is legit? That might be the case for the Linux kernel gatekeeper, but not for one of the several small contributors to any project.
Let's not do some "appeal to authority" BS with a myopic post from 7 years ago.
> He completely misses the point: who's going to check after pulling that the whole history of dozen/hundreds of commits is legit? That might be the case for the Linux kernel gatekeeper, but not for one of the several small contributors to any project.
Git does that automatically for you. Git self-verifies the DAG (which is why it takes a while to warm up on the kernel repo because of the insane number of commits). So he's not missing the point.
1) It's unclear what we should sign on a commit, therefore we shouldn't. I'm not sure what to make of that.
2) Since a key may be revoked, any signatures should not go into the commit. Presumably Torvalds thinks that tags should be deletable, but commits should not, or something.
3) That if all commits are signed, then it become automated, and therefore somehow not worthwhile. Meanwhile, automated signing seems awesome because we know that at least the signer had possession of the key.
I can't really extract much more logic out of that post, and none of it sways me in the slightest.
This isn't about using GPG siged vs nothing, it's "GPG signed in the commit" vs "GPG signed commit and metadata".
The bit about revoking keys is also interesting. He's pointing out that if you keep the signature in the commit, you have permanently bound that particular signature with the commit data. It's mixing two different types of data into one storage location.
The alternative is using tags, which shadow the main branch. This way you have the option of using those tags, or you can create new tags (with new signatures) with a new key. This gives flexibility for the future.
> insults
I don't see Linus insulting anyone in that post. Explaining the problems with a method is not an insult.
> therefore we shouldn't
Is that post really that hard to read? I find it obvious that he is saying you should use the builtin signing with "git tag -s".
[1] It's worth observing that this is a comment from 2009; I'm curious if he has changed his opinion on SHA1 in over the last 6 years.
[2] -s, --sign Make a GPG-signed tag, using the default e-mail address’s key
Fuck that shit.
The other thing is that the SHA1 weaknesses are collisions. It's not likely to be possible in the foreseeable future to create a commit that matches an arbitrary hash - to generate colliding commits, the committer would need to generate the good and evil commit at the same time.
Before SHA1 is broken, you can't change things that I have committed, unless you have access to my account. If a SHA1 preimage attack becomes available, you can. If a SHA1 bithday attack becomes available (for the low low price of $75k, croudfund one today!), you can, by sending me some particularly evil forms of merge requests.
Actually, no. You need a SHA-1 second-preimage attack, which is vastly more difficult. However, even if there was one, it would cause git pull to fail, as downloading the objects would likely require terabytes of garbage data.
Consider the case of taking the SHA1 of a message only 1024 bits (128) characters long (shorter than a tweet, let alone this reply). There are 2^1024 such messages, but only 2^160 hashes. So on average, there are 2^(1024-160) = 2^864 ~= 1.2x10^260 messages of length 1024 that must share each SHA1 hash.
It would be a very interesting result if you could prove that you had some 1024-bit message that didn't share a SHA1 hash with any other message of length less than 1TB ~= 8x10^12 bits, and this property couldn't be true of very many 1024-bit messages, because this means you "take away" from the hash space of messages of length 1032 to length 8x10^12. If you took away just 2^140 out of the 2^160 hashes, for instance, we'd readily see hash collisions in those in-between sizes, because there would just be 2^20 hashes left. But even then, this special property you imagine would only apply to 1 in 2^884 ~= 1 in 1.2x10^254 of all 1024-bit messages so you'd have trouble finding one of them in the first place, and more to the point it's almost impossible that your victim would have selected one of these hashes that would require your hypothetical terabyte file to collide the SHA1 hash.
In theory, nothing I know forbids the possibility of finding an algorithm which yields a 160-bit solution within a few years. In practice, such an algorithm is tremendously more difficult (and potentially impossible) to find than simply finding a SHA-1 second-preimage attack, which is already vastly more difficult to find than a preimage attack, which is also vastly more difficult than getting a non-bruteforce collision.
It would certainly be a phenomenal achievement, however. We still don't have a practical preimage attack on MD5, 10 years after the first actual collision found…
Regardless, we're only now beating the first challenge after 20 years. It is likely to take more to go from there to a second-preimage attack, and more still to have it be a 160-bit solution within a practical amount of time. git will double its lifetime multiple times before then. Something shinier will have replaced it.
I'm not sure "terabytes" is justified either, but the point's a good one.
Apart from that, half the patches I prepare go out via git format-patch rather than a push and a pull request. So I'll need to confirm if the signature survives that workflow, and if it produces unwanted noise or other effects on development mailing lists.