Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?
Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?
A SHA-256 sum, though very good, only assures you with great confidence that you're looking at the same thing you looked at before, or that someone else is looking at elsewhere.
It is not a digital signature, and we don't want digital signatures to serve the role of content hashes.
Speaking of signatures, we have support for them in Git; you can use gpg to sign commits, and set it up to be done automatically.
Nobody is going to fake your commit such that the fake has the same SH-1 hash and your GPG signature.
The worry there is that the key holder (whether the legitimate one, or a malicious party who got a hold of the key) somehow does this: creates a new commit, signed with their key, which somehow has the same SH-1 as an existing signed commit. The git hash includes the GPG signature, so there is a significant layer of difficulty there which is likely harder than faking an unsigned SHA-256 commit.
The attack is I pre-author `Makefile => foo: echo "hello"; bar: echo "world"` along with `Makefile => foo: echo "hello"; bar: rm -rf / ; /* $ELDRITCH_SHA1_SPIRITS_GO_HERE */` that both hash to `ff1234...`
I then prepopulate the repo with `echo "hello"`, wait 6-9 months, then submit a commit for `echo "hello" ; echo "world"` and keep (in my back pocket) the alternate implementation that also includes $ELDRITCH_SPIRITS to force a collision and MY predetermined change in functionality.
I then have free choice as to whether I serve them "hello world" or "hello && rm -rf", and THAT's the plausible problem to avoid: the ability to "cloak" content anywhere within the repo if you have enough $ELDRITCH_SPIRITS and GPU's.
You have _really_ good points, but are woefully confused. The proper answer is (would have been) to include `tree ff12354...` along with `tree-sha256 abc123456789...` for another 20 years along with a `[git.hash_strictness]: default/lazy/strict`, and some oddball `git-rerere` type packfile extension which lets you map `sha1:ff1234... => sha256:abc123456789...` "transparently" rather than the horrific situation you're laying out (correctly!) that forks the ecosystem in to "longhash" and "shorthash" when most repos don't even care in the end.
Yes, I didn't understand that the GPG signing just operates on the top level object in the commit and trusts the SHA-1 hashes contained in it.
The signing process doesn't recursively traverse the bytes of the commit to pull them into GPG, like you would expect.
It's like, imagine you made a "bill of materials" of your project's files consisting of their names and CRC-32 checksums, and then signed this file, and called your project securely signed, LOL.
This aspect can be fixed without foisting new hashing scheme into the content tracker. In fact, it must be fixed; users on SHA-1-based repos deserve secure signing.
It's really sneaky that the SHA-1 business (not intended to be a security mechanism) was embroiled into the signing implementation; that GPG is demoted to the strength of SHA-1.
Was that just to save some cycles? It's certainly faster just to sign the commit object!
No I wouldn't expect this. If by recursive you mean traversing the entirety of git history, that would be prohibitively expensive performance-wise (imagine rehashing the entire multi-gigabyte history of the Linux kernel every time to sign and verify a commit) and destroy git functionality such as shallow clones and blob-less clones.
If you by recursive you are only referring to the current working directory, as I lay out here https://news.ycombinator.com/item?id=49930048 it doesn't work. Indeed, without attestation of parent commits, a malicious attacker can actively frame any pre-existing vulnerability as someone else's handiwork by simply inserting a new commit that purports to be the parent of another commit that does not contain the vulnerability, which then makes the original commit look like the source of the vulnerability.
"I didn't introduce the vulnerability, he did! Look I can even prove it with my signed commit!"
I misspoke. It's worse. "I didn't introduce the vulnerability, he did! Look I can even prove it with his signed commit!"
You can produce two commits that have the same SHA-1 and hash them. One has the vulnerability and one doesn't.
Then someone else bases a new commit on the one with the vulnerability, which has yours as the parent. You then swap in the one without the vulnerability and now it looks like the child is introducing that change.
Their only defenses are (1) plausible deniability (we don't trust parent hashes) and (2) still having the original commit they based theirs on to show that a swizzle took place and that the real original already had the vuln.
If you are arguing that CRC-32 is a cryptographically secure hash function you might.
SHA-1 is (or rather, used to be) one, so you actually can. This is literally how git works.
> It's really sneaky that the SHA-1 business (not intended to be a security mechanism) was embroiled into the signing implementation; that GPG is demoted to the strength of SHA-1.
What's sneaky about it? It's a completely reasonable design decision, making git orders of magnitude more efficient than the counterfactual you're arguing for.
> Was that just to save some cycles? It's certainly faster just to sign the commit object!
Sure, "just" read potentially gigabytes of data, potentially over the network, and send them through your cryptographic hash function as opposed to just the objects you're touching in your commit. Absolutely the same effort.
We can round up the bits that make up a commit in a SHA-1-based repo, and sign those bits securely; this is a thing that is possible.
Regardless, commit hashes should be a security mechanism IMO. And not just commits. I should be able to treat _any_ content addressing system as having secure addresses. If you can engineer collisions you need to patch your system.
(Note that the above does not necessarily imply support for unconditionally forcing a fork of the entire git ecosystem.)
Not "must", but it would be stupid to use two sets of hashes without a compelling reason.
(Inside GPG, there are configurable choices. It's possible to be using SHA-512, so in a SHA-256 git repo, you can still be using two hashes.)
For anything else, it depends on the exact scheme:
If git sent GPG the bits for the current commit specifically because it expects a better hash to be used, that would be silly to design on purpose (git should use a hash that meets all requirements). And it wouldn't protect history, which matters.
If git rounded up the bits of history, that would perform unacceptably badly.
If git used a better hash to secure history for signing purposes, and still had the normal hash elsewhere, that would be very silly to design on purpose because it should use the better hash for everything.
In our reality, the git hash is load bearing. You can disagree with that design choice, but you can't pretend to live in that alternate reality and design your solutions for this reality according to that.
> can't pretend to live in that alternate reality
Discussions about solutions that don't exist or requirements not implemented are everyday occurrences and necessary.
I mean, if you're talking with a contractor about how your bathroom should look, that is a "fictional alternate reality", but you're not pretending to be living in it now.
I don't understand the purpose of the above engagement style, but it doesn't look like a great fit for HackerNews.
Yes but as my other comment explains, this is not particularly useful in and of itself.
you're arguing a position which indicates you don't know git's physical (textual) commmit structure.
Spend some quality time with:
git cat-file -p HEAD
You'll get something like: tree ff1234...
parent ccddef...
fix(BUG-1245): my ai fixed it
Continue to `cat-file -p $TREE` and you'll get: file efef12... Makefile
tree a1b2c3... src/
file f0ea12... README.md
...and then it's turtles all the way down. It is (was!) safe to sign $HEAD (and only head!) because... it's turtles all the way down. Signing HEAD attaches IDENTITY (eg; torvalds@linux.com) to CONTENT (eg: src/@a1b2c3...) and TRANSITIVELY all the way down.Your homework is to go run:
time (
git ls-files | xargs -n 1 sha256sum
) | sha256sum
...and then make a commit (nee: tag/note) containing that content as the commit message.Your INSTINCT isn't wrong, your mechanics run counter to the practicalities of how git is designed to work, and the practicalities of the crucial defense that git-core is trying to divert: the ability to arbitrarily alter the signed(!!!) past, signed by third parties, with $ELDRITCH_SHA1 attacks.
You _still_ have to trust that GitHub.com or kernel.org won't get popped and start erroneously serving "signed but faulty" files and trees, but the urgency of moving to sha256 is about preventing faulty commits in the first place, which removes the requirement of "trust me bro!" relationship with the serving provider (or MITM).
That is false; hopefully everyone here understands that it's a graph linked by hash "pointers", such as SHA-1. I have not seen evidence otherwise.
> time (
Of course I understand that it's (possibly vastly) more expensive to digest the content than just a top-level node. So what? Maybe someone would pay that, if they had the option.
Yes, but git's signing scheme(s) do, as do many third-party ones, so what are you arguing for, exactly?
The non-necessity of a fix in an alternate reality in which nobody uses git as documented? One in which git launched with a big disclaimer of "never trust our cryptographically secure hash function to be cryptographically secure" in its documentation and CLI outputs?
Because one thing is certainly a big no-go in our reality: git walking back these explicit and implicit (Hyrum's law) contracts because fixing a load-bearing hash function is hard.
You refer to PGP signed Git objects, but you also argue:
> Git hashes are not supposed to be a security mechanism
Guess what the Git PGP signature is signing.
I don't really have the crypto chops to declare a fact here, but I have a speculation or intuition. In this day of supply chain worries, I think a proper signing algorithm should not be signing this tower of hashes, or not just this tower.
It should incorporate a canonical stream of all the actual commit content. It is the integrity of this content from the author's working copy that they can and should attest, not some derived byproduct of the storage scheme. Edit: Of course, I mean a secure hash of this stream, not a signature including a copy of the entire content!
My intuition is that the content-addressable store is used to reconstitute the commit content, but the verification should be over the original content, not the internal addressing of the store.
Wouldn't this make it harder to do these exploits? You would have to find alternate content that simultaneously produces collisions in the internal addressing hashes and for the overall canonical stream hash.
If you also carry size info alongside each hash, would this also make it much more difficult to produce useful collisions?
The GPG signature signs some kind of hash calculated over the commit, minus the GPG header, which is thereby added.
The git hash is then calculated over the whole thing. The git hash is on the outside, and not part of the signing.
It kind of is - it’s signing the hash of the tree object, which is the actual thing that you’d attack with a hash collision
The actual git ‘tree’ object, which is the thing a commit actually points to, referenced by a hash in the commit. That is signed by the GPG signature.
That digest can be the SHA-256; since the infrastructure is there for it, signing should use SHA-256 regardless of what hash is used by the repository for identifying and linking content.
> if it takes an unsigned commit and signs only its SHA-1 hash and then creates a new commit with GPG headers
It does indeed. The bytes passed to GPG when constructing a signed commit look something like: tree eebfed94e75e7760540d1485c740902590a00332
parent 04b871796dc0420f8e7561a895b52484b701d51a
author Alice <alice@example.com> 1465981137 +0000
committer Alice <alice@example.com> 1465981137 +0000
Headline
Message
where the contents being signed are entirely represented by the oids of the tree object and parent commit object. This string is very similar to the content that is fed to the hash function to produce a normal git commit object id.The "bytes passed to GPG" of course get hashed by GPG, using something better than SHA-1.
All bytes that comprise the commit should be hashed by GPG, rather than depending on the content referencing hash in the object tracking system.
This is something that is possible; it is not a logically deductive necessity that we just scan the topmost object and trust the hashes it contains.
It kind of is. Otherwise the whole idea of signing a commit with a backing git history (rather than just a snapshot of a working directory) collapses. The only guarantee you have that the git history is what is claimed by the cryptographic signature is some sort of Merkle tree structure. Either the original one, or you have to construct a whole new parallel one with a better hash, in which case, as I bring up in a cousin comment, why not just use a better hash in your original one?
Parent commits can have their own signatures, and it's true that there is an attack possible there where the same parent hash could point to two different commits, that have valid signatures of some kind (possibly from the same sneaky developer who is a bona fide project member).
Be that as it may, it's perfectly okay for the signature on a banana to validate only the banana, and not the gorilla that is holding it, or the vine the gorilla is swinging on, and the whole jungle.
It would be worth it to have better quality commit signing for SHA-1-based repos.
This is a significant degradation of the implicit guarantees given by a cryptographic signature, to the point that basically all personal use cases I have for signed commits would be invalidated.
Keep in mind that git does not have diffs as first-class objects. Every commit is just a snapshot of some state of the working directory. That means that without attesting to the integrity of the parents of a commit, the only thing a commit X signed by a person A says is "at some point on A's computer, the state of the repo looked like X".
Almost all the relevant questions I would want to ask are not answered by this. E.g. there is a malicious function F that is present in X. Did A write it? Don't know. Did someone else write it? Can't know for certain. Who introduced a certain feature? Don't know. Did A sign off on a new bugfix? Don't know.
All you know is that at some point the codebase looked like X on A's computer. A might not have made any relevant changes at all!
You can only back out a diff and therefore actually attribute a change to someone (either explicitly through `git blame` or informally by looking at git logs) if you have attestation of the parents.
The Merkle tree structure of git repos is interwoven through basically ever useful thing git does. Without cryptographic signatures implicitly carrying a promise of validity for that structure, this would make commit signing useless (depending on just how broken SHA-1 is) for the needs of any org I've ever worked at.
Because now, commit X signed by person A does not even attest that the repo looked like X on A's computer, or not very well, due to weaknesses in the hash.
For some people, an improved signing scheme over the SHA-1-linked repository format would meet their requirements based on their assessment of their security threats.
The scheme you just called a "screw up"?
I don't think it's massively stupid. Unless you want to re-hash the entire Merkle tree structure to sign your commit, you basically have to trust the hashes in the Merkle tree (or have a separate parallel Merkle tree) at some point in what you sign, which means you do have to trust the SHA-1 hashes. Otherwise even with a cryptographic signature you can always spoof at least the git repo history (e.g. even if you try to directly hash the entire contents of the current commit).
Re-hashing the entire Merkle tree structure seems prohibitively expensive to generate (even with a lot of caching) and pretty complicated for e.g. verifying a signature. Or you can do that incrementally, but then you're just generating a whole new parallel Merkle tree structure.
Regardless, at the end of the day, you need to trust the integrity of the Merkle tree structure. And you can either do that by trusting the hashes of the current Merkle tree, or you have to completely recreate a new one with more trustworthy hashes, in which case why not just use better hashes in your original tree?
You just have new header fields (or values for existing fields) which indicate that this is such and such a new style signature, and that's mostly it. Older git won't validate it or produce it. Newer git validates and produces old and new.
Importantly, older git installations will be able to read the commits, etc, just not validate the signatures.
This is not to avoid replacing the compromised hash function as such (taking away the choice from people making new git projects). We have that already.
But existing repos aren't based on that new function. Not everyone wants to switch their existing repos to to start using SHA-256; that means the repos don't work with older git.
The disruption is real.
The only substantial difference is that the author vouched for these particular snapshots of code, in a way where nobody else can intercept the communication and substitute a completely different hash than one the author previously signed.
In fact if you can find a SHA collision, you can peal the signature off of the legitimate payload and slap it on the colliding one.