HNHacker News
TopNewBestAskShowJobs

schacon

2,746 karma · joined March 12, 2008

works at GitButler. loves kittens, ruby and git.
submissionscomments
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
akschually! I argue in a footnote at the end that sha1dc is an unnecessarily expensive shim protecting codebases in a similarly unnecessary way and should be removed so that our pushes and fetches can be much faster.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
Well, this could change at any time. Part of the point was to argue as if SHA1 was completely broken and I still feel that its the correct argument.

But from your chosen-prefix argument, having some object thats been around a while is a problem, because any viable attack needs to be fairly new, since git wont replace objects it already thinks it has. Maybe fresh-shallow-clone scenarios like GitHub actions, but that still always has a non-sha based authentication protection (in other words, actions never run on untrusted code and that trust is never based on signed artifacts but on source provenance)

schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
Impractical for everyone. Because if you want to do this, there are better attack vectors, even if second preimage was easy, even though it is impossible.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
I did generate that screenshot as I am not in the beta for this - I don't think many outside of GitHub are. However, this is exactly how GitLab does it and I would be _really_ surprised if this is not almost exactly how it's implemented. There is almost no other reasonable way to do it.

You can't select the repository format based on the first thing pushed to it, because both sides have to be initiated before a transfer can happen.

So you either start the project on the server and then clone an almost empty repository to start working (which I think is rare) - in which case you need to choose the format like this.

Or you initialize it locally and start your project and then push it to GitHub at some point, in which case you need to initialize a server side version that matches the format to push it to. There is no "initialize a new thing on push" and there never has been.

schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
So, I haven't worked at GitHub in some time, but we never had a flat namespace for objects. There are a lot of SHAs in the DB but objects are always namespaced by repository. Forks shared an object database for efficiency, and have technically added reachable objects to a shared database via fork, but it's never been a real problem afaik.

But the very wrong assumption here is: "if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes"

All parts of this are incorrect.

You can push a colliding commit to a fork, but Git will see that it's already there and ignore it - first write does win. Also, "backend doesn't check uniqueness" is also wrong. The server will check for collisions and if this particular case happens, the server will see this and warn you _AND_ not write the object.

schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
I specifically argue that it doesn't matter if it's $1 and base my argument and solution around that. So it's irrelevant where it is in 2035.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
(2 counter) is impractical because all nodes of git will not replace objects if it thinks it already has it. So any attack has to assume this is the first time the node fetched, which is difficult before trust is established, which is difficult. This is part of the argument Linus originally outlined for this vector, which is that it only works for _very recent_ objects.

(1/3 counter) is not what I argued. I argued from the worst-case position that collision and preimages were theoretically cheap and fast. Even in that case, I feel my arguments hold.

The main issue here is that you assume you can replace an existing object with a replaced one, which you cannot. Not only that, but in all known cases, the sha1dc variant of SHA1 that Git uses will even _tell_ you that someone tried to do this, which singles out the source quickly.

schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
It doesn't reject, but it will not replace. Same for a fetch/pull. That is another issue with this attack vector (that Linus also mentions) - it has to be the _first_ time that a node has seen this object. It makes the attack even more difficult than it already is (in like 4 different major ways)
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
It was theoretical - the point was that maybe some paper is published or some new tech or issue comes up. Now we have to do this again. If we separate the concerns, then we don't have to deal with both as though they're one problem. We can deal with one thing for content addressing and another for trust and security.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
Also, functionally, this is incredibly easy to add to Git.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
Actually, this entire blog post came out of a short chat at Git Merge a few weeks ago with Jeff King. I argued more or less this and he didn't _entirely_ disagree, though he has good counterarguments on the list over the last few years, so I don't really know how he thinks about it ultimately.

I would write this to the mailing list, but I thought a conversation that includes people outside that list is more interesting to me. Ultimately I'm not sure if I'm dumb about this or the whistle blower that's willing to actually say "maybe this isn't the right call"

schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
Being a founder of GitHub doesn't make my opinion more interesting. I hope the argument stands no matter who wrote it. :)
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
I do believe this. But that doesn't mean I can't be convinced otherwise by a good argument. The point of this post is to see if anyone has a great counterargument to change my mind.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
They do give the user the choice, but the default is changing. My point is not necessarily to rip out the SHA-256 option, but simply to not make it the default. Because then people will create repos in that format that do not understand the ramifications, where the opposite should be true.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
Technically, git's design (thanks to very smart people trying to solve this problem like brian and others) is _very_ easy to modify to different hashing algorithms now. A lot of amazing work has gone into this in recent years.

However, it's not a git problem. It's an ecosystem problem. It's that every git repo has to choose one and they're entirely incompatible with each other. That is the cost and the difficulty.

schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
Actually, I think sha-256 is possibly faster than the sha1dc variant that Git currently uses.

I just sent a patch series to the list that enables sha1dc to be accelerated on modern CPU architectures to close to normal SHA1 speeds, but since it was ported from a Rust project by an agent, it will never be applied.

https://lore.kernel.org/git/20260929112544.86511-1-scott@git...

schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
That is essentially only a second preimage problem, which is basically impossible.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
There are plans to keep sha1s around in a database, but as far as I know, no way to transmit those, so they seem specific to individual forges. They can be recomputed, sure, but again, any signatures break and it's possible that in the case of an actual replacement, the recomputation is now wrong and not easily comparable. So what is the point?
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
My point is not that it's the end of the world (or the end of Git), but that it will be painful and unclear and confusing to lots of people. That would be fine if it made a huge difference in trust or protection, but it's the wrong way to do that.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
Also, interestingly, Git today does _not_ use a straight SHA1 because of these attacks. It uses `sha1dc`, a slower collision detecting variant that specifically checks for this vector of attacks. So currently, Git's SHA-1 variant is not susceptible to the SHAttered/Shambles attacks.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
I do believe they're smarter than me, but sometimes very smart groups talk themselves into ultimately impractical solutions because they're all smart. Sometimes you need a dumb guy to come in and say "are you sure this is right?"
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
I mean, there are two things here. One is how difficult it is to have a different hashing mechanism. Brian and other heroes in the Git core group have done amazing work to make this _technically_ possible on a repo level. To test some of my theories, I trivially implemented MD5 and an insanely dumb and easily breakable hash backend. It's not _hard_ to change the mechanism now. It's about the community.

Fossil isn't difficult to change not because it's technically harder for Git but because Git has a community and ecosystem that Fossil does not. The cost is not in the individual project for Git, the cost is because there is _so much_ in Git and this bifurcates everything.

schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.

2) I specifically argue that even if both attacks were practical and cheap, it's still not the problem we should be focusing on.

3) Have you read this email (that I linked to)? It is almost the same general message (20 years ago) that this blog post is. It literally goes though a theoretical object replacement attack and how dumb this scenario is and so SHA-1 is fine.

https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.1890...

schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
I do mention this in like the first paragraph. I don't feel great about it, but I've listened to these issues for years now during contributor summits and Git Merge talks and while it's always seemed problematic, I thought they would come up with a good solution. This last Git Merge confirmed that it's close to the switch and not in any way solved or improved. I don't want to just go with it for groupthink reasons. I never thought it was a good idea and I have said that, but we have a last chance to rethink this, so I'm curious if I'm alone or in the silent majority.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
I do actually literally write in this that if it was MD5 it also would not be a problem.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
No, my argument is that the change should not happen at all and nobody wants it and it gains the community very, very little but the default change is forcing it on everyone and most will be _entirely_ unaware - now having to solve problems that are difficult to understand. Defaults also matter when they are the wrong defaults.
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
Emily's talk does a pretty good job of summarizing the issues with intermixing the hashes: https://youtu.be/eJJp0RE7cd4
schacon··on Git 3.0's upcoming SHA-256 default will be a costly mistake
You can certainly do this, as I said, this is Google's backup plan. But defaults matter. People will start running this and getting repos that are uselessly incompatible with other repos, tools, libraries and server instances. Having it as an option is one thing. Making it a default will cause a lot of pain for people who don't want to care about this.
schacon··on Looking forward to Git 2.56 – and 3.0
Yes, Patrick Steinhart (GitLab) has been working not only on reftables and pluggable backends for the references data, but also pluggable backends for object storage, so that you can use any database backend format (sqlite, s3, special large file storage options, etc) to store objects if you want (in addition to loose objects and packfiles).

This is work that Patrick and GitLab have been doing for years now and it's very impressive and nearly complete.

schacon··on Looking forward to Git 2.56 – and 3.0
It's much, much worse than that.

Yes, you do need to do that. However, there is also much more work after that.

Git will not intermingle SHA-256 and SHA-1 enabled repositories, even in things like submodules, so anything used in that manner will need to keep both versions into the indefinite future. If you rely on a submodule that has not yet converted, you will have to convert it yourself and try to keep it up to date, or the forge will have to automatically keep a bidirectional mirror (if you have submodules in various forges, you'll have to wait for all of them to do it), etc.

This means that every SHA referenced anywhere on the internet, in commit messages, in issues, in code comments is now invalid and needs a mapping to find the rewritten one for forever.

It also means that every commit signature ever made is now invalid and will probably have to be stripped from the rewritten new 256 history because it's impossible to resign everything.

Companies like Google and GitHub are working on keeping two versions of each repository so that there can be long stages of ecosystem migrations, but no matter what, it's going to be a huge pain for millions of developers for years to come.

Page 1 of 10Next →