Whatever happened to SHA-256 support in Git?
lwn.net
lwn.net
https://lwn.net/subscribe/Info
And the Wikipedia page for LWN, if you’re not familiar with it:
I don't think the LWN article can be said to take anything out of context. But I think it's worth empathizing that this is a thread on the Git ML in response to a user who's asking if Git/SHA-256 is something "that users should start changing over to[?]".
I stand by the comments that I think the current state of Git is that we shouldn't be recommending to users that they use SHA-256 repositories without explaining some major caveats, mainly to do with third party software support, particularly the lack of support from the big online "forges".
But I don't think there's any disagreement in the Git development community (and certainly not from me) that Git should be moving towards migrating away from SHA-1.
There is a compute cost for that, but it should be minimal relative to the security benefits?
"And more" because to detect a collision with a background SHA-256 you'll need both objects, whereas SHA1DC detects attempts to spoof SHA1 in a way that leads to collisions. So it won't pass along an object that collides with another one, even though it only has 1/2 objects.
That distinction is something that's generally considered important, e.g. there's been past exploits in Git where you could trick a client into doing something bad by e.g. a crafted .gitmodules file.
The fix has not only been to patch clients, but also to patch "git fsck" to detect and reject such bad contents, so that e.g. the forges can't be used to relay a repository exploit to users running older versions.
A viable hash collision exploit in the wild might likewise want to make use of such an attack scenarios, so having servers capable of detecting collisions without having both sides is preferable to doing so by re-hashing with SHA-256.
This is the kind of detail I would have loved to see quoted directly in the article. Sure enough, it's prominently displayed on the project's README::about section, but your very succinct explanation here made it clear in immediate context.
The idea of counter-cryptanalysis is eye-opening.
> I'm happy to answer any questions here that people might have.
Is there any way to achieve a gradual, staged rollout of SHA256?
What's the impact of converting a repo to SHA256 - will old commit IDs become invalid? Would signed commits' signatures be invalidated?
The design document for that is shipped as part of git.git, and available online. Here's the relevant part: https://git-scm.com/docs/hash-function-transition/#_translat...
Basically the idea is that you'd have a say a SHA-256 local repository, and talk to a SHA-1 upstream server. Each time you'd "pull" or "push" we'd "rehash" the content (which we do anyway, even when using just one hash).
The interop-specific magic (covered in that documentation) is that we'd use a translation table, so you could e.g. "git show" on a SHA-1 object ID, and we'd be able to serve up the locally packed SHA-256 content as a result.
But the hard parts of this still need to be worked out, and problems shaken out. E.g. for hosting providers what you get when you "git clone" is an already-hashed *.pack file that's mostly served up as-is from disk. For simultaneously serving clients of both hash formats you'd essentially need to double your storage space.
There's also been past in-person developer meet-up discussion (the last one being before Covid, the next one in fall this year) about the gritty details of how such a translation table will function exactly.
E.g. if linux.git switches they'd probably want a "flag day" where they'd transition 100% to SHA-256, but many clients would still probably want the SHA-1<->SHA-256 translation table kept around for older commits, to e.g. look up hash references from something like the mailing list archive, or old comments in ticketing systems.
Currently the answer to how that'll work exactly is that we'll see when someone submits completed patches for that sort of functionality, and doubtless issues & edge cases will emerge that we didn't or couldn't expect until the rubber hits the road.
For any internal part of Git the hash type is already known, e.g for the wire protocol dialog that "fetch" runs. Other commentators have pointed out that you can use the hash length to disambiguate the two, but the way it's implemented internally we could tell them apart even if we hypothetically had two hashes of the same length.
But that leaves abbreviated hashes, e.g. if you do "git show deadbeef" is that a SHA-1 or SHA-256 hash? We don't know.
It's forseen that in those cases we'll look up both, and in the case of ambiguity do the same thing as e.g. "git show dead" does now (try it on a non-trivially sized repo). Then just as you can do e.g. "git show dead^{commit}" now to disambiguate, you'd use the ^{sha1} or ^{sha256} peel syntax (still unimplemented).
If we changed the hash format from [0-9a-f]{4,40} to something that didn't match the first part of that regex there's even more downstream systems that would need adjusting. A lot of programs that work with Git's output use some variant of that regex. For those that don't limit themselves to "40" things usually Just Work.
So that's basically the reason, it's also less future-proof, as you'll run out of magical prefix characters faster than you'll run out of hash names. Although I'm hoping to be dead way before that small space would be exhausted :)
I'm curious how closely (if at all) they've been following this effort
I myself do some work on upstream Git on behalf of GitLab, although none of it's been on anything related to the SHA-256 transition.
As to why no big forge has SHA-256 support, I think it's a bit of a chicken & egg problem (and these comments are entirely my own, and not on behalf of anyone).
I think it's safe to say that all of the forges are expecting the transition, e.g. I don't think there's anyone creating CHAR(40) database tables for Git hashes anymore (or if they are, someone is planning to deal with it).
Another is that for a successful transition for anything except entirely new repository networks (which already use SHA-1) you really need the "git" client to play along, see my other comment discussing hash interop plans. Some of that same code then needs to run on the server-side.
That code isn't part of git yet, and it's really needed for any sort of viable migration plan.
I mean, it's not really needed, at some point a lot of people reading this migrated from CVS and/or SVN to Git. But a full export/import with a lot of users is painful. We really want it to suck less for Git, to the point that it should Just Work for most or all users.
And a major one is the human factor. For things to happen in free software development someone needs to submit patches, brian m. carlson has been performing a heroic amount of effort over the year on the transition over the years. As he notes in the linked ML thread he's had life reasons for why he hasn't been able to work on it as actively recently as he did in the past.
There's some details on this in https://git-scm.com/docs/hash-function-transition/#_fetch; part of the expected hash transition document assumes that you could have say a SHA-1 *.pack, but maintain both a SHA-1 and SHA-256 index into its contents.
For blob objects you could serve them up as-is, but for the other types which refer to other objects (commits, trees and tags) you'd either need to store two copies and do something close to a sendfile(), or rewrite them on-the-fly as you stream them, using your SHA-1<->SHA-256 translation table.
So I got a bit ahead of myself there, but none of this interop code exists yet, so the trade-offs the forges will have to make are unknown at this point.
example.com/example v0.0.0-20171218180944-5ea4d0ddac55 h1:jbGlDKdzAZ92NzK65hUP98ri0/r50vVVvmZsFP/nIqo=
Where "h1" is an upgradeable hash (h1 is SHA-256). If there's ever a problem with h1, the hash can be simply upgraded.
Git's documentation describes how to sign a git commit:
$ git commit -a -S -m 'signed commit'
When signing a git commit using the built in gpg function the project is not rehashed with a secure hash function, like SHA-256 or SHA3-256. Instead gpg signs the SHA-1 commit digest directly. It's not signing the result of a secure hash algorithm.
SHA-1 has been considered weak for a long time (about 17 years). Bruce Schneier warned in February 2005 that SHA-1 needed to be replaced. Git development didn't start until April 2005. Before git started development, SHA-1 was identified as needing deprecation.
This then makes the signing code use its own form of hashing that is different from the rest of git's commmit hashing, and seems like a novel way to introduce tooling issues / bugs / etc.
Git stores content, not diffs. So the signature verifies all content stored in that commit. It does t verify anything that came before it, unless those are specifically signed as well.
But the "contents" is just pointers to tree roots with a trusted hash. If the hash is no longer secure, you can't garantee that any such trees are your content, or safe.
E.g. say you have `5baa61e4c9b93f3f0682250b6cf8331b7ee68fd8`. What version is that?
Well. It's exactly as long as a SHA1 hash. It doesn't start with "sha256:" or "md5:" or "h1:" or "rot13:". So it's SHA1. Easy and totally unambiguous.
Versioning can almost always begin with version 2.
Me, reaping: "Each record begins with a single byte indicating the record format version. In version 0, this is followed by a 3 octet BE value indicating the record length."
for those, you have to leave versioning room up-front. even 1 bit is enough, since a `1` can imply "following data describes the version", if a bit wastefully in the long run.
If you had: "each record begins with an 8 character record length, in hexadecimal, giving 32 bits", you have no problems. The new version has a 'V' character in byte 0, which is rejected as invalid by the old implementation.
inconsistencies in how data is presented (optional version number) is a pain to deal with in code
Git's entire foundation relies on SHA1 hashes. Each commit is its own hash, and contains a list of the hashes of all files that are a part of it. Branches have hashes, tags have hashes. Everything has a hash. A repository that uses a different hash algorithm is a completely different repository, even if the contents and commits are otherwise identical. You can't even store your code on someone else's server (well, aside from manually copying the repository data over, though that won't be too useful) unless that server has upgraded their git version.
Well, Fossil's database is much better designed, you reply.
That it is!
The intent behind it is obsolescence and phasing out, resulting in an endless make-work treadmill for the users.
If there is ever a "problem with h1", and you neglect to upgrade your data right there and then, five to ten years, it will be unreadable.
stepping through required versions a common operation
It's a more robust, well-specified, interoperable version of this concept.
Though it's probably overkill if you control both the consumer and producer side (i.e. don't need the interoperability) and are just looking to make hash upgrades smoother, in that case a simple version prefix like Go's approach described above has lower overhead.
A minor correction: when signing a commit, gpg does not sign the SHA-1 digest of that commit. This is impossible since the signature becomes part of the commit header which is one of the inputs to the hash function that produces the oid.
Instead, GPG signs the serialized data (parents,headers,tree,message) which would otherwise be the input to SHA-1. Then the sig is inserted into the buffer at the end of the header and the string is digested to produce an oid.
Source: https://github.com/git/git/blob/39c15e485575089eb77c769f6da0...
Git signatures are more or less only as secure as SHA-1, although the properties it relies on are not yet compromised and several factors mitigate the real-world risk.
A practical preimage attack on SHA-1 would seriously undermine the security of git signatures since the string being signed includes two or more SHA-1 hashes representing a content snapshot and the commit ancestry. Arbitrary preimage attacks would make it possible to modify a repo’s contents or history without invalidating oids or signatures.
In practice only collision attacks have been found, all of which have a detectable signature that git has been modified to detect.
Disclosure: I work at GitHub, but am speaking for myself.
"Fossil started out using 160-bit SHA-1 hashes to identify check-ins, just as in Git. That changed in early 2017 when news of the SHAttered attack broke, demonstrating that SHA-1 collisions were now practical to create. Two weeks later, the creator of Fossil delivered a new release allowing a clean migration to 256-bit SHA-3 with full backwards compatibility to old SHA-1 based repositories. [...] Meanwhile, the Git community took until August 2018 to publish their first plan for solving the same problem by moving to SHA-256, a variant of the older SHA-2 algorithm. As of this writing in February 2020, that plan hasn't been implemented, as far as this author is aware, but there is now a competing SHA-256 based plan which requires complete repository conversion from SHA-1 to SHA-256, breaking all public hashes in the repo."
[0]: https://fossil-scm.org/home/doc/trunk/www/fossil-v-git.wiki#...
Joking aside, expected from a developer whose work is the recommended storage format for Library of Congress.
I’ve added to my todo list a reminder to raise this issue with mine. In fact, I’m going to give them a deadline for when we will start evaluating competitors that do support SHA256.
I suspect that most people on HN do not interact with their MS account team. That relationship is probably managed by your CIO or IT department. They probably have monthly or quarterly “business review” meetings. You should get this issue on the agenda of that meeting.
The article says "none of the Git hosting providers appear to be supporting SHA-256", and while GH is not mentioned by name (and I applaud them for indeed not strengthening this "git == github-the-brand" trap), I can't imagine GH was left out of scope when checking the major hosting providers.
Gitlab also appears to be lacking support [0], and the same with Gitea [1].
so it's a grey area where Git itself supports SHA-256-based repos, but without the major Git hosting services also supporting them, the support in core Git is somewhat useless.
Luckily I've versioned the internal hash so the upgrade path back to sha256 should be as smooth as the downgrade was. I'm still bitter about it though.
Linus shouldn't have used SHA-1 in the first place, it was already being deprecated by the time git got its original release. Then every time a new milestone is reached to break SHA-1 we see the same rationalization about how it's not a big deal and it's not a direct threat to git and blablabla.
It'll keep not mattering until it matters and the longer their wait the more churn it'll create. Let's rip the bandaid that's been hanging there for over 15 years now.
Using SHA-1 to begin with was fine. However, commit hashes should have been prepended with a version byte to make it easier to transition to the next hash algorithm.
This would mean an old Git client could report an error to the user of the nature “please upgrade your software to support cloning from this Git server” instead of failing with an error that’s inseparable from “the Git server is broken” when trying to clone a Git repo using SHA-256.
Once a system expects to handle SHA-1, then you have to deal with old assets that have deprecated signatures, and that's a fight I 1) didn't want to have and 2) was fairly sure I wouldn't be around to win.
Git was still brand new, largely unproven at that point, and I don't understand why he picked SHA-1.
So strictly speaking Linus and subsequent maintainers weren't being amateurish in the beginning. (You didn't say that explicitly, but it would be a fair criticism given what was known about SHA-1 at the time, including known by Linus--he knew and made a choice.) Rather, in the beginning it was naivety in believing that people wouldn't begin to depend on Git's apparent security properties.
Honestly, I think it's fair to say that hashes isn't meant to be a security feature.
But signed tags/commits/etc. probably need a better hash.
The first OpenSSL release that has general SHA-256 support seems to have been 0.9.8, released on July 5th, 2005, the code first appeared in OpenSSL's source tree in May of 2004.
Perhaps Linus has commented on it. I don't know, but I wouldn't be surprised if the actual reason is that Git was thrown together as a weekend project, that he vaguely knew SHA-256 was preferable, but his distro's OpenSSL didn't have it yet.
So the initial version used SHA-1 instead, and the rest is history...
Seems like this company could just use the current SHA-256 support then? Especially if it's the type of company that does all its development in-house and there's no need for SHA-1 interoperability.
> I have very recently seen customers move to older much less functional (or useful) VCS platforms just because of SHA-1.
A company this dysfunctional has problems far beyond their choice of revision control system.
0: I use the word "security" only because the teams themselves are named like that. You can probably infer my opinion from the tone.
So, we sold people software that would run on some fridge-sized Sun machine running Solaris, to ensure that their Solaris machine wasn't about to get infected with the latest Windows virus.
The occasional support calls with technically minded *nix admins were amusing. We knew that what we were selling them was completely useless and made to secret of that fact, they likewise knew that the software they were running was useless to them. The one thing they cared about was that it didn't contribute to the load, and we did our best.
But some PHB somewhere in their organizations had decreed that all computers everywhere must have an anti-virus scanner, and if you're sufficiently motivated to buy something eventually someone will sell it to you, even while telling you that you don't need it :)
But blanket-banning an obsolete and insecure hash algorithm isn't a bad thing, it's entirely reasonable. In this case, as the article makes clear, it's git that's at fault.
Reminds me of the time a security audit (which literally just involved running some scanning tool and dumping the results on us) complained that some code I had written was using MD5 - but in a use case in which we weren’t relying on it for any security purposes. I ended up replacing MD5 with CRC-32 - which is even weaker than MD5, but made the security scanning tool mark the issue as remediated. It was easier than trying to argue that it was a false positive.
The big problem with using sha1/md5 in non-secure contexts is:
*Someone later might think its secure and rely on that when extending the system.
*it can make it difficult for security people to audit code later as you have to figure out if each usage is security critical
Using a non crypto hash makes both those concerns go away since everyone knows crc32 is insecure. The alternative of using sha256 also works (performance wise it is close enough, so why not just use the secure one and be done with it.)
- Change the binary file format in repos to support arbitrary hash algorithms, in a way which unambigously makes old software fail.
- Increment the Git major version number to 3.0
- Make the new version support both the old version repos and the new ones. Make it a per-repo config item that allows/disallows old/new hash formats. In theory, there's nothing wrong with having objects hashed with mixed algorithms as long as the software knows how to deal with that.
- The old format will probably have to be supported forever because of Linux.
Most user-facing utilities don't care what the hash algo actually is, they just use the hash as an opaque string.
And why is this an issue? Release the new version that can read new repo formats, but doesn't write them yet. Wait a year. Release new version that can write new repo formats and encourage users to upgrade.
Anyone who hasn't upgraded in the past year probably doesn't care about security and should be left behind. Besides, once they google the error message they'll figure it out soon enough. It's not like git is known for its great UX anyway.
That's an interesting idea, actually. I'm not sure they plan to support that, though? That would make things a lot easier on existing repositories; without support for mixed hashes, repos would have to have their history entirely rewritten, which would invalidate things like signed commits/tags.
there is one hash version, plus a translation table for the other format. no history rewrite.
new repos will use the new hash. old repos will eventually fully convert to the new hash, then all old hash links after the transition period will become obsolete.
Okay, but that's a pretty big reason! A git repo that can't be pushed to github/lab is... not always useless, but certainly extremely impaired.
git init --bare public_html/mything.git
cd public_html/mything.git/hooks/
mv post-update.sample post-update # runs git update-server-info on push
(This assumes that your public_html directory exists and is mapped into webspace, as with the usual configuration of Apache, NCSA httpd, and CERN httpd. If you don't have an account on such a thing you can get such PHP shared hosting accounts with shell access anywhere in the world for a dollar or two a month.)And then on your dev machine, it's precisely the same as for pushing to Gitlab or whatever, except that you use your own username instead of git@:
git remote add someremotename user@myserver:public_html/mything.git
git push -u someremotename master # assuming you want it to be your upstream
Then anyone can clone from your repo with a command like this: git clone https://myserver/~user/mything.git
They can also add the URL as a remote for pulls.If you want them to be able to push, you'll need to give them an account on the same server and either set umasks and group ownerships and permissions appropriately or set a POSIX ACL. Alternatively they can do the same thing on their server and you can pull from it. There are reportedly permission bugs in recent versions of Git (the last five years) that prevent this from being safe with people you don't trust (https://www.spinics.net/lists/git/msg298544.html).
Of course source control is only part of the overall development project workflow, so for many purposes adding SHA-256 support to Gogs or Gitlab or Gitea or sr.ht is probably pretty important: you want a Wiki and CI integration and bug tracking and merge requests. But the git repo still works fine with a bog-standard ssh and HTTP server, though slightly less efficiently. It's easier than setting up a new repo on GitLab etc.
Running a git repack -an && git update-server-info in the repo on the server can help a lot with the efficiency, and for having a browseable tree on the server as well as a clonable repo I put this script at http://canonical.org/~kragen/sw/dev3.git/hooks/post-update:
#!/bin/sh
set -e
echo -n 'updating... '
git update-server-info
echo 'done. going to dev3'
cd /home/kragen/public_html/sw/dev3
echo -n 'pulling... '
env -u GIT_DIR git pull
echo -n 'updating... '
env -u GIT_DIR git update-server-info
echo 'done.'
That's very far from being GitLab (contrast http://canonical.org/~kragen/sw/dev3 with any GitHub tree view), and it's potentially dangerously powerful: if you're doing this in a repo where you pull from other people, and the server is configured to run PHP files or server-side includes in your webspace (mine isn't!) or CGI scripts (mine is!), then just dropping a file in the repo can run programs on the server with your account privileges. This is great if that's what you want, and it's a hell of a lot better than updating your PHP site over FTP, but that code has full authority to, for example, rewrite your Git history.In theory you can do other things from your post-update hook as well, like rebuild a Jekyll site, send a message on IRC or some other message queueing system, or fire off a CI build in a Docker container. (Some of these would run afoul of guardrails common in cheap PHP shared hosting providers and you'd have to upgrade to a US$5/month VPS.)
Hydraulic's first product is a packaging tool for desktop apps and we're mostly a JetBrains shop. The gist is:
- A large dedicated machine in a cheap colo provider (Hetzner), with
- Gitolite with some custom configs
- YouTrack for tickets
- TeamCity master, some agents and a Windows VM for testing.
- Dedicated Mac hardware in the office for Mac CI testing, also running TeamCity agents.
The workflow is a homegrown one that we call "git oriented review". I'll briefly describe it and then discuss why we use it:
1. (Almost) Every git repository is owned by someone specific. There are no shared repositories. All code flows upwards to my repository via merges, kernel style, and that's the one that's used for releases.
2. Gitolite is configured to allow users to push into each others repositories but only under a special branch namespace (rr/$user/whatever). Other branches are protected and yours alone. You also have a personal set of build configs in TeamCity, so if you want to create a customized CI setup you can, and so branches you push to your personal repo don't interfere with the greenness of anyone else's builds.
3. To submit code for review, you push a "review request" branch in the rr/$user namespace of the reviewer's repository. The commit message(s) have a command in them that's interpreted by YouTrack once the build goes green to update the ticket, which in turn then notifies the reviewer that new code exists and is green. The notification can be via email, or IDE notifications, or Slack etc. Lots of options, up to you.
4. The reviewer then makes a code review by adding commits to the rr branch. For small changes, the reviewer just makes the change directly. For larger changes they add a //FIXME comment to the code. FIXMEs may not be merged into non-rr branches, so they are a request for the original author to remove them by e.g. fixing the issue, or adding a comment to explain why in reality it's not meant to be fixed. Thus code and commits are used as a type of discussion forum. You can of course also just jump into a CodeWithMe session to do a bit of pair programming on it for more complex discussions (everyone is currently remote at Hydraulic so everything is done via tools).
5. When satisfied the reviewer merges the rr branch to their own master or dev branches and deletes it. The merge commit contains another command to mark the ticket as fixed, again, it's only applied if the build goes green. At this point they "own" the result because it's in their personal repository. Finger pointing isn't allowed.
I designed this workflow due to lack of satisfaction with the GitHub PR based workflows used at my previous firm, which was a fairly typical centralized one involving a single shared repository, a protected master branch with people pushing ad-hoc branches into it and then opening PRs for reviews. The problems I wanted to fix were:
• Reviewers would only respond with comments because that's what the GitHub workflow promotes. Often many comments were small and it would have been both much faster and also more collaborative for the reviewer to make the change directly, but lack of clear ownership over branches made conflicts likely, and people were reluctant to do this.
• Sometimes what the reviewer wanted wasn't obvious. Again if they could have made the change directly it would have been better.
• Reviewers would often get tired after enough comments, or more than a few rounds of review, especially if the dev wasn't actually applying all the requested fixes. So they'd end up waving code through that wasn't really fully fixed.
• PRs notified reviewers before CI had tested the code. This would often lead to races in which the review was completed before CI pointed out that the code was broken, wasting a review cycle (unfortunately CI was quite slow at the old company due to it being a database engine with lots of IO heavy regression tests). This problem has led GitHub to create "Draft PRs" which don't make conceptual sense.
• We ended up with many branches where it wasn't entirely clear if they were abandoned or not. People became reluctant to delete branches in case they were being used to back up important but unfinished work, and again, there was no clear ownership of who was supposed to do this (I didn't get to pick the management approach and would have fixed this stuff if sufficiently empowered).
• Relatedly we lost clear ownership of the codebase. At first ownership was mine because I approved all reviews, but as that stopped scaling the firm transitioned to a system in which coders could pick their own reviewers and ownership effectively became collectivized. The codebase wasn't really laid out with CODEOWNERS files in mind, so devs just had to get a review from someone and then they could commit. This led to a lot of externalization of costs and juniors reviewing each other's code, often letting serious problems through without realizing.
Git oriented review solves these problems. Code ownership is always concrete and well defined by repository, which avoids needing to mangle the codebase itself to try and reflect shifting reporting lines in the directory hierarchy. Reviewers become collaborators on a branch, relying on git's merging features to avoid conflicts. Discussions use commit messages, or whatever is more appropriate when that's insufficient, instead of being tied to a relatively poor and low-featured ad-hoc "discussion forum" like a GitHub PR is. People can organize their own branch namespaces. Reviewers are informed there's work to do only when a build goes green, and they can control how those notifications work. CI and ticketing are closely integrated so tickets have work logs. Finally, the history of the code review is backed up in a portable and vendor-neutral git repository.
Downsides? Not many found so far. It's unfamiliar to new devs and requires a bit of training especially if their git skills aren't fluent. Gitolite is powerful but has a few awkward limitations and was a bit of a bear to configure. The JetBrains tools are great and cheap/free for small operations like ours, but do require payment later. To view logs without the review history polluting things you have to know about the `git log --first-parent` flag which many people don't realize exists, and which doesn't have any equivalent in the IntelliJ git view. Overall these things are pretty easy to fix. IntelliJ git is open source so we could even add that feature ourselves if necessary, but so far it wasn't.
There's also clarity over ownership that way - the `rr` branch is owned by the owner of the repository and they can push to it, to make it into whatever form they want.
Devs may also explicitly send a message indicating there's work to be reviewed, but they don't have to.
It would have made sense to say that in 01990 when the hardware cost US$12000 and the software required constant hand-feeding. But now, virtually every home internet connection has a server built into the cable modem, you can rent a VPS for US$5 a month, and you can bring up a running nginx configuration with a single docker command.
Running a server isn't any more difficult than running an Ubuntu laptop — in fact, it's mostly the same tasks, except that you can version-control the server setup in Git — and considerably more educational.
So I would say that most developers don't run their own server, and that's a criminal failure of education that imperils the future of civilization.
It seems like that could ease much of the migration problems if it's not a problem?
Generally though length ext attacks have a solution - HMAC, which is much more secure than truncate.
The more you truncate, the more vulnerable you are to birthday attacks (practically speaking you would have to truncate quite a lot)
Also i think the length of the input matters when comparing sha256 vs sha512.
There's a point where truncating starts to make it weaker, but when you first start chopping off bytes the benefits outweigh the drawbacks.
If for some reason length was an issue, a base64 encoded 256 bit string, like a SHA-256 digest, is 43 characters. That too can be truncated to 40 characters, which has 238 bits of security. SHA-256 is not only a better hashing algorithm than SHA-1 but it could also result in higher effective security even when truncated.
160 bit output, without a cryptographic weakness, is good for about 30 trillion commits per second continuously for 1000 years.
For SHA the cryptographic strength isn't primarily from the length of the hash, but from the internal number of rounds is (e.g. 160-bit SHA-1 with fewer rounds has been badly broken way earlier, and 160-bit SHA-1 with more rounds would be safer).
Cryptographic hashes are designed to be safe to truncate and still have all the safety the truncated length can provide. It's basically a requirement for them being cryptographically strong. Even in the SHA-2 family, the SHA-224 and SHA-384 are just truncated versions of larger hashes.
SHA-1 is quite broken at this point. SHA-256 is not. There aren't any practical non-generic attacks on full sha-256 and thus there wouldn't be any on the truncated version. The Wikipedia article goes into the different attacks on the two algorithms.
That said, if your concern is length extension attacks - strongly reccomend using sha-512/256 instead of trying to do your own custom thing.
If that was all that was left, we could at least be using sha256 for new repositories.
It seems to me the big missing piece is support in libgit2, which is at least showing signs of progress:
If everyone started using sha256 then all these problems would be addressed practically overnight.
Sha256 can only be computed in a single sequential stream (thread) by definition.
For large files this is increasingly becoming a performance limitation.
A Merkle tree based on SHA512 would have significant benefits.
SHA512 is faster than SHA256 on modern CPUs because processes 64 bits per internal register instead of 32 bits.
A tree-structured hash can be parallelised across all cores.
For repositories with files over 100MB in them on an SSD this would make a noticeable difference…
SHA256 is actually a lot faster on modern CPUs due to https://en.wikipedia.org/wiki/Intel_SHA_extensions (and similar on Arm), which are implemented for SHA-256 but not for SHA-512, e.g. openssl speed sha256 sha512 on M1:
type 16 bytes 64 bytes 256 bytes 1024 bytes 8192 bytes
sha256 89474.97k 283341.15k 901724.41k 1730980.24k 2339109.86k
sha512 66160.19k 262139.03k 365675.96k 487572.26k 545142.91kBut again, due precisely to their size, large files take a disproportionate amount of time to process.
Don’t confuse the typical use-case with the fundamental concept: versioning.
Git could be a general purpose versioning system with many more use-cases, but limitations like this hold it back unnecessarily…
Granted, signed tags do depend on this collision resistance, but I don’t use that feature. Signing entire releases from a trusted repo seems like a better approach.
Sure, if you use git with a very closed development model, this doesn't necessarily affect you much. But it's (potentially) a big problem for collaborative open-source projects, because it requires trust in every single contributor. And the trust requirement can't necessarily be mitigated using ordinary means like code reviews.
Besides, the known collision attack generates files with blocks of binary garbage, which makes it difficult to trick someone into accepting. It won't look like source code, and if someone accepts binary blobs of executable code, you don't need collisions to pwn them.
IDK, I could see this happening in multiple ways.
1. Images / media artifacts stored for display purposes
2. Cached files - 'zero install' config for yarn comes to mind, where every dependency has its file cached in git.
Plus binary files aren't displayed in git diffs so it seems somewhat easy to sneak in.
Otherwise, yeah, agree. Most people don't rely on Git's security model, they rely on Github's.
They are (albeit not as prominently as they should). And you can add your own diff engines to show full diffs for different binary formats.
Upstream Git client says "Binary files a/filename and b/filename differ" whenever it detects changes a binary file. This is mentioned in output of 'git diff', 'git status', 'git show' and other commands.
As noted in the article, an SHA-1 collision attack does not appear practical now, but that is a situation that can change.
Not to say that this attack is in any way practical, yet. Just that some providers don't require active involvement to try and attempt it.
I don't know if this would survive additional commits on top as I'm not familiar enough with git's internals.
"Given the threat that the SHA-1 hash poses, one might think that there would be a stronger incentive for somebody to support this work. But, as Bjarmason continued, that incentive is not actually all that strong. The project adopted the SHA-1DC variant of SHA-1 for the 2.13 release in 2017, which makes the project more robust against the known SHA-1 collision attacks, so there does not appear to be any sort of imminent threat of this type of attack against Git. Even if creating a collision were feasible for an attacker, Bjarmason pointed out, that is only the first step in the development of a successful attack. Finding a collision of any type is hard; finding one that is still working code, that has the functionality the attacker is after, and that looks reasonable to both humans and compilers is quite a bit harder — if it is possible at all."
But you could also hide it as a fake lookup table or inline XPM or something like that.
This seems true yet there are no demos or documented attacks using this method.
I think practically speaking it’s kind of a pain to do.
Sha1 has a collision attack. We are far away from a preimage attack
To elaborate a bit: One thing that makes a viable attack against Git especially hard is that aside from the hash it's using has a behavior of never replacing an already hashed object[1].
So let's say I have a tool that can take a given file & SHA-1 pair and produce a collision, the next step is quite hard. I could in this scenario produce a file with an exploit whose hash matches that of Linus's kernel/pid.c or whatever.
But how do I get that object to propagate among forks of linux.git to distribute my exploited code?
If I e.g. push it to a fork of linux.git on a hosting provider that Linus uses the the remote "git-index-pack" process will hash my colliding object, but before it stores it check whether such an object ID exists in its object store, if it does it'll drop it on the floor. You don't need to store data you've already got in a content-addressable filesystem.
Which is not to say that a hash collision is a non-issue, and Git should certainly be migrating from SHA-1. There's no disagreement about that in the Git development community.
But it matters for how much you should panic how the software you're using could be exploited in the case of a hash collision.
Also, the scenario above presupposes a preimage attack, which is a much worse attack on a hash function than a collision attack. Currently no viable preimage attack on SHA-1 exists, only a collision attack.
Which means that before any of the above I'd have to have produced a viable version of say kernel/pid.c that Linus was willing to merge, knowing that my evil twin of that version is something I intended to exploit people with.
Then I'd need to patiently wait for that version to make it into a release, knowing that even a one-byte change to the file would foil my plans...
1. On the topic of running with scissors: I wrote a patch to disable that collision check for an ex employer, it helped in that I/O-bound setup, and we were confident in the lessened security being a non-issue for us in that particular setup. The patch never made it into git's mainline. The patch won't apply anymore, but the embedded docs elaborate on the topic: https://lore.kernel.org/git/20181113201910.11518-1-avarab@gm...;
But, even supposing a libgit2 that didn't use SHA1DC I think most users would be protected in practice if the "git" they use used SHA1DC. Hosting providers, local editors etc. use libgit2 for a lot of things, but I think in most cases (certainly in the case of the popular hosting providers) it's some version of "/usr/bin/git" that's handling your push, and actually propagating your objects.
For stopping a colliding hash it's enough that any part of the chain of propagation is able to stop it.
Whereas the GPG signature facility in git is creating a "tag" object with a signature of the preceding part of the object envelope, and that envelope refers to another object type. E.g. in the case of signed tags usually a commit object.
So if you are able to replace the blob the GPG signature will still be valid, since it relies on signing the SHA-1 DAG.
There is e.g. the third-party git-evtag[1] which gets around this problem by recursively unpacking the signed content, which does give you the advantages of the stronger hash. But this isn't how signed content in Git itself currently works.
I think there was some recent-ish discussion of having something like that in git itself, and perhaps I've missed something, but I'm pretty sure I didn't miss the tag signing code doing a full object walk, which is what you'd need for signing a SHA-1 Git repository in a way that didn't piggy-back on SHA-1's security.
But in that case you'd still have the same collision, as surely what you're producing a collision for is the tree or blob object(s) you're asking to be merged in, as opposed to the hash for the commit object (which will change as you amend it). No?
I think the actual issue here is environment accreditation not allowing the use of sha-1 at all, but that is still rare. It'll become a much larger issue if a future FIPS standard ever disallows sha-1, because that will impact a ton of environments. It means git won't even work on your servers any more.
Example: `git commit -a -S -m 'signed commit'` signs the SHA-1 hash directly.
Even if the SHA-1 digest is rehashed with a secure hashing algorithm, SHA-256, it would hide the fact that the reference is to an insecure hashing algorithm. The project itself needs to be rehashed with a secure hashing algorithm for signing to be secure.
How does that work? My understanding was that a git gpg signature only signs the project at that commit state.
It says nothing about past (or future) commits outside of a digest reference to past commits, which if that digest wasn't upgraded, would be considered insecure.
Said another way: Git does not rehash past commits, or the present commit, when gpg signing. A commit itself only includes the SHA-1 digest of the previous commit.
The old signatures will still be SHA-1. But if you try to replace any part of a commit, the SHA-256 won't match. So the combination of "the commit is an ancestor of multiple securely signed commits in this repo" and "the SHA1 on the signature matches" is enough to know you have the right data in most use cases.
This is very similar to the following: Instead of rehashing, i.e. replacing old hashes with new hashes, add the new hashes alongside the old ones, and sign the new hashes, together with the time mark, by a trusted authority. The old hashes and signatures then remain valid indefinitely as long as the new hashes and signatures are verified successfully.
Here’s the paper: https://marc-stevens.nl/research/papers/C13-S.pdf
If someone malicious can make commits to your repo, you have a much worse problem than a possible hash collision.
Most languages these days have a built-in hashtable type, and they use some kind of non-security-critical hash function, albeit usually with a random seed, that's a lot weaker than SHA-1 - often you only need something like 32-64 bits of output in the first place as no-one allocates more than 2^64 buckets as far as I know. From a security perspective that's even worse than MD5!
I'd argue (and so does the article, in part) that the use of SHA-1 in git is closer to the use of a hash in a hashtable than to the use of a cryptographic checksum - using SHA-1 for subresource integrity on the web is an obvious security risk (luckily, it's not allowed there) but I can't see an obvious attack on git's use that doesn't also assume you can do much worse things.
This is exactly how GitHub works though: all repositories in a fork network share a single underlying repository, only the ref namespaces are separated.
- what happens when the source repository rewrites all the commits (to remove one); should the forks also be re-written? - what happens when the source repository disappears?
_All_ of the problems go away when you actually do a "git clone" of the source repository and have your own objects. More storage, yes, but way fewer corner cases to think through.
The news has been proclaimed loudly and often: the SHA-1 hash algorithm is terminally broken and should not be used in any situation where security matters.
There are organizations where SHA-1 is blanket banned across the board - regardless of its use
I think it's about time people understood the different types of attacks and what they mean for various uses of hash functions instead of blindly cargo-culting as the security (aka paranoia) industry seems to want to do. Then again, the people in those same organizations who promote this sort of anti-thinking would probably suddenly not care if you just renamed all occurrences of "SHA1" in everything they see to something else, because they are such incompetent idiots anyway --- true stories from experience...
Either way, anything relying on hashes for data integrity should at least be flexible to the option of multiple hash algos. But with git, it's going to be hard enough as is to change to SHA-256, and I don't know how parametric it'll be.
Also, BLAKE3 is so significantly faster than it's probably worth waiting for (waiting for it to get vetted cryptanalysis-wise).
(I have seen proofs of concept [1], but never actually heard of an exploit in the wild using it; for example, on: digital certificate signatures, email PGP/GPG signatures, software vendor signatures, software updates, ISO checksums, backup systems, deduplication systems, Git, etc.)
https://twitter.com/rauchg/status/834770508633694208 > a SHA-1 "Pinata" [...] claimed
https://news.ycombinator.com/item?id=13723892 > Make your own colliding PDFs
https://news.ycombinator.com/item?id=13917990 > Collision Detection
The most in the wild one i have ever heard of was when webkit accidentally broke their svn repo by checking in a collision.
However you can look at the history of md5 which had a similar flaw which was exploited by the flame malware.
If a practical pre-image attack on SHA-256 comes around we have bigger problems than git.
Yes, a practical preimage weakness in SHA-256 is a nightmare scenario with huge implications to the rest of internet security beyond just get. It's why I sometimes can't sleep at night knowing how much energy bitcoin spends daily on a continuous massively distributed partial preimage attack on SHA-256.
I would not be concerned about this. The way the asics operate is they discard the results. Also, the hashes are random strings which don't compress very well, so storing trillions upon trillions of them (for later analysis) is not practical.
Maybe you find "cold comfort" that because we can watch it in real time if someone discovers a weakness we will also watch its repercussions and the subsequent horrifying fall in real time, too, but I certainly don't.
A new hash algorithm for Git https://lwn.net/Articles/811068/
Updating the Git protocol for SHA-256 https://lwn.net/Articles/823352/
Just have the malicious code in a file and append a multiline comment block, then have your collision generator insert random junk into that comment block
> Even if creating a collision were feasible for an attacker, Bjarmason pointed out, that is only the first step in the development of a successful attack. Finding a collision of any type is hard; finding one that is still working code, that has the functionality the attacker is after, and that looks reasonable to both humans and compilers is quite a bit harder — if it is possible at all.
* Add support for new hash
* Migrate all data to new hash, dropping support for clients that don't support it.
It appears we're in the waiting phase between the two bullet points. I imagine that could be many years, because many people don't update their git clients often.
if backwards compatibility is so hard, why not break it? replace git with git2 in all places (command line, URLs, protocol, etc) and make a git1to2 utility, and you are golden
Sounds like there's money in this.
I give -3 flying ducks about this, and don't want the Git storage format to be diddled with in any way. Git in 2122 should read and write a git repo made in 2010.
Git is not a public crypto system.
If you think a commit is important and needs to be signed, you need to sign the files and add the signature to the commit.