SHA1 collisions make Git vulnerable to attacks
metzdowd.com
metzdowd.com
"After sitting through an endless flood of headless-chicken messages on multiple media about SHA-1 being fatally broken, I thought I'd do a quick writeup about what this actually means. In short:
Reports of SHA-1's demise are considerably exaggerated.
What CWI/Google have done is confirmed what we've known for a long time, that SHA-1 is shaky. Using a nation-state's worth of resources and a year of time (https://security.googleblog.com/2017/02/announcing-first-sha...), they've shown that, with a very carefully-crafted document, you can create a collision. Their presentation of the results is detailed and accurate, it's the panicked misinterpretation of those results that are the problem."
Continues here: http://www.metzdowd.com/pipermail/cryptography/2017-February...
[edit: typo]
110,000 USD is in the ballpark of state level players when we are talking about forging documents to avoid any sort of tampering detection. It has practically zero use of small time hackers or script kiddies. Why would anybody invest 110K into a collision? What is the practical use of it?
https://www.bishopfox.com/resources/tools/other-free-tools/m...
https://arstechnica.com/security/2016/08/cisco-confirms-nsa-...
https://www.theregister.co.uk/2010/12/15/openbsd_backdoor_cl...
If there's ever a practical use for it (i.e. money to be made) 110$k is totally accessible to the private sector. It's definitely not "a nation-state's worth of resources" which is the quote I was replying to.
Fortunately there doesn't appear to be a whole lot of practical use for these collisions for the time being.
Why would anybody invest 110K into a collision?
The thing people fear is (1) A collision that lets you have good code pass review, then have evil code released to users; (2) That happening to Linux/Android/Firefox/Chrome; (3) The cost of creating a remote code execution exploit being lower than the market value of that exploit on the black market.I don't know how /realistic/ this fear is. Certainly, if everyone PGP signs all their commits, it's a much-reduced risk - but how many projects mandate that?
or some less scrutinised but widely deployed package
Suppose you are on the verge of completing a major sale to some large, nervous purchaser -- perhaps a major world military. This is a decent-sized but not huge sale: $2 billion, with profits of around $200 million. The other major competitor for this contract is built around Linux and your offering relies on a custom operating system.
Your head of sales thinks that the the purchasing agent seems particularly concerned about security issues with the operating system -- keeps asking questions like "So, can you document that your system is less vulnerable than some 'open source' system?". The head of sales makes a rough guess that a news story about vulnerabilities in Linux might sway the chance of winning the contract by around 5%.
So: that's $10 million in value to your company that might created by generating publicity about the vulnerability of Git so long as that publicity is generated at the right moment in time. What's the chance that 1% of that can be "found" to make it happen?
The thing is: $110,000 is actually a very SMALL amount of money, relative to the amounts of money that many influential people manage on a daily basis. The use doesn't have to be very practical for it to be well worth it.
This attack still do not allow for inserting a arbitrary data in arbitrary places to make attack on Linux possible. Finally SHA1 in git also take size into consideration and make this attack even more expensive[2].
People should really chill out. There are cheaper attack vectors that collisions.
[2] https://public-inbox.org/git/CA+55aFxJGDpJXqpcoPnwvzcn_fB-za...
One has to agree, an entity willing to drop a cool $100k on finding a single SHA1 collision to try and attack your git repo is a lot closer to nation-state level than the for-the-lulz level.
There's a lot of people for whom $100k is "fuck you" money.
English is fun that way. :)
There are plenty of non-state organizations for whom $110,000 is barely even pocket change.
Edit: wikipedia to the rescue! https://en.wikipedia.org/wiki/Nation_state#United_Kingdom
(My understanding of the method is it might be extendable to modifying a comment mid-file and then introducing later code, instead of modifying a JPG inside a PDF)
Currently the attack vector only works when you can get both documents to "work towards each other" to produce a valid identical SHA1 value.
SHA1(P | A | anything) = SHA1(P | B anything)
The Merkle-Damgård construction (used in MD4, MD5, SHA1 and SHA2 but not in SHA3 and some other modern hashes) invariably means length extension is possible, if you can collide two documents then you can add a suffix to both and also get a collision.
This is how there's already a web site where you feed it images and it makes a "different" colliding PDF, it's just using Google's result with a different suffix after the 128-byte collision near the start.
And there are already patches on the mailing list for that.
If you're using some weird way of getting a binary that you have already verified, but that could somehow differ, and you're hopping that git will catch the difference, you're doing it wrong to begin with.
You could just place a malicious one from the get go and no one would know (or they would know just as much -- blob do rely on virtually unconditional trust)
They're basically building that into git so that if this specific collision attack is ever used, git will notice and throw a warning/error.
"But if you use git for source control like in the kernel, the stuff you really care about is source code, which is very much a transparent medium. If somebody inserts random odd generated crud in the middle of your source code, you will absolutely notice. "
, which I think is a very weak argument.
The other thing which people seem to miss is that it requires 6,500 years of GPU computation for the _first_ phase of the SHA1 attack, and 110 years of the GPU compatation for the _second_ phase of the attack. You need to do both phases in order successfully carry out this attack. And even if you do, Google released code so that someone can easily tell if the object they were hashing was one created using this parituclar attach, which required 6,500 + 110 years of GPU computation.
But alas, it's a lot more fun to run around screaming that the sky is falling.....
"But if you use git for source control like in the kernel, the stuff you really care about is source code, which is very much a transparent medium. If somebody inserts random odd generated crud in the middle of your source code, you will absolutely notice. " , which I still think is a very weak argument.
It might or might not be true for any particular developer, and his argument does not refute the claim that the SHA1 integrity checks for that code is being rendered useless. I specifically recall that Linus previously described the hashed chain of commits as something which would prevent malicious insertion of code. And this has now, at least to some degree, been compromised.
He did provide some solid countermeasures and migration plans, but I think he could have been more acknowledging to all the people who predicted this attack. It would have been a good idea to prepare for changing hash function eventually.
http://stackoverflow.com/a/34599081 has actually gone about doing it, but it has been over a year since that, and as linus says, there has been multiple collision mitigations added as well, so tests should probably be re-done
Fixing this is going to require breaking backwards compatability with every program that works with git -- it's going to be a huge undertaking, because early in git's design they didn't support multiple hash functions.
The only way to detect without error two identical files is by comparing the files. This comparison can be speed up by comparing compressed version of the files.
The other functionality of hashes is to build a presumably unique file identifier. The byte sequence of the compressed file could serve as identifer.
So instead of using the file system as index with the sha1 name as file name, we would have to build a specific database organized as a set whose values (compressed files) would be the keys. A hash index could be used to speed up the search and equality test. Here a very fast hash would do the trick. Sha1 or a faster hash would be ok. The file system could then be used to organize the hash buckets as does git.
File comparision would of course first compare compressed and uncompressed file size. Or use other hashes or longer hash values to detect different files. When all these values are identical, then a file comparison must be performed to detect if we have a collision.
File compression can only get better and faster.
So basically git would only need to add hash collision detection and the capacity to support different objects with the same hash identifier.
And what about old releases that encounter a new repo?
And what about URLs and emails that reference commit hashes? Think archives of mailing lists that suddenly become useless unless there's a way to keep both hashes around.
Yes. These are all solvable problems (maybe not the old-release needing to handle new-style repo gracefully), but the complexity is much higher than upgrading a global constant.
We do not yet know if one way functions truly exists, so from a theoretical standpoint, any hash function is a weakpoint if you do not properly handle malicious collisions.
> SHA-1 is considered weak for many years, git should've migrated from it already.
Linus addressed this many years ago, when de was working on the first version of Git. It was chosen, despite the fact that it was known to be weakened. I don't know if they lost sight of this, or the geniuenly still believe that malicious colliding hashes are not a problem. I do not know enough about the intimate details of Git to comment on that fact.
https://arstechnica.com/security/2017/02/watershed-sha1-coll...
-S[<keyid>], --gpg-sign[=<keyid>]
GPG-sign commits. The keyid argument is optional and defaults to the committer identity; if specified, it must be stuck to the option without a space.
Yes it is.So a collision in a blob that represents a file (or any other internal git object) will cause in your old signature still being valid for the new file that corresponds to the git collision.
It's certainly possible to create a valid patch file that causes a collision, but it seems really hard to make a collision that looks like a valid pull request.
I understand your concern (I think) but consider all the extra stuff that has to happen for someone to accept a pr.
Edit.
I do agree it's time to start thinking about moving to the fire exits
Here, this is a good read: https://github.com/git/git/commit/e83c5163316f89bfbde7d9ab23...
For example, a file could be created that appears to be a normal source file at the front but contains some other behavior further down.
Git allows signing tags and commits, and those features are now broken because all object names use SHA1.
My understanding was you're using a public key type of encryption such as PGP at that point. I feel I may be missing the point here. (apologies if so)
However, as far as I know, when you sign a git commit you are actually sign the hash of the commit. With SHA-1 broken in the current way it essentially means someone with 110K to burn could forge a commit and reuse my PGP signature.
In addition, you'd have to have everyone else not notice it, all the insanely cheaper exploits not been tried on your current setup, and all the other stars aligning...
That might be a hint that Git isn't something which you should allow it to handle the security. Literally the first step to the entire thing: pick any email or name...
Please provide reasonable security policies in your repos--and if someone is exploited, you've probably got far bigger problems than someone duplicating a sha-1. Not necessarily, but highly likely your system is owned.
This kind of hairy distinction of what a signature was supposed to mean and what it actually covers is what you get with (semi-)broken cryptographic primitives. It's awful and, frankly, unnecessary.
1) you've got to rely on a lot better security than the minimal if at all security provided by git (it assumes a web of trust). If people are signing off with PGP sigs but not watching diffs, you've got big problems.
2) You're probably far more likely to be exploited by far cheaper methods at this point. If they have access to a trusted contributor's keys, it's far more cost effective to slip in other tricks than sha-1 collisions right now. I'd say this is the main point so far, but admittedly maybe not in the future.
3) It sounds like Linus and the git devs have admitted they need to migrate from sha-1, but also I haven't seen any cheap, exploitable PoC for git yet based on this due to how they actually mix in other info instead of raw sha-1 hashes of the files.
4) As far as I know, and I'm sure I'm subject to correction, but there hasn't been a WebKit svn repo-esque calamity yet like what they've experienced dropping 2 sha-1 collision PDFs into the repo in a Git context yet.
Again, I'm totally open to new info, but the sky-is-falling attitude right now is what I'm mainly arguing against.
"But the _real_ security comes from the fact that git is distributed, which means that a developer should never actually use a public tree for his development."
And then GitHub happened.
https://github.com/amoffat/masquerade/commit/9b0562595cc479a...
Yes, if you say you are billg@microsoft.com and make a commit to some repo on github, github will look up the username associated with billg@microsoft.com and show that user as the committer. Should it do that? Eh, probably not but this has come up a few times and github hasn't changed it. So by now we should just start to educate ourselves that this is how github is intended to work.
"In the next few years, nasty people will teach him the threat model"
I'd like to see those very forceful claims substantiated. Git hasn't said moving forward it will never change from sha-1 and 2005 was a far different era than 2017 for crypto. Let's keep that in perspective.
2) No, even easy "malicious" collisions in SHA-1 will still not break most of Git's usages. You're already trusting the repo you're pulling because of TLS, you're already trusting the commits you're getting because of peer-review (you read the commit) and a web-of-trust (you trust your collaborators). (And you're trusting commits even more when they're signed.)
The object store could be modified to support file collision. One way to disambiguate collision is to use a randomly generated byte sequence as SHA1 seed or hashed before the file data. This random byte sequence would behave like a salt and disable any forged collisions. A single seed for the whole repository would be enough. It should remain secret to prevent forging a collision with the two hashes. It's harder but not impossible.
To test if a given file is in the object store, one first compute the SHA1 key to use as file name. If no collisions ever occured with an object a file with that SHA1 name will be present in the store. That file contains the usual data plus the second hash computed with the random seed. This second hash could be added as needed to keep backward compatibility and provide silent automatic upgrade.
When one need to test if the file is present in the store, one computes the normal SHA1 key and the secondary hash with the seed. We locate the object in the store uisng the first SHA and test for file equality with the randomly seeded hash. Using a faster hash like blacke2 to compute the random seeded hash could mitigate the price to compute two hashes. It should be parameterized this time and the hash size should be variable.
If a collision is detected, that is the secondary hash differ, the file is replaced by a directory with the common SHA1 as name. The colliding files would be stored in the directory using the secondary hash as name. Or the files could be packed in a single tar like file with the secondary hash used as file identifier.
This should be enough to protect against forged collisions which is the only real problem. The required change to git would be limited. The only serious disadvantage is the need to compute the randomly seeded hash.
It should be trivial to add an additional checksum to git. Not to replace how SHA1 is currently used, but to add essentially per-commit checksum, which is a checksum of the entire commit contents (including the checksum of the previous commit). It wouldn't be as elegant as using SHA256 in place of SHA1, but at least you could, with some effort, validate the source tree in a cryptographically secure way.
This program is free software; you can redistribute it
and/or modify it under the terms of...commit -m "Added dutch license translation"
This attack won't work on plain text source because it won't look like source code.
If the mere existence of collisions is not acceptable to your VCS, then your VCS can't use a hash, period.
If you're worried about an intentional attack, it's no closer today than it was last week: the attacker doesn't control the output hash of the collision, or either input.
What? It could use a secure hash. Collisions take more energy than the universe has to discover.
(The point being that a VCS should handle collisions gracefully no matter what has is used.)
https://forums.whonix.org/t/security-git-general-verificatio...
In the retrospect many decisions which might have simplified current moment's problems are obvious, but in reality of those decisions in the past they never are. This feeling (it's called the retrospective predictability) is not a function of current problem or previous wrong decisions, but the random reality of events in complex systems, open source software implementing fresh ideas in a new way is being random and complex enough for this.
Gitlab blog: working on UX
Bitbucket: branch permissions
All of the teams should handle this issue as an emergency, and at least have a blog entry about their points of view.