The biggest issue is that git still uses it, which presents a problem if you want to protect a repo from active integrity attacks.
The biggest issue is that git still uses it, which presents a problem if you want to protect a repo from active integrity attacks.
There is also work to support SHA-256, though that seems to have stalled: https://lwn.net/Articles/898522/
The fundamental problem is that get developers assumed that hash algorithms would never be changed, and that was a ridiculous assumption. It's much wiser to implement crypto agility.
Cryptographic agility makes this problem worse, not better: instead of having a "flag day" (or release) where `git`'s digest choice reflects the State of the Art, agility ensures that every future version of `git` can be downgraded to a broken digest.
E.g. you will want to be able to read some sha-1-only repo from disk that was last touched a decade ago. That's a different thing than some protocol which requires both parties to be on-line, say wireguard, in which instance it's easier to switch both to a new version that uses a different cryptographic algorithm.
Git has such protocols as well, and maybe it can deprecate sha-1 support there eventually, but even there it has to support both sha-1 and sha-2 for a while because not everyone is using the latest and greatest version of git, and no sysadmin wants the absolute horror of flag days.
I don't mean that as some ridiculing criticism, I just am genuinely puzzled.
Because switching to a different hash algorithm would break compatibility with all existing Git clients and repositories.
Please pardon my ignorance but could you elaborate on what time (e.g. the year) are you referring to?
> The basic attack goes like this:
>
> - I construct two .c files with identical hashes.
Ok, I have a better plan.
- you learn to fly by flapping your arms fast enough
- you then learn to pee burning gasoline
- then, you fly around New York, setting everybody you see on fire, until
people make you emperor.
Sounds like a good plan, no?
But perhaps slightly impractical.
Now, let's go back to your plan. Why do you think your plan is any better
than mine?
https://git.vger.kernel.narkive.com/9lgv36un/zooko-zooko-com...Git not being prepared for this is going to cost a lot of time and money for a very large amount of people, and it could have been trivially mitigated if security were taken seriously in the first place, and if Torvalds was mature enough to understand the he is not an expert on cryptography topics.
git's first release was in 2005, so I guess technically SHA-1 issues could've been known or suspected during development time.
More generously, it could've been somewhat simultaneous. It sounds like it was considered a state-sponsored level attack at the time, if collisions were even going to be possible. Don't know if the git devs knew this and intentionally chose it anyway, or just didn't know.
[1] https://en.wikipedia.org/wiki/SHA-1
[2] https://www.schneier.com/blog/archives/2005/02/cryptanalysis...
EDIT: sibling comment has evidence that Linus did in fact know about it and considered it an impractical vector at the time
https://git.vger.kernel.narkive.com/9lgv36un/zooko-zooko-com...
Yet, the requirement of the hashing algorithm for Git is not broken, it's not cryptographic but merely stochastic, and Linus knows this.
Why bother to produce a collision, when you have the power to get your changes pulled into a release branch? Your attack might be noticed, and your cover blown.
Instead, simply try to get a bug merged that results in a zero day. In case somebody discovers it, at least you have plausible deniability that it happened on accident.
If by “perfectly fine” you mean “subject to attacks that generate somewhat targeted collisions that are practical enough that people do them for amusement and excuses to write blog posts and cute Twitter threads”, then maybe I agree.
Snark aside, SHA-1 is not fine for deduplication in any context where an attacker could control any inputs. Do not use it for new designs. Try to get rid of it in old designs.
Not every tool needs to be completely resilient to an entire Internets’ worth of attacks.
sometimes, i really feel like people in crypto just can't detach themselves enough to see that just because they have a hammer, not everything in the world is a nail.
Your comparison is flawed. It's more like if you have a nail and next to it a workbench with two hammers - a good hammer and a not as good hammer. This isn't a hard choice. But for reasons that are unclear to me, people in this thread are insisting on picking the less good hammer and rationalizing why for this specific nail it isn't all that much worse. Just pick the better hammer!
Any analysis about how hard it is for an attacker to get a file on your local file system via a cloned got repo, cached file, email attachment, image download, shared drive, etc is just a distraction.
BLAKE 3 is faster only in wall clock time, on an otherwise idle computer, because it fully uses all CPU cores, but it does not do less work.
BLAKE 3 is preferable only when the computer does nothing else but hashing.
On a modern intel CPU, one core of SHA1 does about 500MB/s worth of hashing. Blake3 on the same core is 1.5GB/s or faster.
Edit: Before anyone lecture me on SHA-1 being slow, yes, I use BLAKE2 for new projects.
If you are just using sha1 as a heuristic you dont fully trust, i suppose sha1 is fine. It seems a bit of an odd choice though as something like MurmurHash would be much faster for such a use case.
While such a scenario may be plausible for a public file repository, so SHA-1 is a bad choice for a version control system like GIT, there are a lot of applications where this is impossible, so it is fine to use SHA-1.
I also think working out all the possibilities is really hard, and using sha256 is really easy.
If we're a group of devs with a not insignificant percentage of those devs being frontend/UI/UX types, then having the same image in multiple sizes, formats, etc is going to be pretty common. Looking for multiples of the exact file is only going to reduce so much. Knowing you have a library of images with a source and then all of the derivatives is going to get you a lot less files as long as you know you have the source, then running image based sameness is much more beneficial. Sure, this is niche territory, but yeah, and, so?
Maybe there's someone new(-ish) that hasn't really had to deal with cleaning up thousands of images to this extent. One would hope the same image in its various forms within a dev's env would be similarly named, but that's not guaranteed. If we could depend on filenames, we wouldn't need hashing, right?
I agree that image editing workflows are a different use case more suited to perceptual hashes than cryptographic hashes.
But I don't know for sure that's the case.
DO NOT USE SHA-1 UNLESS IT’S FOR COMPATIBILITY. NO EXCUSES.
With that out of the way: SHA-1 is not even particularly fast. BLAKE2-family functions are faster. Quite a few modern hash functions are also parallelizable, and SHA-1 is not. If for some reason you need something faster than a fast modern hash, there are non-cryptographic hashes and checksums that are extraordinarily fast.
If you have several TB of files, and for some reason you use SHA-1 to dedupe them, and you later forget you did that and download one of the many pairs of amusing SHA-1 collisions, you will lose data. Stop making excuses.
Is it still true that CRC32 is only about twice as fast as SHA1?
Yeah I know the XX hashes are like 30 times faster than SHA1.
A lot depends on instruction set and processor choice.
Maybe another way to put it is I've always been impressed that on small systems SHA1 is enormously longer but only twice as slow as CRC32.
For a lot of interoperable-maxing non-security non-crypto tasks, CRC32 is not a bad choice, if its good enough for Ethernet, zmodem, and mpeg streams its good enough for my telemetry packets LOL. (IIRC iSCSI uses some variant different formulae)
For files, it is useless. Even if that was expected, I have computed CRC32 for all the files on an SSD. Of course, I have found thousands of collisions.
32 bits is too small to do the entire job of duplicate detection, but if it's fast enough then you can add a more thorough second pass and still save time.
These types of applications are usually using a cryptographic hash as one of a set of comparison functions that often start with file size as an optimization and might even include perceptual methods that are intentionally likely to produce collisions. Some will perform a byte-by-byte comparison as a final test, although just from a performance perspective this probably isn't worth the marginal improvement even for hash functions in which collisions are known to occur but vanishingly rare in organic data sets (this would include for example MD5 or even CRC at long bit lengths, but the lack of mixing in CRC makes organic collisions much more common with structured data).
SHA2 is significantly slower than SHA1 on many real platforms, so given that intentional collisions are not really part of the problem space few users would opt for the "upgrade" to SHA2. SHA1 itself isn't really a great choice because there are faster options with similar resistance to accidental collisions and worse resistance to intentional ones, but they're a lot less commonly known than the major cryptographic algorithms. Much of the literature on them is in the context of data structures and caching so the bit-lengths tend to be relatively small in that more collision-tolerant application and it's not always super clear how well they will perform at longer bit lengths (when capable).
Another way to consider this is from a threat modeling perspective: in a common file deduplication operation, when files come from non-trusted sources, someone might be able to exploit a second-preimage attack to generate a file that the deduplication tool will errantly consider a duplicate with another file, possibly resulting in one of the two being deleted if the tool takes automatic action. SHA1 actually remains highly resistant to preimage and second preimage attacks, so it's not likely that this is even feasible. SHA1 does have known collision attacks but these are unlikely to have any ramifications on a file deduplication system since both files would have to be generated by the adversary - that is, they can't modify the organic data set that they did not produce. I'm sure you could come up with an attack scenario that's feasible with SHA1 but I don't think it's one that would occur in reality. In any case, these types of tools are not generally being presented as resistant to malicious inputs.
If you're working in this problem space, a good thing to consider is hashing only subsets of the file contents, from multiple offsets to avoid collisions induced by structured parts of the format. This avoids the need to read in the entire file for the initial hash-matching heuristic. Some commercial tools initially perform comparisons on only the beginning of the file (e.g. first MB) but for some types of files this is going to be a lot more collision prone than if you incorporate samples from regular intervals, e.g. skipping over every so many storage blocks.
When hashing hundreds of GB or many TB of data, the hash speed is important.
When there are no active attackers and even against certain kinds of active attacks, SHA-1 remains secure.
For example, if hashes of the files from a file system are stored separately, in a secure place inaccessible for attackers (or in the case of a file transfer the hashes are transferred separately, through a secure channel), an attacker cannot make a file modification that would not be detected by recomputing the hashes.
Even if SHA-1 remains secure against preimage attacks, it should normally be used only when there are no attackers, e.g. for detecting hardware errors a.k.a. bit rotting, or for detecting duplicate data in storage that could not be accessed by an attacker.
While BLAKE 3 (not BLAKE 2) can be much faster than SHA-1, all the extra speed is obtained by consuming proportionally more CPU resources (extra threads and SIMD). When the hashing is done in background, there is no gain by using BLAKE 3 instead of SHA-1, because the foreground tasks will be delayed by the time gained for hashing.
Only when a computer does only hashing, BLAKE 3 is the best choice, because the hash will be computed in a minimal time, by fully using all the CPU cores.
If you know you have other threads that need to do work, then yes, multithreading BLAKE3 would just pointlessly compete with those other threads. But I don't think the same is true of SIMD. If your process/thread isn't using vector registers, it's not like some other thread can borrow them. They just sit idle. So if you can make use of them to speed up your own process, there's very little downside. AVX-512 downclocking is the most notable exception, and you'd need to benchmark your application to see whether / how much that hurts you. But I think in most other cases, any power draw penalty you pay for using SIMD is swamped by the race-to-idle upside. (I don't have much experience measuring power, though, and I'd be happy to get corrected by someone who knows more.)
It depends what you are doing, but deduplication where a collision means you loose data, seems like an inapropriate place for sha-1.
I believe the GP's point hinges on the word "attacker". If you aren't in a hostile space, like just your won file server and you are monitoring your own backups it's fine. I still use MD5s to version my own config files. For personal use in non-hostile environments these hashes are still perfectly fine.
That's what the developers of subversion thought, but they didn't anticipate that once colliding files were available people would commit them to SVN repos as test cases. And then everything broke: https://www.bleepingcomputer.com/news/security/sha1-collisio...
You use digests to quickly detect potential collisions, then you verify each collision report, then you delete the actual duplicates. Human involvement still very much required because you're curating your own data.
if you want to dedupe images, some sort of phashing would be much better so that the actual image is considered vs just the specific bits to generate the image.
(I suspect this would be a good compromise for git, since so much tooling assumes a 160 bit hash, and yet we don't want to continue using SHA1)
SHA1 (and MD5) need to be treated the same way you would treat O(n^2) sorting in a code review for a PR written by a newbie.
SHA1 and MD5 are the most widely accessible, though, and I agree it's fine to use them if you don't care about security.
Thus, a construct like hash(key + message) can be used similar to SHA3 [1]
If you dont care about security, use a faster hash. If you care about security use sha256 (which is about the same speed anyways).
The only valid reason to still use it in non-security critical roles is backwards compat.
The emphasis being on "for security"
I've also used SHA-1 over the years for binning and verifying file transfer success, none of those are security related.
Sometimes, if you make a great big pile of different systems, what's held in common across them can be weird, SHA-1 popped out of the list so we used it.
I'm well aware its possible to write or automate the writing of dedicated specialized "perfect" hashing algos to match the incoming data, to bin the data more perfectlyier, but sometimes its nice if wildly separate systems all bin incoming data the same highly predictable way thats "good enough" and "fast enough".
It could. If you want to verify that the file has not been tempered by someone, it is security related.
Verified as in "is this file completely transferred or not?"
non-security critical data, I just want a general idea if its valid or the file transfer failed half way thru or the thing sending it went bonkers and just sent trash to us.
Another funny file transfer use: Send me a file of data every hour. Is the non-crypto-hash new or the same old hash? If its the same old hash, those clowns sent me the same file twice, I'm supposed to get a new one. Yes I know I can dedupe "easily" but not as "easily" as sha-1. And some application layer software like MySQL can directly generate SHA1 as a function in the query. Its really quite handy sometimes!
Absolutely agree, especially when speed is a workable trade-off and accepting real world hash collisions are unlikely and perhaps an acceptable risk. For financial data, especially files not belonging to me I would have md5+sha1+sha256 checksums and maybe even GPG sign a manifest of the checksums ... because why not. For my own files md5 has always been sufficient. I have yet to run into a real world collision.
FWIW anyone using `rsync --checksum` is still using md5. Not that long ago I think 2014 it was using md4. I would be surprised if rsync started using anything beyond md5 any time soon. I would love to see all the checksum algorithms become CPU instruction sets.
Optimizations:
no SIMD-roll, no asm-roll, no openssl-crypto, asm-MD5
Checksum list:
md5 md4 none
Compress list:
zstd lz4 zlibx zlib none
Daemon auth list:
md5 md4The SHA-1 collision attack can only work if you take a specially-crafted file from the attacker and commit it to your repository. The file needs to have a specific structure, and will contain binary data that looks like junk. It can't look like innocent source code. If you execute unintelligible binary blobs from strangers, you're in trouble anyway.
There is no preimage weakness in SHA-1, so nobody is able to change or inject new data to an arbitrary repo/commit that doesn't already contain their colliding file.
Pretty sure you can anyway, I haven't thought deeply about the git file formats involved and such.
Note that this attack isn't _that_ serious. There's not a lot of cases where this would make sense.
As far as I know, with current public SHA-1 vulnerabilities, you can create two new objects with the same hash (collision attack), but cannot create a second object that has the same hash as some already existing object (preimage attack).
Given the limitations, really not too practical.
In particular, CRC guarantees detection on all bursts of the given length. CRC32 protects vs all bursts of length 32 bits.
For the sake of the next person who has to maintain your code though, please choose algorithms that adequately communicate your intentions. Choose CRCs only if you need to detect random errors in a noisy channel with a small number of bits and use a length appropriate to the intended usage (i.e. almost certainly not CRC64).
CRC says that you never intended security from the start. It's timeless, aimed to prevent burst errors and random errors.
--------
BTW, what is the guaranteed Hamming distance between SHA1? How good is SHA1 vs burst errors? What about random errors?
Because the Hamming distances of CRC have been calculated and analyzed. We actually can determine, to an exact level, how good CRC codes are.
CRC should be better for any error detection code issue. Faster to calculate, more studied guaranteed detection modes, and so forth.
SHA1 has no error detection studies. It's designed as a cryptographic hash, to look random. As it so happens, it is more efficient to use other algorithms and do better than random if you have a better idea of how your error looks like.
Real world errors are either random or bursty. CRC is designed for these cases. CRC detects the longest burst possible for it's bitsize.
The same is true of CRCs over a large enough input as an aside.
Basically after ~10 rounds the output is always indistinguishable from randomness which means hamming distance is what you'd expect (about half the bits differ) between the hashes of any two bitstreams.
If you are only worried about random errors, might as well use chksum.
THIS IS FALSE. Please do not ever do this. Why not? For example, by controlling any four contiguous bytes in a file, the resultant 32bit CRC can be forced to take on any value. A CRC is meant to detect errors due to noise - not changes due to a malicious actor.
Program design should not be done based upon one's feelings. CRCs absolutely do not have the required properties to detect duplication or to preserve integrity of a stored file that an attacker can modify.
And SHA1 is now broken like this, with collisions and so forth. Perhaps it's not as simple as just 4 bytes, but the ability to create collisions is forcing this retirement.
If adversarial collisions are an issue, then MD5 and SHA1 are fully obsolete now. If you don't care for an adversary, might as well use the cheaper, faster CRC check.
------
CRC is now more valid use case than SHA1. That's the point of this announcement.
No. That isn't the point of this announcement. This announcement codifies a transition time for SHA-1 use in SBU and has nothing to do with a CRC.
This is relevant to when you'd choose CRC, because it also has no real security.
That's also false. There is a large body of knowledge here that you aren't expressing in your comments. That leads me to see that you are unfamiliar with the purposes of hash functions and their utility in real world situations.
The announcement refers to the transition timeline to stop using SHA-1, preferring the SHA-2 and SHA-3 families. However, the recommendations for years from NIST have been not to use SHA-1. For example SP 800-131Ar2 discusses not to use SHA-1 for digital sig gen and that digital sig ver is only acceptable for legacy uses.
The recommendation would have been for years to not use SHA-1 at all, except for this carve-out to handle already stored data that uses SHA-1. The remaining use cases cover protocol use, such as TLS, where SHA-1 is used as a component in constructs and not solely as a primitive.
Don't post like this, please.