NIST is announcing that SHA-1 should be phased out by Dec. 31, 2030
nist.gov
nist.gov
Does anyone in the field know if there's a SHA-4 competition around the corner, or is SHA-3 good enough? It would be interesting if one of the SHA-4 requirements was Zero Knowledge Proof friendlyness. MiMC and Poseidon are two of the big contenders in that area, and it would be really great to see a full NIST quality onslaught against them.
https://www.schneier.com/blog/archives/2013/10/will_keccak_s...
From your first link:
I misspoke when I wrote that NIST made “internal changes” to the algorithm. That was sloppy of me. The Keccak permutation remains unchanged. What NIST proposed was reducing the hash function’s capacity in the name of performance. One of Keccak’s nice features is that it’s highly tunable.
I do not believe that the NIST changes were suggested by the NSA. Nor do I believe that the changes make the algorithm easier to break by the NSA. I believe NIST made the changes in good faith, and the result is a better security/performance trade-off.
jk I’m sure it’s fine… probably
I still think it's probably fine, but I feel better about insisting on SHA-2+mitigations or blake3 instead now, even if the main problem with SHA-3 is its being deliberately designed to encourage specialized hardware accelleration (cf AES and things like Intel's aes-ni).
(To be clear, the fact that Schneier claims to "believe NIST made the changes in good faith" is weak but nonzero evidence that they did not. I don't see any concrete evidence for a backdoor, although you obviously shouldn't trust me either.)
Do you think he’d disclose it if he did?
Making late (in this case, after the competition was already over) changes to a crytographic primitive - without extensive documentation both of why that's necessary (not just helpful) and why it's not possible (not just you promise it doesn't) for that to weaken security or insert backdoors - is a act of either bad faith or sufficiently gross incompetence that it should be considered de facto bad faith.
Schneier claiming to believe that it's good faith implies that he either doesn't understand or (presumably more likely given his history?) doesn't care about keeping the standardization process secure against corruption by the NSA, which suggests either incompetence or bad faith on Schneier's part as well. (Or, in context, that someone was leaning on him after the earlier criticism.)
This is particularly inexcusable since reducing security parameters on the pretense that "56 bits ought to be enough for anybody" is a known NSA tactic dating back to fucking DES.
The new padding prepends either 01, 11, or 1111 depending on the variant of SHA-3. That way the different variants don't sometimes give the same [partial] hash.
It was weird to toss that in at the last second but there's no room for a backdoor there.
But I've had a soft spot for Schneier since this (pre-Snowden) blog post on Dual_EC_DRBG: https://www.schneier.com/blog/archives/2007/11/the_strange_s....
> NIST's current proposal for SHA-3, namely the one presented by John Kelsey at CHES 2013 in August, is a subset of the Keccak family. More concretely, one can generate the test vectors for that proposal using the Keccak reference code (version 3.0 and later, January 2011). This alone shows that the proposal cannot contain internal changes to the algorithm.
In fact, one actually has done almost exactly that: https://en.wikipedia.org/wiki/Data_Encryption_Standard#NSA's...
Blake3 was designed from the ground up to be highly optimized for vector instructions operating on four vectors, each of 4 32-bit words. If you already have the usual 4x32 vector operations, plus a vector permute (to transform the operations across the diagonal of your 4x4 matrix into operations down columns) and the usual bypass network to reduce latency, I think it would rarely be worth the transistor budget to create dedicated Blake3 (or Blake2s/b) instructions.
In contrast, SHA-3's state is conceptually five vectors each of five 32-bit words, which doesn't map as neatly onto most vector ISAs. As I remember, it has column and row operations rather than column and diagonal operations that parallelize better on vector hardware.
SHA-2 is a Merkle-Damgard construction where the round function is a Davies-Meyer construction where the internal block cipher is a highly unbalanced Feistel cipher. Conceptually, you have a queue of 8 words (32-bit or 64-bit, depending on which variant). Each operation pops the first word from the queue, combines it in a nonlinear way with 6 of the other words, adds one word from the "key schedule" derived from the message, and pushes the result on the back of the queue. The one word that wasn't otherwise used is increased by the sum of the round key and a non-linear function of 3 of the other words. As you might imagine, this doesn't map very well onto general-purpose vector instructions. This cipher is wrapped in a step (Davies-Meyer construction) where you save a copy of the state, encrypt the state using the next block of the message, and then add the saved copy to the encrypted result (making it non-invertible, making meet-in-the middle attacks much more difficult). The key schedule uses a variation on a lagged Fibonacci generator to expand each message block into a larger number of round keys.
This is true, and the BLAKE family inherits this structure from ChaCha, but there's also more to it than that. If you have enough input to fill many blocks, you can run multiple blocks in parallel. In this situation, rather than dividing up the 16 words of a block into four vectors, you put each word in a different vector, and the words of each vector represent the same position in different blocks. (I.e rather than representing columns or rows, the vectors point "out of the page.) There are several benefits to this arrangement:
1. You don't need to do that diagonalization operation anymore.
2. If your CPU supports "instruction-level parallelism" for vector operation, working across the four words/vectors in a row gets to take advantage of that.
3. Best of all, you're no longer limited to 4-word vectors. If you have enough input to fill 8 blocks (AVX2) or 16 blocks (AVX-512), you can use those much larger instruction sets.
This is all easy to take advantage of in a stream cipher like ChaCha, because each block is independent. With a hash function, things are more complicated, because you usually have data dependencies between different blocks. That's why the tree structure of BLAKE3 (or somewhat similarly, KangarooTwelve) is so important for performance. It's not just about multithreading; it also about SIMD. See section 5.3 of the BLAKE3 paper for more on this.
I am shocked how fast it is. Just tried it on several files. It is blazing fast.
brew install b3sum on mac.
Edit: Never mind, grand-uncle comment beat me to it:
For what applications is it not good enough, and how/why do we believe that?
(You could contort yourself into an argument that "SHA2 isn't good enough for protocols where you need a keyed hash without HMAC", but 1) that isn't true given SHA2-384 and 2) there really are no such protocols).
The SHA1 break doesn't threaten SHA2. The two hashes are different in a significant way that breaks the SHA1 attack.
Hmm, I think you might be overstating this a bit. While the SHA1 attack indeed does not break SHA2, the whole reason SHA3 exists at all is that SHA1 and SHA2 are similar enough in their structure that we were worried that the methods used in the SHA1 break could be extended to attack SHA2.
So far the answer seems to be no, but it was a serious concern for a while.
Is it a serious concern among cryptographers outside of NIST?
late edit
I removed the scare quotes around "differential paths" after skimming the first Stevens paper and the Chabaud-Joux paper and confirming they were using "differential" they way I understand the term. :)
Also, I'm more confident that the message schedule is pretty central to the attack, and I guess the whole line of research that led up to it?
Is it a serious concern among cryptographers outside of NIST?
Today? Not really, because we've had years of research showing us that SHA2 still seems to be safe. But when the major breaks of SHA1 were happening? Absolutely. In my "everything you need to know about crypto in 1 hour" talk I explicitly said "use SHA2 but be ready to move to SHA3 if needed because the attacks on SHA1 are scary and we're worried they could generalize to SHA2 as well".
As for your points about anticipated benefits of the SHA2 message schedule: isn't part of the point of SHA2 that it's more ARX-y? Which is basically what makes the Shattered line of attacks not viable?
I think we're in the same place in terms of recommendations! Except: I don't know that "the SHA1 attacks could generalize" is all that valid a concern? Regardless, the "why" of this is super interesting and I don't think anyone has broken it down super clearly (I'm not doing it by myself; I don't have the chops).
I don't know that the NSA has ever published the internal discussions which resulted in SHA2, but I always thought the primary design purpose of SHA2 was to produce larger hashes (and thus avoid the 2^80 birthday attack on SHA1).
Except: I don't know that "the SHA1 attacks could generalize" is all that valid a concern?
It's not, now. It was in 2005/2006. Remember there wasn't just one SHA1 attack; there was a whole series of them. (And I'm guessing the NSA was particularly concerned given that some of the attacks were discovered by Chinese researchers.)
It will also be really nice on amd_64 CPUs for doing large amounts of hashing if they ever get a sha3 instruction.
That said, hashing is rarely your bottleneck so there's not much urgency.
SHA-3 is by far and away the most elegant modern cryptographic hash algorithm.
Given that NIST itself warns PQC algorithm may be unsafe after 2035, this should be considered SHA-4.
Sha is not a digital signature algorithm. That is a different type of crypto primitive.
I thought post-quantum cryptography was supposed to be futureproof? Or am I misunderstanding
What would this requirement look like and why would it be important for a hash function?
My boring hash function of choice is Sha-384. The Sha-512 computation is faster on Intel hardware, and ASICS to crack it are far more expensive than Sha-256 because of bitcoin.
If you're hashing passwords or something, use a "harder" hash like Argon2 or Scrypt.
On Intel Atom starting with Apollo Lake (2016) and on Intel Core starting with Ice Lake (2019) and on all AMD Zen CPUs (2017), SHA-256 is implemented in hardware and it is much faster than SHA-512.
SHA-384 is a truncated SHA-512. From the claims of sec people it does not offer more security when it comes to length attacks. But from how the algo works I would assume that it does.
Nist is also plain wrong about their calculations. Cause how long it takes to calculate a specific hash depends on the hardware available, not what theory books says. It may in practice be faster to calculate a hash with more bits.
Depends on the type of break. If the break only allows finding a hash with 128+k leading zeroes in 2^{128+k/2} time, that would still be quite useless for bitcoin mining.
The break would have to cover the bitcoin regime of around 78 leading 0s.
The biggest issue is that git still uses it, which presents a problem if you want to protect a repo from active integrity attacks.
In particular, CRC guarantees detection on all bursts of the given length. CRC32 protects vs all bursts of length 32 bits.
For the sake of the next person who has to maintain your code though, please choose algorithms that adequately communicate your intentions. Choose CRCs only if you need to detect random errors in a noisy channel with a small number of bits and use a length appropriate to the intended usage (i.e. almost certainly not CRC64).
CRC says that you never intended security from the start. It's timeless, aimed to prevent burst errors and random errors.
--------
BTW, what is the guaranteed Hamming distance between SHA1? How good is SHA1 vs burst errors? What about random errors?
Because the Hamming distances of CRC have been calculated and analyzed. We actually can determine, to an exact level, how good CRC codes are.
CRC should be better for any error detection code issue. Faster to calculate, more studied guaranteed detection modes, and so forth.
SHA1 has no error detection studies. It's designed as a cryptographic hash, to look random. As it so happens, it is more efficient to use other algorithms and do better than random if you have a better idea of how your error looks like.
Real world errors are either random or bursty. CRC is designed for these cases. CRC detects the longest burst possible for it's bitsize.
The same is true of CRCs over a large enough input as an aside.
Basically after ~10 rounds the output is always indistinguishable from randomness which means hamming distance is what you'd expect (about half the bits differ) between the hashes of any two bitstreams.
If you are only worried about random errors, might as well use chksum.
THIS IS FALSE. Please do not ever do this. Why not? For example, by controlling any four contiguous bytes in a file, the resultant 32bit CRC can be forced to take on any value. A CRC is meant to detect errors due to noise - not changes due to a malicious actor.
Program design should not be done based upon one's feelings. CRCs absolutely do not have the required properties to detect duplication or to preserve integrity of a stored file that an attacker can modify.
And SHA1 is now broken like this, with collisions and so forth. Perhaps it's not as simple as just 4 bytes, but the ability to create collisions is forcing this retirement.
If adversarial collisions are an issue, then MD5 and SHA1 are fully obsolete now. If you don't care for an adversary, might as well use the cheaper, faster CRC check.
------
CRC is now more valid use case than SHA1. That's the point of this announcement.
No. That isn't the point of this announcement. This announcement codifies a transition time for SHA-1 use in SBU and has nothing to do with a CRC.
This is relevant to when you'd choose CRC, because it also has no real security.
That's also false. There is a large body of knowledge here that you aren't expressing in your comments. That leads me to see that you are unfamiliar with the purposes of hash functions and their utility in real world situations.
The announcement refers to the transition timeline to stop using SHA-1, preferring the SHA-2 and SHA-3 families. However, the recommendations for years from NIST have been not to use SHA-1. For example SP 800-131Ar2 discusses not to use SHA-1 for digital sig gen and that digital sig ver is only acceptable for legacy uses.
The recommendation would have been for years to not use SHA-1 at all, except for this carve-out to handle already stored data that uses SHA-1. The remaining use cases cover protocol use, such as TLS, where SHA-1 is used as a component in constructs and not solely as a primitive.
Don't post like this, please.
(I suspect this would be a good compromise for git, since so much tooling assumes a 160 bit hash, and yet we don't want to continue using SHA1)
Pretty sure you can anyway, I haven't thought deeply about the git file formats involved and such.
Note that this attack isn't _that_ serious. There's not a lot of cases where this would make sense.
As far as I know, with current public SHA-1 vulnerabilities, you can create two new objects with the same hash (collision attack), but cannot create a second object that has the same hash as some already existing object (preimage attack).
Given the limitations, really not too practical.
Thus, a construct like hash(key + message) can be used similar to SHA3 [1]
If by “perfectly fine” you mean “subject to attacks that generate somewhat targeted collisions that are practical enough that people do them for amusement and excuses to write blog posts and cute Twitter threads”, then maybe I agree.
Snark aside, SHA-1 is not fine for deduplication in any context where an attacker could control any inputs. Do not use it for new designs. Try to get rid of it in old designs.
Not every tool needs to be completely resilient to an entire Internets’ worth of attacks.
sometimes, i really feel like people in crypto just can't detach themselves enough to see that just because they have a hammer, not everything in the world is a nail.
Your comparison is flawed. It's more like if you have a nail and next to it a workbench with two hammers - a good hammer and a not as good hammer. This isn't a hard choice. But for reasons that are unclear to me, people in this thread are insisting on picking the less good hammer and rationalizing why for this specific nail it isn't all that much worse. Just pick the better hammer!
Any analysis about how hard it is for an attacker to get a file on your local file system via a cloned got repo, cached file, email attachment, image download, shared drive, etc is just a distraction.
BLAKE 3 is faster only in wall clock time, on an otherwise idle computer, because it fully uses all CPU cores, but it does not do less work.
BLAKE 3 is preferable only when the computer does nothing else but hashing.
On a modern intel CPU, one core of SHA1 does about 500MB/s worth of hashing. Blake3 on the same core is 1.5GB/s or faster.
Edit: Before anyone lecture me on SHA-1 being slow, yes, I use BLAKE2 for new projects.
If you are just using sha1 as a heuristic you dont fully trust, i suppose sha1 is fine. It seems a bit of an odd choice though as something like MurmurHash would be much faster for such a use case.
While such a scenario may be plausible for a public file repository, so SHA-1 is a bad choice for a version control system like GIT, there are a lot of applications where this is impossible, so it is fine to use SHA-1.
I also think working out all the possibilities is really hard, and using sha256 is really easy.
If we're a group of devs with a not insignificant percentage of those devs being frontend/UI/UX types, then having the same image in multiple sizes, formats, etc is going to be pretty common. Looking for multiples of the exact file is only going to reduce so much. Knowing you have a library of images with a source and then all of the derivatives is going to get you a lot less files as long as you know you have the source, then running image based sameness is much more beneficial. Sure, this is niche territory, but yeah, and, so?
Maybe there's someone new(-ish) that hasn't really had to deal with cleaning up thousands of images to this extent. One would hope the same image in its various forms within a dev's env would be similarly named, but that's not guaranteed. If we could depend on filenames, we wouldn't need hashing, right?
I agree that image editing workflows are a different use case more suited to perceptual hashes than cryptographic hashes.
But I don't know for sure that's the case.
DO NOT USE SHA-1 UNLESS IT’S FOR COMPATIBILITY. NO EXCUSES.
With that out of the way: SHA-1 is not even particularly fast. BLAKE2-family functions are faster. Quite a few modern hash functions are also parallelizable, and SHA-1 is not. If for some reason you need something faster than a fast modern hash, there are non-cryptographic hashes and checksums that are extraordinarily fast.
If you have several TB of files, and for some reason you use SHA-1 to dedupe them, and you later forget you did that and download one of the many pairs of amusing SHA-1 collisions, you will lose data. Stop making excuses.
Is it still true that CRC32 is only about twice as fast as SHA1?
Yeah I know the XX hashes are like 30 times faster than SHA1.
A lot depends on instruction set and processor choice.
Maybe another way to put it is I've always been impressed that on small systems SHA1 is enormously longer but only twice as slow as CRC32.
For a lot of interoperable-maxing non-security non-crypto tasks, CRC32 is not a bad choice, if its good enough for Ethernet, zmodem, and mpeg streams its good enough for my telemetry packets LOL. (IIRC iSCSI uses some variant different formulae)
For files, it is useless. Even if that was expected, I have computed CRC32 for all the files on an SSD. Of course, I have found thousands of collisions.
32 bits is too small to do the entire job of duplicate detection, but if it's fast enough then you can add a more thorough second pass and still save time.
These types of applications are usually using a cryptographic hash as one of a set of comparison functions that often start with file size as an optimization and might even include perceptual methods that are intentionally likely to produce collisions. Some will perform a byte-by-byte comparison as a final test, although just from a performance perspective this probably isn't worth the marginal improvement even for hash functions in which collisions are known to occur but vanishingly rare in organic data sets (this would include for example MD5 or even CRC at long bit lengths, but the lack of mixing in CRC makes organic collisions much more common with structured data).
SHA2 is significantly slower than SHA1 on many real platforms, so given that intentional collisions are not really part of the problem space few users would opt for the "upgrade" to SHA2. SHA1 itself isn't really a great choice because there are faster options with similar resistance to accidental collisions and worse resistance to intentional ones, but they're a lot less commonly known than the major cryptographic algorithms. Much of the literature on them is in the context of data structures and caching so the bit-lengths tend to be relatively small in that more collision-tolerant application and it's not always super clear how well they will perform at longer bit lengths (when capable).
Another way to consider this is from a threat modeling perspective: in a common file deduplication operation, when files come from non-trusted sources, someone might be able to exploit a second-preimage attack to generate a file that the deduplication tool will errantly consider a duplicate with another file, possibly resulting in one of the two being deleted if the tool takes automatic action. SHA1 actually remains highly resistant to preimage and second preimage attacks, so it's not likely that this is even feasible. SHA1 does have known collision attacks but these are unlikely to have any ramifications on a file deduplication system since both files would have to be generated by the adversary - that is, they can't modify the organic data set that they did not produce. I'm sure you could come up with an attack scenario that's feasible with SHA1 but I don't think it's one that would occur in reality. In any case, these types of tools are not generally being presented as resistant to malicious inputs.
If you're working in this problem space, a good thing to consider is hashing only subsets of the file contents, from multiple offsets to avoid collisions induced by structured parts of the format. This avoids the need to read in the entire file for the initial hash-matching heuristic. Some commercial tools initially perform comparisons on only the beginning of the file (e.g. first MB) but for some types of files this is going to be a lot more collision prone than if you incorporate samples from regular intervals, e.g. skipping over every so many storage blocks.
When hashing hundreds of GB or many TB of data, the hash speed is important.
When there are no active attackers and even against certain kinds of active attacks, SHA-1 remains secure.
For example, if hashes of the files from a file system are stored separately, in a secure place inaccessible for attackers (or in the case of a file transfer the hashes are transferred separately, through a secure channel), an attacker cannot make a file modification that would not be detected by recomputing the hashes.
Even if SHA-1 remains secure against preimage attacks, it should normally be used only when there are no attackers, e.g. for detecting hardware errors a.k.a. bit rotting, or for detecting duplicate data in storage that could not be accessed by an attacker.
While BLAKE 3 (not BLAKE 2) can be much faster than SHA-1, all the extra speed is obtained by consuming proportionally more CPU resources (extra threads and SIMD). When the hashing is done in background, there is no gain by using BLAKE 3 instead of SHA-1, because the foreground tasks will be delayed by the time gained for hashing.
Only when a computer does only hashing, BLAKE 3 is the best choice, because the hash will be computed in a minimal time, by fully using all the CPU cores.
If you know you have other threads that need to do work, then yes, multithreading BLAKE3 would just pointlessly compete with those other threads. But I don't think the same is true of SIMD. If your process/thread isn't using vector registers, it's not like some other thread can borrow them. They just sit idle. So if you can make use of them to speed up your own process, there's very little downside. AVX-512 downclocking is the most notable exception, and you'd need to benchmark your application to see whether / how much that hurts you. But I think in most other cases, any power draw penalty you pay for using SIMD is swamped by the race-to-idle upside. (I don't have much experience measuring power, though, and I'd be happy to get corrected by someone who knows more.)
It depends what you are doing, but deduplication where a collision means you loose data, seems like an inapropriate place for sha-1.
I believe the GP's point hinges on the word "attacker". If you aren't in a hostile space, like just your won file server and you are monitoring your own backups it's fine. I still use MD5s to version my own config files. For personal use in non-hostile environments these hashes are still perfectly fine.
There is also work to support SHA-256, though that seems to have stalled: https://lwn.net/Articles/898522/
The fundamental problem is that get developers assumed that hash algorithms would never be changed, and that was a ridiculous assumption. It's much wiser to implement crypto agility.
Cryptographic agility makes this problem worse, not better: instead of having a "flag day" (or release) where `git`'s digest choice reflects the State of the Art, agility ensures that every future version of `git` can be downgraded to a broken digest.
E.g. you will want to be able to read some sha-1-only repo from disk that was last touched a decade ago. That's a different thing than some protocol which requires both parties to be on-line, say wireguard, in which instance it's easier to switch both to a new version that uses a different cryptographic algorithm.
Git has such protocols as well, and maybe it can deprecate sha-1 support there eventually, but even there it has to support both sha-1 and sha-2 for a while because not everyone is using the latest and greatest version of git, and no sysadmin wants the absolute horror of flag days.
I don't mean that as some ridiculing criticism, I just am genuinely puzzled.
Because switching to a different hash algorithm would break compatibility with all existing Git clients and repositories.
Please pardon my ignorance but could you elaborate on what time (e.g. the year) are you referring to?
> The basic attack goes like this:
>
> - I construct two .c files with identical hashes.
Ok, I have a better plan.
- you learn to fly by flapping your arms fast enough
- you then learn to pee burning gasoline
- then, you fly around New York, setting everybody you see on fire, until
people make you emperor.
Sounds like a good plan, no?
But perhaps slightly impractical.
Now, let's go back to your plan. Why do you think your plan is any better
than mine?
https://git.vger.kernel.narkive.com/9lgv36un/zooko-zooko-com...Git not being prepared for this is going to cost a lot of time and money for a very large amount of people, and it could have been trivially mitigated if security were taken seriously in the first place, and if Torvalds was mature enough to understand the he is not an expert on cryptography topics.
git's first release was in 2005, so I guess technically SHA-1 issues could've been known or suspected during development time.
More generously, it could've been somewhat simultaneous. It sounds like it was considered a state-sponsored level attack at the time, if collisions were even going to be possible. Don't know if the git devs knew this and intentionally chose it anyway, or just didn't know.
[1] https://en.wikipedia.org/wiki/SHA-1
[2] https://www.schneier.com/blog/archives/2005/02/cryptanalysis...
EDIT: sibling comment has evidence that Linus did in fact know about it and considered it an impractical vector at the time
https://git.vger.kernel.narkive.com/9lgv36un/zooko-zooko-com...
Yet, the requirement of the hashing algorithm for Git is not broken, it's not cryptographic but merely stochastic, and Linus knows this.
Why bother to produce a collision, when you have the power to get your changes pulled into a release branch? Your attack might be noticed, and your cover blown.
Instead, simply try to get a bug merged that results in a zero day. In case somebody discovers it, at least you have plausible deniability that it happened on accident.
That's what the developers of subversion thought, but they didn't anticipate that once colliding files were available people would commit them to SVN repos as test cases. And then everything broke: https://www.bleepingcomputer.com/news/security/sha1-collisio...
You use digests to quickly detect potential collisions, then you verify each collision report, then you delete the actual duplicates. Human involvement still very much required because you're curating your own data.
if you want to dedupe images, some sort of phashing would be much better so that the actual image is considered vs just the specific bits to generate the image.
SHA1 (and MD5) need to be treated the same way you would treat O(n^2) sorting in a code review for a PR written by a newbie.
SHA1 and MD5 are the most widely accessible, though, and I agree it's fine to use them if you don't care about security.
If you dont care about security, use a faster hash. If you care about security use sha256 (which is about the same speed anyways).
The only valid reason to still use it in non-security critical roles is backwards compat.
The emphasis being on "for security"
I've also used SHA-1 over the years for binning and verifying file transfer success, none of those are security related.
Sometimes, if you make a great big pile of different systems, what's held in common across them can be weird, SHA-1 popped out of the list so we used it.
I'm well aware its possible to write or automate the writing of dedicated specialized "perfect" hashing algos to match the incoming data, to bin the data more perfectlyier, but sometimes its nice if wildly separate systems all bin incoming data the same highly predictable way thats "good enough" and "fast enough".
It could. If you want to verify that the file has not been tempered by someone, it is security related.
Verified as in "is this file completely transferred or not?"
non-security critical data, I just want a general idea if its valid or the file transfer failed half way thru or the thing sending it went bonkers and just sent trash to us.
Another funny file transfer use: Send me a file of data every hour. Is the non-crypto-hash new or the same old hash? If its the same old hash, those clowns sent me the same file twice, I'm supposed to get a new one. Yes I know I can dedupe "easily" but not as "easily" as sha-1. And some application layer software like MySQL can directly generate SHA1 as a function in the query. Its really quite handy sometimes!
Absolutely agree, especially when speed is a workable trade-off and accepting real world hash collisions are unlikely and perhaps an acceptable risk. For financial data, especially files not belonging to me I would have md5+sha1+sha256 checksums and maybe even GPG sign a manifest of the checksums ... because why not. For my own files md5 has always been sufficient. I have yet to run into a real world collision.
FWIW anyone using `rsync --checksum` is still using md5. Not that long ago I think 2014 it was using md4. I would be surprised if rsync started using anything beyond md5 any time soon. I would love to see all the checksum algorithms become CPU instruction sets.
Optimizations:
no SIMD-roll, no asm-roll, no openssl-crypto, asm-MD5
Checksum list:
md5 md4 none
Compress list:
zstd lz4 zlibx zlib none
Daemon auth list:
md5 md4The SHA-1 collision attack can only work if you take a specially-crafted file from the attacker and commit it to your repository. The file needs to have a specific structure, and will contain binary data that looks like junk. It can't look like innocent source code. If you execute unintelligible binary blobs from strangers, you're in trouble anyway.
There is no preimage weakness in SHA-1, so nobody is able to change or inject new data to an arbitrary repo/commit that doesn't already contain their colliding file.
Looks like it'll limp along for a while yet
Sounds like the deprecation schedule is too slow and unsafe.
The big issues with SHA2 are MD structure/length extension (which HMAC addresses, and you can also use the truncated versions; length extension matters pretty much exclusively if you're designing entire new protocols) and speed.
I'd reach for Blake2 right now instead of SHA2 (or SHA3) in a new design, but I wouldn't waste time replacing SHA2 in anything that already exists, or put a lot of effort into adding a dependency to a system that already had SHA2 just to get Blake2.
Or is this claim ignoring progress of quantum computing?
"neither us nor our children will see a SHA-256 collision (let alone a SHA-512 collision)" -- JP Aumasson https://twitter.com/veorq/status/652100309599264768
I have somewhat jovially suggested that the OpenPGP standard should just rename it if it turns out that the name becomes a problem...
In any case, since there are feature flags and algorithm preferences to signal support for these, all of this is in fact backwards compatible. There's little risk that someone will accidentally use an AEAD mode to encrypt a message for a recipient that doesn't support it, since the recipient needs to signal support for it in their public key.
And, offering performance and security benefits for those that care to upgrade their implementations and keys is still a worthy goal, IMHO.
That just makes things worse. The OpenPGP standard covers the offline, static encryption case. You have an encrypted file or message. If your implementation has implemented the encryption method then you can decrypt it. If it doesn't than you can't. Contrast this with an online, dynamic method like TLS where you can negotiate a method at the time of encryption and decryption.
>In any case, since there are feature flags and algorithm preferences to signal support for these, all of this is in fact backwards compatible.
The OpenPGP preferences are not applicable to symmetric encryption. In that case the implementation has to guess what method will be supported by the decrypting implementation. In most cases it would make sense to use the most widely implemented method. That is always going to be the one that has been used since forever.
The OpenPGP preferences are included in the public key at key generation time. They reflect the capabilities of that implementation. The public key then has a life of it's own. It is quite normal to use that key to encrypt material for other implementations. Having optional methods, again, makes things much worse here. This is not a theoretical problem. I have already been involved in a difficult usability issue caused by an implementation producing one of these new incompatible encryption modes. The file was created on the same implementation as the public key. So the implementation saw that the reciepient supported all the same methods that it did. But the actual implementation that did the decryption did not support that mode. This sort of interoperability problem, even if it only happens from time to time, seriously impacts usability. That is the ultimate issue. Why are we making things harder for the users of these systems for no real reason?
I agree, if you don't know the capabilities of the implementation decrypting the message then it makes sense to be conservative, and wait until AEAD is widely implemented before using it.
> I have already been involved in a difficult usability issue caused by an implementation producing one of these new incompatible encryption modes.
I also think that generating AEAD encrypted OpenPGP messages today is irresponsible, since the crypto refresh is still a draft, not an RFC yet. Even if the decrypting implementation can read the message today, the draft could still change (although now that it's in Working Group Last Call it's unlikely), and then you'd have an even bigger problem.
But I think that's the fault of the implementation, not of the (proposed) standard. If we eventually want to have the security and performance benefits of AEAD, we have to specify it today (well, or yesterday, but that's a bit hard to change now ^.^).
The thing is, there doesn't seem to be any security weaknesses with the existing AE method. I have looked hard. There also doesn't seem to be any need for the AD (associated data) part.
OCB seems to be the fastest out of all of them. If the proposal was just to add on OCB as an enhanced performance mode then I might be OK with that. Why make the people encrypting multi-TB files wait? I am mostly grumpy with the idea that we have to drop an existing well established standard for stuff like messaging and less extreme file sizes.
Even if nobody found any concrete security issues, there might be compliance reasons not to use SHA-1, as indicated by the OP. It's easier to switch to something new than to endlessly explain to auditors that actually in this particular case the use of SHA-1 is probably fine. Also note that there's no security proof for the MDC, making it harder to convince such auditors.
> There also doesn't seem to be any need for the AD (associated data) part.
I personally disagree: https://gitlab.com/openpgp-wg/rfc4880bis/-/issues/145. The paper linked there explains why having AD would be useful. But I'll grant you that the crypto refresh doesn't make significant use of it yet, so it can't really be counted as an advantage for AEAD yet.
> OCB seems to be the fastest out of all of them. If the proposal was just to add on OCB as an enhanced performance mode then I might be OK with that. Why make the people encrypting multi-TB files wait? I am mostly grumpy with the idea that we have to drop an existing well established standard for stuff like messaging and less extreme file sizes.
I'm not sure I understand what the concrete difference would be between what you're proposing and what the crypto refresh does? It introduces a new mode and encourages its use, but doesn't disallow the use of the current mode. That being said, once OCB is widely deployed, why would you want to use CFB instead of OCB?
I am not sure what aspect of the MDC you would want to prove. The security properties of such constructions are well understood at this point[1]. It is a simple scheme. It has been under scrutiny for 20+ years. It isn't really possible to prove a scheme secure in general. What happens in practice is that it turns out that an assumption is incorrect. The MDC requires few assumptions and is probably more secure than other more complex schemes.
>The paper linked there explains why having AD would be useful.
I don't find that very compelling. The scheme described talks about a false positive rate. Such a thing would not be acceptable in normal PGP usage. I am also not convinced that such a scheme is impossible without AD.
As already mentioned, I am concerned about the ethics of changing a long term standard for no good reason. The users do not deserve the wave of low level interoperability problems such a proposal would create. Usability is a serious issue for end to end encrypted messaging of all types and should be prioritized.
[1] https://link.springer.com/content/pdf/10.1007/3-540-44987-6_...
It does not use the popular combination of an encryption function acting more or less independently of a MAC (message authentication code). It uses a different method[2]. This seems to cause much confusion.
Now the name will be annoyingly and misleadingly wrong forever, in a way that was totally predictable.
I would however be fine with removing "unlimited" from the name of services which are limited, 'Simple' from protocols which don't need subjective comments about their complexity, 'Ultimate' and 'Supreme' from the names of things which are neither the last nor best, etc.
If we do the same thing as the hash names, I think it's fine. "SHA-1", specifically with the "1", doesn't cause problems.
Or AAC - Advanced Audio Codec. Fast forward a few years, Opus came out and you can reduce the bitrate by another 30-50% with similar performance
Standards evolve, and SHA has grown new generations to replace ones that have become insecure. See also "Secure Sockets Layer."