Might be useful to clarify for people replying with suggested use cases: this repo has two hash functions, describing them as “General-purpose, non-crypto hash” (OK) and “Experimental cryptographic-like hashing” (???)
The question is, what is that second one for? Do they just mean it could eventually become a cryptographic hash?
edit: AH, I see you mean in the comparison table. Yeah that wording could be improved. I'll change it now. The rest of the text is hopefully clearer:
"Consider it (rainstorm) experimental, but designed to be secure. Note on Cryptographic Intent: Rainstorm's design includes concepts from cryptographic hashing, but without formal analysis, it should not be considered secure. Feedback and analysis are welcome."
Although without a bunch of accepted analyses people shouldn't trust a secure cryptohash that merely claims to be a secure cryptohash - which this doesn't claim.
It is funny people get obsessed with wording when what's really needed is evidence.
I will take a look at the wording again now and if there's any wording that's insufficiently clear that rainstorm is proposed as a crypto-hash and requires analysis, I'll update such wording.
For clarity, this lib has 2 hashes: rainbow, and rainstorm. "rainbow" the faster, is not intended to be crypto secure. "rainstorm" the slower, designed with strength, is, but it has not been analyzed publicly.
My intent is that rainstorm is proposed as a cryptohash and eventually is analyzed properly in order to be trusted, and used - or improved! We probably can't remove the "ish" until then.
In terms of saying "more secure"? Without formal cryptanalysis it is hard to say, but if you have good statistical properties (as measured by these tests in SMHasher3) it's fair to say you're more secure as a lot of the attacks on (cryptographic) hashes are to do with finding dependencies between input and output ("differentials").
I wouldn't make sweeping recommendations unless a hash was vetted because that's how we end up trusting primitives, through the trust in the collective work of people trying to attack them. Right now we just have the speed results, the SMHasher3 results and the design, and some limited analysis. But more is coming, some people are doing good stuff on GitHub right now.
In terms of aptness, to use either of my hashes in rain as a cryptohash without having a deep understanding of potential attacks against it would be rolling the dice, but you could probably trust that at least potential attacks would seem to be difficult given the lack of statistical predictabilities. And they're both fine to use as regular non-cryptographic hashes.
Of course I want to recommend people use them, but I couldn't say use them in crypto applications without being irresponsible and I don't want to do that. Also, to be more lawyerly about it, heh, it's provided "as is".
I guess the main draw right now is if you want good statistical properties and hashes that can be 64, 128 or 256 (or even 512 in rainstorm) bits wide, and fairly fast for that.
Pretty standard, and consistent with the rest of the text.
What did you think it meant?
I'm not trying to pick on you, it's just that this keeps coming up with constructions like PCG. There's a cryptographic literature, and then a side-hustle of people building stuff like this. We used to have a sci.crypt discourse about it, but now I worry that HN has lost the antibodies.
Maybe your immune response was wrongly triggered? It happens. But what we really need antibodies against is excessive negativity and vague criticism, combined with attempts to justify the same with "you're not someone who is supposed to do that" gatekeeping.
Could some of this be motivated by the fact that you want to do stuff like this yourself, but instead of doing so, or analyzing it, you just criticize it, inappropriately? I don't think you're trying to pick on me, just confused and powered by righteous protective instinct that is better were it more informed and discerning. Reflect some?
What other ideas can you think of? How about in a construct for password hashing? You can add more rounds to the loop to increase the cost like pbkdf1/2.
For rainbow you can use it in place of whatever other non-crypto hash you have, especially if you want fast 128 or 256 bit widths.
If you're trying to be friendly you got a funny way about it. What’s your philosophical question?
If not...Oh well, you might have to wait until it gets analyzed a lot. Use Rainbow now without fear: fast, statistically great, and unlikely collisions - which makes most uses better. As someone else put it: "fastest 128-bit and 256-bit non-crypto hash, passes all tests, and under 140 source lines of code."
I'm sure they will eventually, not too long. What's your big concern / lack of confidence here?
In the meantime you can take care of it and nurture it by encouraging people to analyze it, using your influential voice for good.
Hopefulyl this gives an overview of some of the historical influences and inspirations.
For a checksum, this isnt a bad thing and actually could be good.
There are applications where hashes are used as identifiers, where it would be impossible to use the original data to resolve possible collisions. One example would be in RTTI (Run Time Type Information), when you want to check if two objects (of unknown-at-compile-time type) are instances of the same type. The only performant way to achieve this is if the compiler emits each type’s RTTI with a unique integer key to efficiently compare against, and this key must be unique across all object files and shared objects. In practice, this is done using hashing. If there is a collision, then the program behavior is undefined, so it is ideal to minimize probability of collision. You can see this issue in the Rust compiler project where there is ongoing discussion about how to handle this hashing, trying to balance compiler speed and collision likeliness:
https://github.com/rust-lang/rust/issues/129014
Another scenario would be like what Backtrace does: a hash is used to identify certain classes of errors in application code. The cost of analyzing each error is quite high (you have to load objects into memory, source maps (if available), perform symbolification, etc), so comparing any two errors every time you want to perform any metrics or retrieve associated data (like issue tracker identifiers) would reduce the performance to a point that any such application would be too slow to be useful. So a lot of thought was given to producing fast, non-cryptographic hashes (these hashes aren’t shared across tenants, and no customer is going to go out of their way to work out initialization parameters so they can be their own adversary) with a known likely hood of collision. You can read about the hash (UMASH) here:
https://engineering.backtrace.io/2020-08-24-umash-fast-enoug...
FWIW I think commonly used checksums are basically ideal ("equally likely" to come to a different value) on single-bit flipped inputs, but don't quote me on that.
I read that as you thinking Git is a motivating use case for this hash.