How I created two images with the same MD5 hash
natmchugh.blogspot.com
natmchugh.blogspot.com
> The chosen prefix collision attack works by repeatedly adding 'near collision' blocks which gradually work to eliminate the differences in the internal MD5 state until they are the same
> This type of collision is has been termed a chosen prefix collision. In this case the image data is the prefix or to be more exact the internal state of the MD5 algorithm after processing the image is. You can't see the added binary data at the end of jpeg images as it is preceded with an End Of Image JPEG marker.
From extensive experimentation with this exact topic, you have to be really careful because there's a lot of handwavey arguments or "just use openssl" but just because md5 is 33% faster than sha1 on 32 bit openssl in i386 doesn't mean whatever random platform and library combo that you're actually using will be exactly 33% by switching from sha1 to md5. Probably, someone out there has a pathological example where md5 is somehow actually slower than sha1.
MD5 is almost too good of a hash function for mere de-duping, but I have to use it because its near universal. On mysql I have the function "md5()". In clojure its "(digest/md5 "something")". On the AS/400 thankfully thats not my area of responsibility but I know its compatible with all of the above. In perl its "use Digest::MD5". The most important part is a md5 calculated on an engineering data record will always be the same hash result number no matter the source, which is very important for de-duping.
The world of programmers has some very peculiar ideas about hashes such that there is usually only one library for hashing and it often only has stuff like md5 or sha1 which is not good enough for crypto and overkill for de-duping and ID creation. The world would be a better place if as a profession we split our hashing libraries on all machines and languages into "really fast insecure de-dupe hash functions" and "really secure glacially slow crypto hash functions"
I just wanted to point out that, for situations where user input and security are important, you want the algorithm to be slow.
I didn't say anything about how to implement it or whether you should use BLAKE2 or what. There's a lot more to it than I could put in a reply here, and even quick Googling would turn up info about salting/iterating/etc.
https://leastauthority.com/blog/BLAKE2-harder-better-faster-...
And that's how my colleague sees the photos of some unknown woman in his Dropbox (and supposedly, she sees his pictures too).
Furthermore, the kind of weakness md5 has now would require that that the attacker generate both usernames. And the usernames would be really really long because they had gibberish tacked onto the end to make the hashes collide. In the jpeg format you can make that part invisible, but not in a username.
MobileSafari, iOS8.
Or if not, maybe more people need to be told not to use blogspot. Second article I've read in a week that was broken by this "improvement".
Conflicting git check ins, breaking cache layers, tampering with downloads, etc
2. git stores the length of a blob before hashing [1] which makes it harder (but not impossible) to perform such attacks [2]
[1] http://www.git-scm.com/book/en/v2/Git-Internals-Git-Objects#... [2] http://blogs.msdn.com/b/oldnewthing/archive/2004/05/19/13493...
I'm aware that they are possible, but they appear to be unlikely, especially when using md5 correctly, even then most people have already migrated to sha1 or sha256...
Do both files need to be altered to find a collision, or just one? Can it be done in a fixed file size?
There was fear that the attacks could be extended to SHA-2, thus we now have SHA-3 too. However, SHA-2 remains secure for now.
_____
¹ Wikipedia: »As of 2012, the most efficient attack against SHA-1 is considered to be the one by Marc Stevens[34] with an estimated cost of $2.77M to break a single hash value by renting CPU power from cloud servers.« I.e. it's quite expensive, but can be done in a reasonable time, especially by adversaries with interest and funds to do so.
SHA-1 and SHA-2 are similar at an architectural level, in some of the same ways that two mid-1990s Feistel ciphers might be similar, and share building blocks, but they aren't the same hash function. They are much more different than, say, DES and 3DES.
SHA-2 remains the best practical choice for most systems today. The truncated variants (like SHA2-512/256) even break length extension exploits.
That knowledge has driven adoption of SHA-2, and was partial motivation for the (completed in 2012) competition to design SHA-3. I don't believe any weaker-than-designed problems have yet been found with the SHA-2 family.
In 2012, Bruce Schneier reported on an analysis by Jesse Walker of Intel about when SHA-1 collisions might be practical to create:
https://www.schneier.com/blog/archives/2012/10/when_will_we_...
That analysis suggests: "A collision attack is therefore well within the range of what an organized crime syndicate can practically budget by 2018, and a university research project by 2021."
But it also notes non-commodity approaches (GPUs, custom chips, etc) could achieve SHA-1 collisions sooner/cheaper.
[1] http://en.wikipedia.org/wiki/Pigeonhole_principle
Edit: source
We're getting there with reduced SHA-1[1] (that is, less than 80 rounds, that means less than 2^80 theoretical operations[2]). But the cost of finding a collision decreases over time[3], and this is why everybody says SHA-1 is obsolete.
[1] http://eprint.iacr.org/2010/413.pdf [2] https://www.schneier.com/blog/archives/2005/02/sha1_broken.h... [3] https://www.schneier.com/blog/archives/2012/10/when_will_we_...
"To search though all possible MD5 values is 2^128 operations which is massive. To be in with a good chance of finding a collision would take ~ 2^64 operations which is again far too big for normal computing."
Seams weird, why so low to find collision compared to full search? or does the author not understand Exponentiation?[1] http://en.wikipedia.org/wiki/Birthday_problem
[2] http://en.wikipedia.org/wiki/Collision_attack#Classical_coll...