[1] https://pibytes.wordpress.com/2013/02/09/deduplication-inter...
[1] https://pibytes.wordpress.com/2013/02/09/deduplication-inter...
So as long as your data on disk hasn't evolved intelligence and actively tries to create two blocks with the same hash, your storage system is absolutely fine.
Source: I wrote that hash-based dedup code
Can you tell why SHA256 was chosen? Was it because of the larger hash, or was it a requirement to have a cryptographic secure hash?
Do you know how the hash is used in the Netapp in regard to the remark MichaelMoser123 made: "(two blocks are considered equal if the hash is equal - most systems don't care to check for collisions" [...]
Is it (overoversimplifed ;):
if sha256(a) == sha256(b) {
performDedup()
[...]
Or is a compare added if the hashes found to be identical: if ( sha256(a) == sha256(b) and
byteCompare(a) == byteCompare(b)
) {
performDedup()
[...]If you know a technique for making SHA-1 collisions and a technique for making MD5 collisions, it's probably possible to combine the techniques and make a collision for SHA-1-MD5.
If you are interested in hash function combiners, have a look at Robust Multi-Property Combiners for Hash Functions [2], a recent paper on the topic. It aims at getting several properties at the same time (preimage resistance, collision resistance). For simpler schemes with a single property, just explore the bibliography at the end of the paper.