Ask HN: Best tool for de-duplication of local files?
There is quite a bit of overlap and many instances exist of the same file having been backed up in different places.
As it stands I’ve collated the dumps into local folders on a fairly fast NVME drive and I’m trying to think of the best way to merge 3-4 local folder trees representing the cloud dumps into a single output folder that will skip binary duplicates by hash.
Ideally I’d also be able to detect duplicates where some files have been compressed (so will not match on hash), but I have a feeling that’s a really optimistic goal for any kind of automation.
Can anyone help with a suggestion?