Go Find Duplicates: A fast and simple tool to find duplicate files
github.com
github.com
But the main problem is not the suspicious performance, it's the lack of explanation. The tool is supposed to "find duplicate files (photos, videos, music, documents)". Does it mean it is restricted to some file types? Does it find identical photos with different metadata to be duplicates? Compare this with rdfind which clearly describes what it does, provides a summary of its algorithm, and even mentions alternatives.
Overall, it may be a fine toy/hobby project (3 commits only, 3 months ago), I didn't read the code (except for finding the command-line options). I don't get why it got so much attention.
It initially groups files that have the "same extension and same size", so you're out of luck if you have two copies named foo.jpg and foo.jpeg.
Then, it cheats by computing a crc32 (!) of the beginning, middle, and end bytes of the file and groups together files that have the same crc32.
So, it'll mostly work, but miss a lot of duplicates, and potentially flag different files as duplicates.
The efficient way to do things is to just read files in parallel and break once they diverge. Basically how `cmp` works.
Example. For fileX.txt being X=1 to 10000 you have three copies of each archive in:
mydir/file-X.txt
mydir/subdir/file-X.txt and
mydir/subdir/copy/file-X.txt.
Fdupes would delete random files in mydir/, mydir/subdir/ and mydir/subdir/copy/. You would end with the remaining files scattered by all the directory tree. A mess with three incomplete copies.
Rdfind correctly guess that what most people would want is to remove entirely all files in two of the directories and keep one copy (files and tree-dir) intact. So it wipes the inner subdirs in a predictable way and keeps the outer dir intact. This is a terrific feature able to disentangle one tree directory cloned and nested into the original copy without distroying it, like in this case:
a/b/c/d/files00X.txt
a/b/a/b/c/d/files00X.txt
Computing incremental checksums in parallel and breaking up once checksums diverge does not work very well on HDDs because the files can be in distant physical locations, and that would cause lot of seeks (seeks are terribly slow on rotational drives).
That's going to be pretty slow on a hard disk because it scatters the reads causing a lot of head movement.
If all equally-sized files on your disk are duplicates, it makes you read both files in their entirety in a scatter version. That’s (expected to be) slow, indeed.
However, if most of them aren’t, chances are you find a difference in the first block read, which means you have to read only two blocks to decide they’re different.
So, what’s that distribution? Also, where is one most likely to find differences in equally-sized files? First block? Last?
But it's a good example why there can't be one optimal strategy for finding duplicates. Depending on the data set composition, various implementations turn out the fastest. It depends heavily on where most the files differ (or not).
But just to read the first block, you need to open both files and move the heads from one to another at least once. If the files are placed in distant cylinders, this will be a significant amount of time (a few ms).
At the beginning, fclones also processed files in each group matching by size together. But files matching by size do not necessarily live physically close to each other. Later fclones got a great speedup on HDD when I switched to processing the files in their physical order, regardless of their logical grouping. It is a bit more complex to do in terms of programming because in each grouping pass you have to "ungroup" the files to a flat list, sort the list by locations and group again, but it is really worth it (on rotational drives). You can read about this here: https://pkolaczk.github.io/disk-access-ordering/
fdupes is the classic (going way way back) but it's really very slow, not worth using anymore.
The four I know are worth trying these days (depending on data set, hardware, file arrangement and other factors, any one of these might be fastest for a specific use case) are:
https://github.com/jbruchon/jdupes
https://github.com/pauldreik/rdfind
https://github.com/jvirkki/dupd
https://github.com/sahib/rmlint
Had not encountered fclones before, will give it a try.
The cynical side of me wants to know what features and safety checks a "blazingly fast" tool has not implemented that the older "glacially slow" tool it is replacing ended up implementing after all the edge conditions were uncovered.
find * -type f -exec md5sum '{}' ';' \
| tee /tmp/index_file.txt \
| gawk '{print $1}' \
| sort | uniq -c \
| gawk '/^ *1 /{ print $2 } \
> /tmp/duplicates.txt
for m in $( cat /tmp/duplicates.txt )
do
grep $m /tmp/index_file.txt
echo ========
done \
| less
Tweak as necessary. I do have a comparison executable that only compares sizes and sub-portions to save time, but I generally find it's not worth it.It takes less time to type this that than it does to remember what some random other tool is called, or how to use it. I also have saved a variant that identifies similar files, and another that identifies directory structures with lots of shared files, but those are (understandably) more complex (and fragile).
I'll check on it/them.
Anyway, thanks for sharing - it is always very exciting to see how far you can go with a few unix utilities and a bit of scripting :)
find ... \! -empty ...
they have the same hash, but they do not need to be treated as duplicate
It finds duplicate files and replaces them with hard links, saving you space. Just make sure you provide it with paths in the same filesystem.
I originally wrote it to save some space from personal files (videos, photos, etc), but it turned out very useful for tar files, docker images, websites, and more. For example I maintain a tar file and a docker image with Kafka connectors which share many jar files. Using duphard I can save hundreds of megabytes, or even more than a gigabyte! For a documentation website with many copies of the same image (let's just say some static generators favor this practice for maintaining multiple versions), I can reduce the website size by 60%+, which then makes ssh copies, docker pulls, etc way faster speeding up deployment times.
I say unfortunately because before writing duphard, I tested fdupes and a couple other utilities (duff and duperemove) but none offered the functionality I needed.
-L --linkhard hard link all duplicate files without promptingUse fclones, fslint, jdupes, rdfind instead, which either use much stronger hashes (128-bit) or even verify files by direct byte-to-byte comparison.
E.g. see this:
* https://www.strchr.com/hash_functions
* https://jpountz.github.io/lz4-java/1.2.0/xxhash-benchmark/
Now to address a few concerns:
# The tool doesn't delete anything -- As the name suggests, it just finds duplicates. Check it out.
# File uniqueness is determined by file extension + file size + CRC32 of first 4KiB, middle 2KiB and last 2KiB
# Above seems not much. But, on my portable hard drive with >172K files (mix of video, audio, pics and source code), I got the same number of collisions as that of "SHA-256 of entire file" (By the way, I'm planning to add an option in the tool to do this)
How does it handle small files?
* being able to detect not only duplicate files but also duplicate dirs (without returning all their sub-contents as duplicates)
* being able to query multiple times without having to re-scan, and to do other types of queries (i.e. I am computing a hash on all files, not only of those with duplicate sizes. This makes scanning slower but enables other use-cases. N.b. beware that I only hash fixed portions of files for files>3MB, which is enough for my use-case considering that I always triple-check the results and is a reasonable tradeoff for performance, but it might not be OK for everyone !)
* being able to tell whether all files in dir1/ are included in dir2/ (regardless of file/dir structure)
* being able to mount the sqlite index as a FUSE FS (which is convenient for e.g. diff -r or qdirstat...)
Still work-in-progress, yet it works for several of my use-cases
The first one i found and still use when it got obvious that fslint is EOL is czkawka [0] (meaning hiccup in polish). Its' speed is an order of magnitude higher than fslint, memory use is 20%-75%.
<;)> Satisfied customer, would buy it again. </;)>
I have had pretty good luck with that one. I used to use 'duplicate commander' but I am not sure that one is out there anymore.
Example:
/some/location/one/January/Photos
/some/location/two/January/Photos
I need a tool that would return a match on January directory.
It would be great to be able to filter things. So for example, if I have backups of my dev folder, I want to filter out all the virtual envs (venv below): /home/HumblyTossed/dev/venv/bin /home/HumblyTossed/backups/dev/venv/bin
https://man7.org/linux/man-pages/man1/find.1.html (-mindepth -maxdepth can also be added to make it stricter)
find some/location -type d -wholename '*/January/Photos'
https://github.com/sharkdp/fd fd -p '*/January/Photos' some/locationThe better approach might be, for same size files, to just Seek(FileSize div 2) and read 32 bytes from there. If those are identical with another file then start a full file comparison until one character diverges then stop. If multiple files are having these same middle bytes then maybe do, for each file, a full SHA256 and compare those.
Also, as other commenters pointed, you might have same info but meta is different (videos, pictures, etc) so that needs to be implemented as well.
I did the obvious trick of binning by size before trying to compute any hashes, and was mildly surprised to find how few out of my ~million files had exactly the same size.
For multiple files with identical size I just did the full file MD5, we only had HDD's back then and we all know how much they like random access.
The author says that they tested it on 172K+ files and it's safe, but I still wouldn't trust it enough to delete files from my filesystem.