Should you remove duplicate files?
eclecticlight.co
eclecticlight.co
What are Mac users doing that makes finding and removing duplicate files such a regular task? I don't use a Mac, but I can honestly say I've never given second thought to duplicate files. I'm sure there are some here and there but my hard drive isn't filling up with duplicates of large files.
So you end up with a lot of duplicates from either trying to edit a read only document in some apps (which spawns an editable copy) or accidentally hitting duplicate (generate a copy of the current file in the same directory) which is the mac equivalent of “Save As”. It’s on you the user to move the new file where you want it.
Here are just a few ways I’ve ended up with enough dupes to want to dedupe, which has happened on Windows and Linux for me as frequently as on Mac: Backing up mobile phones, using version control for multiple repositories that share files, losing access to repository servers and archiving or backup up the files, forking an old archived folder and making modifications to it that I want to keep, not to mention homegrown schemes & scripts for archiving and backing up when I either don’t want to pay for backup software or can’t even find anything appropriate to all my needs (cross platform, saving to a NAS, shared with my family, etc.)
* 7:00 AM find and remove duplicate files
* 8:00 AM reorganize the 2D layout of file icons within folders for maximum visual pleasure
* 9:00 degauss the LCD (really frustrated that you need third party software for this since the CRT days)
* 10:00 find and remove duplicate files
And so on. Finding duplicate files is just as essential as monitor degaussing, but you won't find this listed in the manual. You have to go to user forms where people will convince you of the need and get you to adopt procedures like the above to keep your machine at top performance.
Learned a ton about distros, backups, treating machines as disposable, configuration management and scripting the desktop setup though. It wasn't completely wasted.
https://packages.gentoo.org/packages/search?q=Degauss
Until gentoo has this, I'm not sure if it's ready for the desktop. Maybe next year?
Only if the destination is on a different volume. Like from a hard disk to a USB drive.
Hold down the Command key while dragging a file between volumes if you really want Move instead of Copy.
And on the other side: Apple computers typically don’t have upgradeable hard drives so for anyone who takes photos/video it’s an inevitable countdown to zero free space.
From the project:
> This name is not as bad as it seems, because I was also thinking about using words like żółć, gżegżółka or żołądź, but I gave up on these ideas because they contained Polish characters, *which would cause difficulty in searching for the project.*
Really dodged a bullet there.
They work shockingly well on large libraries, I've managed to check millions of files consuming several TBs in a matter of hours.
I've been a sysadmin managing SunOS and Linux systems at scale.
I've had personal machines running a variety of OSs: FreeBSD, Linux, Windows, Mac OS, OS X, to say nothing off all the different filesystems involved across those OSes.
In all that time, I can't think of a single time I've ever bothered to de-duplicate files.
Once upon a time, you used to occasionally need to defragment some file systems. But it's been more than two decades since I've even thought about defragmenting. Every modern filesystem just handles that automatically.
What even is the use case where folks end up with enough duplicate files that it's worth worrying about?
In college there was, for a glorious, brief, time, a program called mytunes. It would scan the network for iTunes instances that had the network sharing turned on. Then it allowed you to copy over entire iTunes libraries (iTunes only let you play songs from other users).
I would log on to the dorm network which normally had dozens of iTunes instances running at a go, and just copy everything to my, at the time, massive 160gb drive.
Then you would dedupe. That’s the last time I thought about deduping
It turns out the node_modules folders I've collected over the years had several gigabytes of duplication. The same was true to a lesser extend for Python code (because the dependency folders and venvs were a lot smaller).
Non-technical person will copy-paste a folder of documents to make a backup before editing a file. A crude form of version control. I've seen this method lead to the consumption of hundreds of gigabytes for duplicate content, because the folders contain a lot of auxiliary data like PDFs and JPEGs.
Technical person will run npm install and get a massive node_modules directory that is almost identical to the node_modules directory in another project.
Do they? I always assumed they just try to avoid fragments but ignore them once created, and rely on bigger faster disks to hide any problems this may cause.
What's unique about PhotoSweeper is that it gives you deep control over what you mean by "dupe" and supports pro image formats like RAW and DNG. For example, I was able to automatically de-dupe byte-for-byte equivalents, then manually evaluate dupes with the same image data but slightly different metadata. Of course, it can do things like find dupes at different resolutions and keep the highest-quality version, etc. It saved me a ton of time.
If your fs implements hard links, no need inside one mount point. But you can't hard link cross mount.
Copies can be good or bad. It depends.
DragonFlyBSD's HAMMER filesystem supports on-demand/scheduled dedup using the hammer(8) command, with configurable memory and runtime limits.
So if I have a disk image that takes 1,000 blocks, and clone it - I now have two images that total 1,000 blocks. And if I modify one block in one image, I now have two different images that total 1,001 blocks. With a hardlink the modification would affect both images.
The way I visualise the difference, is that a directory entry points to a file entry, and a file entry points to a list of blocks. A hardlink is a new directory entry that points to an existing file entry. a linked clone is a new file entry that points to the same list of blocks.
Btrfs (and more recently also XFS) supports on-demand, batch dedup which doesn't have any of the downsides of ZFS's always-active dedup. This is an implementation issue, not a categorical problem of block (or extent) dedup.
https://btrfs.readthedocs.io/en/latest/btrfs-man5.html#raid5...
Also, reading stuff like "mostly OK" in a file system doesn't inspire confidence.
If Time Machine is tracking file contents for deduplication it would be really cool if it did this to the source files whenever it noticed that they were identical.
While there are many people who weirdly (to me at least) grotesquely overestimate the value of deduplication, it is legitimately useful because we tend not to want to completely rewrite large files very often like that, so even modifications to that file are very likely to either be appends or modifications of something in the middle without shifting the entire file.
Is it intentional?