Hyperspace
hypercritical.co
hypercritical.co
I see it is "pre-release" and sort of low GH stars (== usage?), so I'm curious about the stability since this type of tool is relatively scary if buggy.
Whenever using it on something sensitive that I can’t back up first for whatever reason, I make checksum files and compare them afterwards. I’ve done this many times on hundreds of GB and haven’t seen corruption. Caveat emptor.
There is one huge caveat I should add to the README - block corruption happens. Having a second copy of a file is a crude form of backup. Cloning causes all instances to use the same block, so if that one instance is corrupted, all clones are. That’s fine for software projects with generated files that can be rebuilt or checked out again, but introduces some risk for files that may not otherwise be replaceable. I keep multiple backups of all that stuff in hardware other than where I’m deduping, so I dedup with abandon.
I’m a nobody with no audience. Maybe some attention here will get some users.
I was also really impressed that `make` ran basically instantly.
I love the documentation from FreeBSD and OpenBSD. Only having target one platform and only system libraries makes building simple.
A few notes:
* By default it doesn't scan everything. It ignores all files but those in an allow list. The way the allow list is structured, it seems like Hyperspace needs to understand the content of a file. As an end user, I have no idea what the difference between a Text file and a Source Code file would be or how Hyperspace would know. Hyperscan only found 360MB to dedup. Allowing all files increased that to 842MB.
* It doesn't scan files smaller than 100 KB by default. Disabling the size limit along with allowing all files increased that to 1.1GB
* With all files and no size limit it scanned 67,309 of 68,874 files. `dedup` scans 67,426.
* It says 29,522 files are eligible. Eligible means they can be deduped. `dedup` only fines 29,447. There are 76 already deduped files, which is an off-by-one, so I'm not sure what the difference is.
* Scanning files in Hyperspace took around 50s vs `dedup` at 14s
* It seems to scan the file system, then do a duplication calculation, then do the deduplication. I'm not sure why the first shouldn't be done together. I chose to queue any filesystem metadata as it was scanned and in parallel start calculating duplicates. The vast majority of the time files can be mismatched by size, which is available from `fts_read` "for free" while traversing the directory.
* Hyperspace found 1.1GB to save, `dedup` finds 1.04GB and 882MB already saved (from previous deduping)
* I'm not going to buy Hyperspace at this time, so I don't know how long it takes to dedup or if it preserves metadata or deals with strange files. `dedup` took 31s to scan and deduplicate.
* After deduping with `dedup`, Hyperscan thinks there are still 2 files that can be deduped.
* Hyperspace seems to understand it can't dedup files with multiple hard links, empty files, and some of the other things `dedup` also checks for.
* I can't test ACLs or any other attribute preservation like that without paying. `strings` suggests those are handled. HFS Compression is a tricky edge case, but I haven't tested how Hyperspace's scan deals with those.
With a FOSS project this would have been expected, but with a ShareWare-style model? Idk..
Again, I've got no problem with people selling software or closed source models, but I've never understood using this justification. Maybe in this instance he's a well known public figure with published contact info that people will abuse?
One does not simply “not accept bug reports”.
https://github.com/sqlite/sqlite/blob/e8346d0a889c89ec8a78e6...
Who says you have to deal with support requests if you open source something?
> All his apps are personal itches he scratched and he sells them not to make a profit but to make the barrier of entry high enough to make user feedback manageable.
That makes no sense
Almost anyone who has ever maintained popular open-source software, even if dealing with them means putting up a notice that says "Don't ask support questions" and having to delete angrily posted issues.
My understanding from listening to his explanation is he wants to be able to support users and have an income stream to incentivize that.
As an open-source maintainer of a popular piece of software, I'm very empathetic.
I mean, that's very obviously a false statement. You don't have to post any notices or reply to or delete any issues.
> My understanding from listening to his explanation is he wants to be able to support users and have an income stream to incentivize that.
That's valid, but is basically the opposite of the reasoning provided in gp comment.
So GitHub created a mess, and the whole of open source is considered to be GitHub now.
The solution is the same as being able to avoid tons of Windows-related headaches when you don't use Windows. Just don't use GitHub.
A tar or zip file with source code posted online (or bundled with the program, even) under an open source license is still open source.
There's a lot of merit to Open Source. But there's also a lot of spam, politics and drama that comes with opening up. That negativity is invisible to people who haven't encountered it, or are simply guilty of causing it.
Maintainer burnout is real; more power to John for choosing whatever keeps him focused on building good software.
Just keeping everything closed is really missing the point of how trust in infra that handles critical data is built nowadays.
Apps like this can easily bit rot, and more users does often mean more work e.g. answering or filtering emails, finding more edge cases, etc.
From his perspective that means having a income to dedicate time to this. I don't think he's interested in being an "infra" app as you would think of it.
As someone who maintains critical open-source software, I can strongly empathize, even if it’s not an approach I would take.
Making software proprietary and for-pay, especially such a small tool, doesn't just significantly reduce the number of eyeballs this testing is crowdsourced to, but it also disincentivises issue reporting .. why should I spend the time to for free report sth to somebody who is making money off my testing and doesn't even bother to be transparent about how things work exactly (i.e. the source)?
If you really care about the quality of your work then maximizing the eye ball count and incentivise high qualith issue reporting.
Though if you want to maximize income instead, you keep it closed and ask for a subscription.
Quite obvious which option he chose.
You don’t have to. No one is saying you are compelled to report bugs in software you paid for. Most people don’t. The benefit to you as a customer is it can help get the bug fixed. That is clearly a mutual benefit.
> If you really care about the quality of your work then maximizing the eye ball count and incentivise high qualith issue reporting.
I think you’re vastly overestimating the value in the “higher quality” bug reports you’re getting from free users. You might get some higher quality reports but you’ll mostly get a lot more noise.
There are limits to how practical it is to allow for more and more feedback and that threshold for a solo developer is quite low. Restricting your user base by charging for your work means that there is less noise because the only people sending bug reports are paid users.
The quality of these reports are probably lower than if you had an open issue tracker, but you are substantially reducing the mental overhead and you know the people that are sending feedback are doing so with their own interests in mind.
An issue tracker, on the other hand, requires active engagement from the developer. Every issue, even low quality ones, require some form of processing, be that responding, closing, or categorizing. While tools can assist a person in these tasks, the developer is ultimately still responsible for it.
I'm not saying people should only create closed-sourced paid software, but I strongly disagree with the idea that it's negatively affecting the quality of the software because there's no open issue tracker for people to post to.
It's not just github. It's every single issue tracker where users can submit feedback, some of which are almost entirely opaque, like Apple's feedback system. Look at Mozilla's issue tracker, or look at the mailing lists for linux. It's a lot of effort which simply is not worth it for a lot of people in a lot of cases.
Nope.
Pretty strange thing to say in the context of a closed source app.
My intent was to express sympathy to making a closed source app instead of an open source one.
Something must have got lost in translation because it fits the context exactly.
It's equivalent to saying 'I worked on a closed source project and liked it, therefore open source model sucks.'
App store prices are localized. If the blog post said it costs “$10” or whatever, that doesn’t mean anything to millions of potential customers who live where they don’t use $, and is confusing for millions more that do use $ but don’t know if the price is in their local $ or USD
The local currency argument is wrong btw. I'm located in Europe and use a Spanish IP. The prices shown are in USD.
There are lots of apps called "Hyperspace" in the Apple app store, by the way.
https://apps.apple.com/us/app/hyperspace-lighting/id15371988...
https://apps.apple.com/us/app/hyperspace-gpt-chats-ai-art/id...
...
I ran it over my Postgres development directories that have almost identical files. It saved me about 1.7GB.
The project doesn't have any license associated with it. If you don't mind, can you please license this project with a license of your choice.
As a gesture of thanks, I have attempted to improve the installation step slightly and have created this pull request: https://github.com/ttkb-oss/dedup/pull/6
This got me wondering why the filesystem itself doesn't run a similar kind of deduplication process in the background. Presumably, it is at a level of abstraction where it could safely manage these concerns. What could be the downsides of having this happen automatically within APFS?
I still do not trust de-duplication software.
I've done the entire compare every file via hashing and then log each of the matches for humans to compare, but never has any of that ever been allowed to mv/rm/link -s anything. I feel my imposter syndrome in this regard is not a bad thing.
Any halfway-competent developer can write some code that does a SHA256 hash of all your files and uses the Apple filesystem API's to replace duplicates with shared-clones. I know swift, I could probably do it in an hour or two. Should you trust my bodgy quick script? Heck no.
The author - John Siracusa - has been a professional programmer for decades and is an exceedingly meticulous kind of person. I've been listening to the ATP podcast where they've talked about it, and the app has undergone an absolute ton of testing. Look at the guardrails on the FAQ page https://hypercritical.co/hyperspace/ for an example of some of the extra steps the app takes to keep things safe. Plus you can review all the proposed file changes before you touch anything.
You're not paying for the functionality, but rather the care and safety that goes around it. Personally, I would trust this app over just about any other on the mac.
You are involved. You see the list of duplicates and can review them as carefully as you'd like before hitting the button to write the changes.
I’ve heard many horror stories of dedupe related corruption or restoration woes though, especially after a ransomware attack.
I've been writing a similar thing to dedupe my photo collection and I'm so paranoid of pulling the trigger I just keep writing more tests.
I think that ZFS actually does this. https://www.truenas.com/docs/references/zfsdeduplication/
EDIT: Is this referring to the "fast" dedup feature?
There were about 1000 copies of the same pre-requisite .NET and VC++ runtimes (each build had one) and we only paid for the cost of storing it once. It was great.
It is worth pointing out though, that on Windows Server this deduplication is a background process; When new duplicate files are created, they genuinely are duplicates and take up extra space, but once in a while the background process comes along and "reclaims" them, much like the Hyperspace app here does.
Because of this (the background sweep process is expensive), it doesn't run all the time and you have to tell it which directories to scan.
If you want "real" de-duplication, where a duplicate file will never get written in the first place, then you need something like ZFS
(not really, since it's not fragmentation, but conceptually similar)
ZFS is great if you believe you'll exceed some threshold of space while writing. I don't personally plan my volumes with that in mind but rather make sure I have some amount of excess free space.
WinSvr allows you to disable dedupe if you want (don't know why you would) where as ZFS is a one-way street without exporting the data.
Both have pros and cons. I can live with the WinSvr cons while ZFS cons (memory) would be outside of my budget, or would have been at the particular time with the particular system.
It might also be a little unintuitive that modifying one byte of a large file would result in a lot disk activity, as the file system would need to duplicate the file again.
I believe as soon as you change a single bite you get a complete copy that’s your own.
And that’s how this program works. It finds perfect duplicates and then effectively deletes and replaces them with a copy of the existing file so in the background there’s only one copy of the bits on the disk.
As far as I understand, it works like a reflink feature in the modern linux FSs. If so, thats really cool, and thats also a bit better than the zfs's snapshots. Iam newbie on macos, but it looks amazing
However APFS gives you a number of space related foot-guns if you want. You can overcommit partitions, for example.
It also means if you have 30 GB of files on disk that could take up anywhere from a few hundred K to 30 GB of actual data depending on how many dupes you have.
It’s a crazy world, but it provides some nice features.
I think it stores a delta:
https://en.wikipedia.org/wiki/Apple_File_System?wprov=sfti1#...
If we only have two files, A and its duplicate B with some changes as a diff, this works pretty well. Even if the user deletes A, the OS could just apply the diff to the file on disk, unlink A, and assign B to that file.
But if we have A and two different diffs B1 and B2, then try to delete A, it gets a little murkier. Either you do the above process and recalculate the diff for B2 to make it a diff of B1; or you keep the original A floating around on disk, not linked to any file.
Similarly, if you try to modify A, you'd need to recalculate the diffs for all the duplicates. Alternatively, you could do version tracking and have the duplicate's diffs be on a specific version of A. Then every file would have a chain of diffs stretching back to the original content of the file. Complex but could be useful.
It's certainly an interesting concept but might be more trouble than it's worth.
https://www.vastdata.com/blog/breaking-data-reduction-trade-...
Identifying similar blocks and, maybe sub-rechunking isn’t something I’ve ever considered.
https://www.microsoft.com/en-us/download/details.aspx?id=397...
[0] https://www.truenas.com/docs/references/zfsdeduplication/
Of course, that is true of most filesystems.
ZFS is essentially an object store database at one layer; the checksum-hash deduplication table is an object like any other (file, metadata, bookmarks, ...). There is one deduplication table per pool, shared among all its datasets/volumes.
On reads, one does not have to consult the dedup table.
The mechanism was fairly easy to add. And for highly-deduplicatable data that is streaming-write-once-into-quiescent-pool-and-never-modify-or-delete-what's-written-into-a-deduplicated-dataset-or-volume, it was a reasonable mechanism.
In other applications, the deduplication table would tend to grow and spread out, requiring extra seeks for practically every new write into a deduplicated dataset or volume, even if it's just to increment or decrement the refcount for a record.
Destroying a deduplicated dataset has to decrement all its refcounts (and remove entries from the table where it's the only reference), and if your table cannot all fit in ram, the additional IOPS onto spinning media hurt, often very badly. People experimenting with deduplication and who wanted to back out after running into performance issues for typical workloads sometimes determined it was much MUCH faster to destroy the entire pool and restore from backups, rather than wait for a "zfs destroy" on a set of deduplicated snapshots/datasets/volumes to complete.
Doing deduplication at this level is nice because you can dedupe across file systems. If you have, say, a thousand systems that all have the same OS files you can save vats of storage. Many times, the only differences will be system specific configurations like host keys and hostnames. No single filesystem could recognize this commonality.
This fails when the deduplication causes you to have fewer replicas of files with intense usage. To take the previous example, if you boot all thousand machines at the same time, you will have a prodigious I/O load on the kernel images.
> There is no way for Hyperspace to cooperate with all other applications and macOS itself to coordinate a “safe” time for those files to be replaced, nor is there a way for Hyperspace for forcibly take exclusive control of those files.
Maybe not solving the same problem.
You probably don’t want your phone or watch de-duping stuff.
There are tons of knobs you can tweak in macOS but Apple has always been pretty conservative when it comes to what should be default behavior for the vast majority of their users.
Certainly when you duplicate a file using the Finder or use cp -c at the command line, the Copy-on-Write functionality is being used; most users don’t need to know that.
I wish it were more obvious how to do it with other software. Often there's a learning curve in the way before you can see the value.
After all how many perfect duplicate files do you probably create a month accidentally?
There’s a subscription or buy forever option for people who think that would actually be quite useful to them. But for a ton of people a one time IAP that gives them a limited amount of time to use the program really does make a lot of sense.
And you can always rerun it for free to see if you have enough stuff worth paying for again.
however has anyone been able to find out from the website how much the license actually costs?
For example, I discovered my time machine backup kicked out the oldest versions of files I didn't know it had a record of and thought I'd long since lost, but it destroyed the names of the directories and obfuscated the contents somewhat. Thousands of numerically named directories, some of which have files I may want to hang onto, but don't know whether I already have them or not, or where they are since it's completely unstructured. Likewise, many of them may just have one of the same build system text file I can obvs toss away.
If we took the first 1024 bytes of each file as the lookup key, then our key size would be 1024 bytes. If you have 1 million files on your disk, then that's 128MB of ram just to store all the keys. That's not a big deal these days, but it's also annoying if you have a bunch of files that all start with the same 1024 bytes -- e.g. perhaps all the photoshop documents start with the same header. You'd need a 2-stage comparison, where you first match the key (1024 bytes) and then do a full comparison to see if it really matches.
Far more efficient - and less work - If you just use a SHA256 of the file's contents. That gets you a much smaller 32 byte key, and you don't need to bother with 2-stage comparisons.
I don't think it would be far more efficient to do hash the entire contents though. If you have a million files storing a terabyte of data, the 2 stage comparison would read at most 1GB (1 million * 1KB) of data, and less for smaller files. If you do a comparison of the whole hashed contents, you have to read the entire 1TB. There are a hundred confounding variables, for sure. I don't think you could confidently estimate which would be more efficient without a lot of experimenting.
... okay, so as long as you always feed chunks of data into your hash in the same deterministic order, it doesn't matter for the sake of correctness what that order is or even if you process some bytes multiple times. You could hash the first 1kB, then the second-through-last disk blocks, then the entire first disk block again (double-hashing the first 1kB) and it would still tell you whether two files are identical.
If you're reading from an SSD and seek times don't matter, it's in fact probable that on average a lot of files are going to differ near the start and end (file formats with a header and/or footer) more than in the middle, so maybe a good strategy is to use the first 32k and the last 32k, and then if they're still identical, continue with the middle blocks.
In memory, per-file, you can keep something like
- the length
- h(block[0:4])
- h(block[0:4] | block[-5:])
- h(block[0:4] | block[-5:] | block[4:32])
- h(block[0:4] | block[-5:] | block[4:128])
- ...
- h(block[0:4] | block[-5:] | block[4:])
etc, and only calculate the latter partial hashes when there is a collision between earlier ones. If you have 10M files and none of them have the same length, you don't need to hash anything. If you have 10M files and 9M of them are copies of each other except for a metadata tweak that resides in the last handful of bytes, you don't need to read the entirety of all 10M files, just a few blocks from each.A further refinement would be to have per-file-format hashing strategies... but then hashes wouldn't be comparable between different formats, so if you had 1M pngs, 1M zips, and 1M png-but-also-zip quine files, it gets weird. Probably not worth it to go down this road.
Also, use the length of the file for a fast check.
For each candidate file, you need some "key" that you can use to check if another candidate file is the same. There can be millions of files so the key needs to be small and quick to generate, but at the same time we don't want any false positives.
The obvious answer today is a SHA256 hash of the file's contents; It's very fast, not too large (32 bytes) and the odds of a false positive/collision are low enough that the world will end before you ever encounter one. SHA256 is the de-facto standard for this kind of thing and I'd be very surprised if he'd done anything else.
At that point maybe it’s better to just compare byte by byte? You’ll have to read the whole file to generate the hash and if you just compare the bytes there is no chance of hash collision no matter how small.
Plus if you find a difference in bytes 1290 you can just stop there instead of reading the whole thing to finish the hash.
I don’t think John has said exactly how on ATP (his podcast with Marco and Casey), but knowing him as a longtime listener/reader he’s being very careful. And I think he’s said that on the podcast too.
Wonder what the distribution is here, on average? I know certain file types tend to cluster in specific ranges.
>maybe it’s better to just compare byte by byte? You’ll have to read the whole file to generate the hash
Definitely, for comparing any two files. But, if you're searching for duplicates across the entire disk, then you're theoretically checking each file multiple times, and each file is checked against multiple times. So, hashing them on first pass could conceivably be more efficient.
>if you just compare the bytes there is no chance of hash collision
You could then compare hashes and, only in the exceedingly rare case of a collision, do a byte-by-byte comparison to rule out false positives.
But, if your first optimization (the file size comparison) really does dramatically reduce the search space, then you'd also dramatically cut down on the number of re-comparisons, meaning you may be better off not hashing after all.
You could probably run the file size check, then based on how many comparisons you'll have to do for each matched set, decide whether hashing or byte-by-byte is optimal.
To have a mere one in a billion chance of getting a SHA-256 collision, you'd need to spend 160 million times more energy than the total annual energy production on our planet (and that's assuming our best bitcoin mining efficiency, actual file hashing needs way more energy).
The probability of a collision is so astronomically small, that if your computer ever observed a SHA-256 collision, it would certainly be due to a CPU or RAM failure (bit flips are within range of probabilities that actually happen).
Context is everything.
This is the default for ZFS deduplication and git does something similar with size and far weaker SHA-1. I would add a test for SHA-256 collisions, but no one has seemed to find a working example yet.
…and you’re not worried about shark attacks, are you?
Hashing the whole file after that is wasteful. You need to read (and hash) only as much as needed to demonstrate uniqueness of the file in the set.
The tree concept can be extended to every byte in the file:
https://github.com/kornelski/dupe-krill?tab=readme-ov-file#n...
I have one data set where `dedup` was 40% faster than `dupe-krill` and another where `dupe-drill` was 45% faster than `dedup`.
`dupe-krill` uses blake3, which last I checked, was not hardware accelerated on M series processors. What's interesting is that because of hardware acceleration, `dedup` is mostly CPU-idle, waiting on the hash calculation, while `dupe-krill` is maxing out 3 cores.
Thanks for the link!
https://crypto.stackexchange.com/questions/47809/why-havent-...
- compute SHA256 hashes for each file on the source side
- copy files which are not already known to a "canonical copies" folder on the destination (this step uses the hash itself as the file name, which makes it easy to check if I had a copy from the same file earlier)
- mirror the source directory structure to the destination
- create hardlinks in the destination directory structure for each source file; these should use the original file name but point to the canonical copy.
Then I got too scared to actually use it :)
In any case, a good design is to ask the kernel to do the dedupe step after user space has found duplicates. The kernel can double-check for you that they are really identical before doing the dedupe. This is available on Linux as the ioctl BTRFS_IOC_FILE_EXTENT_SAME.
Restic and Borg can do this at the block level, which is more effective but requires the tool to be installed when I want to check out something.
Of course, engineering being what it is, it's possible that only one of these has hardware support and thus might end up actually being faster in realtime.
You can group all files into buckets, and as soon as a bucket is empty, discard it. If in the end there are still files in the same bucket, they are duplicates.
Initially all files are in the same bucket.
You now iterate over differentiators which given two files tell you whether they are maybe equal or definitely not equal. They become more and more costly but also more and more exact. You run the differentiator on all files in a bucket to split the bucket into finer equivalence classes.
For example:
* Differentiator 1 is the file size. It's really cheap, you only look at metadata, not the file contents.
* Differentiator 2 can be a hash over the first file block. Slower since you need to open every file, but still blazingly fast and O(1) in file size.
* Differentiator 3 can be a hash over the whole file. O(N) in file size but so precise that if you use a cryptographic hash then you're very unlikely to have false positives still.
* Differentiator 4 can compare files bit for bit. Whether that is really needed depends on how much you trust collision resistance of your chosen hash function. Don't discard this though. Git got bitten by this.
Using sha256 was a no-brainer, at least for me.
You misunderstood the article, as it's basically doing the opposite of what you said.
This tool finds duplicate data that is specifically not duplicated via copy-on-write, and then turns it into a copy-on-write copy.
libcopyfile also supports cloning via two flags: COPYFILE_CLONE and COPYFILE_CLONE_FORCE. The former clones if supported (same volume and filesystem supports it) and falls back to actual copy if not. The force variant fails if cloning isn't supported.
I modify A_0. Does this modify A_1 as well or just kind of reify the new state of A_0 while leaving A_1 untouched?
https://en.wikipedia.org/wiki/Copy-on-write#In_computer_stor...
But if you have the same 500MB of node_modules in each of your dozen projects, this might actually durably save some space.
I'm not sure if this is what you intended, but just to be sure: writing changes to a cloned file doesn't immediately duplicate the entire file again in order to write those changes — they're actually written out-of-line, and the identical blocks are only stored once. From [the docs](^1) posted in a sibling comment:
> Modifications to the data are written elsewhere, and both files continue to share the unmodified blocks. You can use this behavior, for example, to reduce storage space required for document revisions and copies. The figure below shows a file named “My file” and its copy “My file copy” that have two blocks in common and one block that varies between them. On file systems like HFS Plus, they’d each need three on-disk blocks, but on an Apple File System volume, the two common blocks are shared.
[^1]: https://developer.apple.com/documentation/foundation/file_sy...
So APFS supports it, but there is no way to control what an app is going to do, and after it’s done it, no way to know what APFS has done.
I don't understand what makes you think there's a significant risk of corruption. Are you talking about the risk of something modifying a file while the dedupe is happening? Or do you think there's risk associated with just having deduplicated files on disk?
> Finally, at WWDC 2017, Apple announced Apple File System (APFS) for macOS (after secretly test-converting everyone’s iPhones to APFS and then reverting them back to HFS+ as part of an earlier iOS 10.x update in one of the most audacious technological gambits in history).
How can you revert a FS change like that if it goes south? You'd certainly exercise the code well but also it seems like you wouldn't be able to back out of it if something was wrong.
Run the real thing, throw away the results, report all problems back to the mothership so you have a high chance of catching them all even on their multi-hundred million device fleet.
https://asciiwwdc.com/2017/sessions/715
Let’s say for simplification we have three metadata regions that report all the entirety of what the file system might be tracking, things like file names, time stamps, where the blocks actually live on disk, and that we also have two regions labeled file data, and if you recall during the conversion process the goal is to only replace the metadata and not touch the file data.
We want that to stay exactly where it is as if nothing had happened to it.
So the first thing that we’re going to do is identify exactly where the metadata is, and as we’re walking through it we’ll start writing it into the free space of the HFS+ volume.
And what this gives us is crash protection and the ability to recover in the event that conversion doesn’t actually succeed.
Now the metadata is identified.
We’ll then start to write it out to disk, and at this point, if we were doing a dry-run conversion, we’d end here.
If we’re completing the process, we will write the new superblock on top of the old one, and now we have an APFS volume.
I then tried again including my user home folder (731K files, 127K folders, 2755 eligible files) to hopefully catch more savings and I only ended up at 1.3GB of savings (300MB more than just what was in the NodeJS folders.)
I tried to scan System and Library but it refused to do so because of permission issues.
I think the fact that I use pnpm for my package manager has made my disk space usage already pretty near optimal.
Oh well. Neat idea. But the current price is too high to justify this. Also I would want it as a background process that runs once a month or something.
You "only" found that 12% of the space you are using is wasted? Am I reading this right?
Ex - On my non-apple machines, 8GB is trivial. I load them up with the astoundingly cheap NVMe drives in the multiple terabyte range (2TB for ~$100, 4TB for ~$250) and I have a cheap NAS.
So that "big win" is roughly 40 cents of hardware costs on the direct laptop hardware. Hardly worth the time and effort involved, even if the risk is zero (and I don't trust it to be zero).
If it's just "storage" and I don't need it fast (the perfect case for this type of optimization) I throw it on my NAS where it's cheaper still... Ex - it's not 40 cents saved, it's ~10.
---
At least for me, 8GB is no longer much of a win. It's a rounding error on the last LLM model I downloaded.
And I'd suggest that basically anyone who has the ability to not buy extortionately priced drives soldered onto a mainboard is not really winning much here either.
I picked up a quarter off the ground on my walk last night. That's a bigger win.
You do realize that this software is only available on macOS, and only works because of Apple's APFS filesystem? You're essentially complaining that medicine is only a win for people who are sick.
There are lots of other file systems that support this kind of deduplication...
Like ZFS that the author of the software explicitly mentions in his write up https://www.truenas.com/docs/references/zfsdeduplication/
Or Btrfs ex: https://kb.synology.com/en-id/DSM/help/DSM/StorageManager/vo...
Or hell, even NTFS: https://learn.microsoft.com/en-us/windows-server/storage/dat...
This is NOT a novel or new feature in filesystems... Basically any CoW file system will do it, and lots of other filesystems have hacks built on top to support this kinds of feature.
---
My point is that "people are only sick" because the company is pricing storage outrageously. Not that Apple is the only offender in this space - but man are they the most egregious.
If you read the rest of the comment he only saved another 30% running his entire user home directory through it.
So this is not a linear trend based on space used.
When I run it on my home folder (Roughly 500GB of data) I find 124 MB of duplicated files.
At this stage I'd like it to tell me what those files are - The dupes are probably dumb ones that I can simply go delete by hand, but I can understand why he'd want people to pay up first, as by simply telling me what the dupes are he's proved the app's value :-)
For reference, from the comment they’re talking about:
> I then tried again including my user home folder (731K files, 127K folders, 2755 eligible files) to hopefully catch more savings and I only ended up at 1.3GB of savings (300MB more than just what was in the NodeJS folders.)
You misunderstood my comment. I ran it on my home folder which contains 165GB of data and it found 1.3GB is savings. That isn't significant for me to care about because I currently have 225GB free of my 512GB drive.
BTW I highly recommend the free "disk-inventory-x" utility for MacOS space management.
You wrote: but it only found 1GB of savings on a 8.1GB folder.
It’s quite a saving and that’s what everyone understood from your comment.
His comment is pretty understandable if you've done frontend work in javascript.
Node_modules is so ripe for duplicate content that some tools explicitly call out that they're disk efficient (It's literally in the tagline for PNPM "Fast, disk space efficient package manager": https://github.com/pnpm/pnpm)
So he got ok results (~13% savings) on possibly the best target content available in a user's home directory.
Then he got results so bad it's utterly not worth doing on the rest (0.10% - not 10%, literally 1/10 of a single percent).
---
Deduplication isn't super simple, isn't always obviously better, and can require other system resources in unexpected ways (ex - lots of CPU and RAM). It's a cool tech to fiddle with on a NAS, and I'm generally a fan of modern CoW filesystems (incl APFS).
But I want to be really clear - this is people picking spare change out of the couch style savings. Penny wise, pound foolish. The only people who are likely to actually save anything buying this app probably already know it, and have a large set of real options available. Everyone else is falling into the "download more ram" trap.
When I ran it on my home folder with 165GB of data it only found 1.3GB of savings. This isn't that significant to me and it isn't really worth paying for.
BTW I highly recommend the free "disk-inventory-x" utility for MacOS space management.
True
> and dedupes automatically
Also true.
But the way you put them after each other, makes it sound like npm does de-duplication, and since pnpm tries to be a drop-in replacement for npm, so does pnpm.
So for clarification: npm doesn't do de-duplication across all your projects, and that in particular was of the more useful features that pnpm brought to the ecosystem when it first arrived.
It’s a free app because you don’t have to buy it to run it. It will tell you how much space it can save you for free. So you don’t have to waste $20 to find out it only would’ve been 2kb.
But that means the parts you actually have to buy are in app purchases, which are always hidden on the store pages.
macOS has a sealed volume which is why you're seeing permission errors.
https://support.apple.com/guide/security/signed-system-volum...
In general, you don’t want to mess with that.
My main motivation was for the packages of Python virtual envs, where I often have similar packages installed, and even if versions are different, many files would still match. Some of the packages are quite huge, e.g. Numpy, PyTorch, TensorFlow, etc. I got quite some disk space savings from this.
https://github.com/albertz/system-tools/blob/master/bin/merg...
If it works it's a no-brainer so why isn't it the default?
https://learn.microsoft.com/en-us/windows/dev-drive/#dev-dri...
Each worktree I usually work on is several gigs of (mostly) identical files.
Unfortunately the source files are often deep in a compressed git pack file, so you can't de-duplicate that.
(Of course, the bigger problem is the build artefacts on each branch, which are like 12G per debug/release per product, but they often diverge for boring reasons.)
There is also ".git/objects/info/alternates", accessed via "--shared"/"--reference" option of "git clone", that allows only sharing of object storage and not branches etc... but it is has caveats, and I've only used it in some special circumstances.
Am I missing something, or isn't it a "file de-duplicator" with a nice UI/UX? Sounds pretty simple to describe, and tells you why it's useful with just two words.
> Q: Are clone files the same thing as symbolic links or hard links?
> A: No. Symbolic links ("symlinks") and hard links are ways to make two entries in the file system that share the same data. This might sound like the same thing as the space-saving clones used by Hyperspace, but there’s one important difference. With symlinks and hard links, a change to one of the files affects all the files.
> The space-saving clones made by Hyperspace are different. Changes to one clone file do not affect other files. Cloned files should look and behave exactly the same as they did before they were converted into clones.
Once you open one of the reference handles and modify the contents, the copy-on-write process is invoked by the filesystem, and the underlying data is copied into a new, separate file with your new changes, breaking the link.
Comparing with a hardlink, there is no copy-on-write, so any changes made to the contents when editing the file opened from one reference would also show up if you open the other hardlinks to the same file contents.
With APFS Clones, the contents start off identical, but can be changed independently. If you change a small part of a file, those block(s) will need to be created, but the existing blocks will continue to be shared with the clone.
https://hypercritical.co/hyperspace/#how-it-works
APFS apparently allows for creating "link files" which when changed, start to diverge.
I know that internally it isn't actually "removing" anything, and that it uses fancy new technology from Apple. But in order to explain the project to strangers, I think my tagline gets the point across pretty well.
The duplicates aren't removed, though. Nothing changes from the POV of users or software that use those files, and you can continue to make changes to them independently.
That’s why Copy-on-write clones are completely different than hardlinks.
In times where documentation is often an afterthought, and technical details get hidden away from users all the time ("Ooops some error occurred") this should be celebrated.
Symbolic links, hard links, ref links are all part of the file system interface, not the implementation.
The idea is not new, of course, and I've written one of these (for Linux, with hardlinks) years ago but in the end just deleted all the duplicate files in my mp3 collection and didn't touch the rest of the files on the disk, because not a lot of size was reclaimed.
I wonder for whom this really saves a lot of space. (I saw someone mentioning node_modules, had to chuckle there).
But today I learned about this APFS feature, nice.
Also, I get a data security itch having a random piece of software from the internet scan every file on an HD, particularly on a work machine where some lawyers might care about what's reading your hard drive. It would be nice if it was open source, so you could see what it's doing.
> It would be nice if it was open source
> I get a data security itch having a random piece of software from the internet scan every file on an HD
With the source it would be easy for others to create freebie versions, with or without respecting license restrictions or security.
I am not arguing anything, except pondering how software economics and security issues are full of unresolved holes, and the world isn't getting default fairer or safer.
--
The app was a great idea, indeed. I am now surprised Apple doesn't automatically reclaim storage like this. Kudos to the author.
Edit: this might not work with the payment option actually. I don't think you can IAP without the internet.
I probably wouldn't have risked it, either.
Here’s a question though: how does this work with transparently compressed files on APFS?
In my past experience, using reflinks is fine and using transparent compression is fine, but combining them leads to hard-to-debug file corruption.
Wait a minute, what happens to copies on different physical drives. Are they cloned too?
https://web.archive.org/web/20210506130542/https://github.co...
> Q: Does Hyperspace preserve file metadata during reclamation?
> A: When Hyperspace replaces a file with a space-saving clone, it attempts to preserve all metadata associated with that file. This includes the creation date, modification date, permissions, ownership, Finder labels, Finder comments, whether or not the file name extension is visible, and even resource forks. If the attempt to preserve any of these piece of metadata fails, then the file is not replaced.
Like with Hyperspace, you would need to use a tool that can identify which files are duplicates, and then convert them into reflinks.
When you asked about the filesystem, I assumed you were asking about which filesystem feature was being used, since hyperspace itself is not provided by the filesystem.
Someone else mentioned[0] fclones, which can do this task of finding and replacing duplicates with reflinks on more than just macOS, if you were looking for a userspace tool.
You only get CoW on APFS if you copy a file with certain APIs or tools.
If you have a program that does it manually, you copied a duplicate to somewhere on your desk from some other source, or your files already existed on the file system when you converted to APFS because you’ve been carrying them for a long time then you’d have duplicates.
APFS doesn’t look for duplicates at any point. It just keeps track of those that it knows are duplicates because of copy operations.
It tries, but there are some things it can't perfectly preserve like the last access time. Instances where it can't duplicate certain types of extended attributes or ownership permissions it will not perform the operation.
https://podcasts.apple.com/podcast/id617416468?i=10006919599...
No word about alternate data streams. I'll pass for now.. Although it's nice to see how much duplicates you have
Q: Does Hyperspace preserve file metadata during reclamation?
A: When Hyperspace replaces a file with a space-saving clone, it attempts to preserve all metadata associated with that file. This includes the creation date, modification date, permissions, ownership, Finder labels, Finder comments, whether or not the file name extension is visible, and even resource forks. If the attempt to preserve any of these piece of metadata fails, then the file is not replaced.
If you find some piece of file metadata that is not preserved, please let us know.
Q: How does Hyperspace handle resource forks?
A: Hyperspace considers the contents of a file’s resource fork to be part of the file’s data. Two files are considered identical only if their data and resource forks are identical to each other.
When a file is replaced by a space-saving clone during reclamation, its resource fork is preserved.
He points out its dangerous but could be worth it cause space savings.
I wonder if the implementation is using a hash only or does an additional step to actually compare the contents to avoid hash collision issues.
It's not open source, so we'll never know. He chose a pay model instead.
Also, some files might not be identical but have identical blocks. Something that could be explored too. Other filesystems have that either in their tooling or do it online or both.
sudo tmutil listlocalsnapshots /
sudo tmutil deletelocalsnapshots <date_value_of_snapshot>This uses a specific feature of APFS that allows the creation of copy-on-write clones. [1] If a clone is written to, then it is copied on demand and the original file is unmodified. This is distinct from the behavior of hardlinks or symlinks.
Sources: https://unix.stackexchange.com/questions/631237/in-linux-whi... https://forums.veeam.com/veeam-backup-replication-f2/openzfs...
> If some eligible files were found, the amount of disk space that can be reclaimed is shown next to the “Potential Savings” label. To proceed any further, you will have to make a purchase. Once the app’s full functionality is unlocked, a “Review Files” button will become available after a successful scan. This will open the Review Window.
I half remember this being discussed on ATP; the logic being that if you have the list of files, you will just go and de-dupe them yourself.
If you can do that, you can check for duplicates yourself anyway. It's not like there aren't already dozens of great apps that dedupe.
Most of those delete rather than use the features of APFS.
Back in the MS-DOS days, when the RAM was sparse, there was a class of so-called "memory optimization" programs. They all inevitably found at least few KB to be reclaimed through their magic even if the same optimizer was run back to back with itself and allowed to "optimize" things. That is, on each run they always find extra memory to be freed. They ultimately did nothing but claim they did the work. Must've sold pretty well nonetheless.
You may find this interesting: Investigations into SoftRAM 95 by Raymond Chen [1] and Mark Russinovich [2] respectively.
"They implemented only one compression algorithm.
It was memcpy."
[1] https://devblogs.microsoft.com/oldnewthing/20211111-00/?p=10...
[2] https://www.drdobbs.com/parallel/inside-softram-95/184409937
QEMM worked by remapping stuff into extended memory - in a time that most software wasn't interested in using it. It worked as advertised.
Quarterdeck made good stuff all around. Desq and DesqView/X were amazing multitaskers. Way snappier than Windows and ran on little to nothing.
(Obviously it won't show the list until payment. That part is expected behavior.)
That said, keeping track of blocks and extents for deduplication would be a much more expensive problem to solve.
brew install fclones
Thanks for the recommendation! Just installed it via homebrew.< You can also create [APFS (copy on write) clones] in Terminal using the command `cp -c oldfilename newfilename` where the c option requires cloning rather than a regular copy.
`fclones dedupe` uses the same command[1]:
if cfg!(target_os = "macos") {
result.push(format!("cp -c {target} {link}"));
[1] https://github.com/pkolaczk/fclones/blob/555cde08fde4e700b25...But then I ran this command and saved over 20GB:
brew install fclones
cd ~
fclones group . | fclones dedupe
I've used fclones before in the default mode (create hard links) but this is the first time I've run it at the top level of my home folder, in dedupe mode (i.e. using APFS clones). Fingers crossed it didn't wreck anything.This tool to enable compression is free and open source
https://github.com/RJVB/afsctool
Also note about APFS vs HFS+, if you use HDD e.g. as backup media for Time Machine, HFS+ is must have over APFS as it is optimised only for SSD (random access).
https://bombich.com/blog/2019/09/12/analysis-apfs-enumeratio...
https://larryjordan.com/blog/apfs-is-not-yet-ready-for-tradi...
Not so smart Time Machine setup utility forcefully re-creates APFS on a HDD media, so you have to manually create HFS+ volume (e.g. with Disk Utily) and then use terminal command to add this volume as TM destination
`sudo tmutil setdestination /Volumes/TM07T`
The author talked about being very conservative on launch; skipping directories like the Photo library or others apps that actively manage data or looking across user directories. He stumbled into writing this app because he noticed the duplicated data of shared Photo libraries between different users on the same machine. That use case isn't even supported in this version. He said he plans future development to safely dedup more data--making a one time purchase less sustainable for them.
rmlint -c sh:handler=reflink .
I'm not sure if reflink works out of the box, but you can write your own alternative script that just links both filesTL;DR - $49 for a lifetime subscription, or $19/year or $9/month.
It could definitely be easier to find.
Edit: Apparently both is possible in the end: https://hypercritical.co/hyperspace/#purchase
I worry about that with Procreate. It feels like it's priced too low to be sustainable.
This app though? No chance. Parent comment says “if you want to support the app’s development” but not all apps need to be “developed” continuously, least of all system utilities.
And because it has no Internet access yet (and because I prompted it to use a workaround like this in that circumstance), the first thing it asked me to do (after hallucinating the functionality first, and then catching itself) was run `curl https://hypercritical.co/hyperspace/ | sed 's/<[^>]*>//g' | grep -v "^$" | clip`
("clip" is a bash function I wrote to pipe things onto the clipboard or spit them back out in a cross-platform linux/mac way)
clip() {
if command -v pbcopy > /dev/null; then
[ -t 0 ] && pbpaste || pbcopy;
else
if command -v xclip > /dev/null; then
[ -t 0 ] && xclip -o -selection clipboard || xclip -selection clipboard;
else
echo "clip function error: Neither pbcopy/pbpaste nor xclip are available." >&2;
return 1;
fi;
fi
}When doing something with any risk potential I first ask the model for potential risks with the output, and then I manually read the code.
I also "recreated" this tool with Sonnet 3.7. The initial bash script worked (but was slow), and after a few iterations we landed on an fclones one-liner. I hadn't heard of fclones before, but works great! Saved a bunch of disk space today.
Since it's the kind of thing you will likely only need every couple of years, $10 each time feels fair.
If putting all your data online or into an SSD makes more sense, then this app isn't for you and that's okay too.
(no, it's not a symlink)
"CoW is used as the underlying mechanism in file systems like ZFS, Btrfs, ReFS, and Bcachefs"
Obligatory: https://en.wikipedia.org/wiki/Copy-on-write
macOS 15 was released in September 2024, this feels far too soon to deprecate older versions.
The problem is SwiftUI. It's very new, still barely usable on the Mac, but they are adding lots of new features every macOS release.
If you want to support older versions of macOS you can't use the nice stuff they just released. Eg. pointerStyle() is a brand new macOS 15 API that is very useful.
They are working on it, and making it better every year. I've started using it for small projects and it's pretty neat how fast you can work with it -- but not everything can be done yet.
Since they are still adding pretty basic stuff every year, it really hurts if you target older versions. AppKit is so mature that for most people it doesn't matter if you can't use new features introduced in the last 3 years. For SwiftUI it still makes a big difference.
I used to be a hardcore Apple/Mac guy, but I'm kind of giving up on the ecosystem. Even the dev tools are keeping everyone on the treadmill.
If we are to believe ChatGPT itself: "The ChatGPT macOS desktop app is built using Electron, which means it is primarily written in JavaScript, HTML, and CSS"
https://github.com/tarkah/iced_table is a third-party widget for tables, but you can roll out your own or use other alternatives too
It's in Rust, not Swift, but I think switching from the latter to the former is easier than when moving away from many other popular languages.
I'd argue there's a lot more to iced than just being a quick toolkit. the Elm Architecture really shines for GUI apps
That being said, it's not quite an apples to apples comparison, because SwiftUI or UIKit can work with basically an infinite number of rows, whereas HTML will eventually get to a point where it won't load.
For me on the German store it looks like this:
Unlock for One Year 22,99 €
Unlock for One Month 9,99 €
Lifetime Unlock 59,99 €
So it supports both one time purchases and subscriptions. Depending on what you prefer. More about that here: https://hypercritical.co/hyperspace/#purchaseOf course the Copy-on-Write clone functionality has been available since APFS became the default file system in 2019.
So it’s neither novel or new.
I think author being “famous” around tech circles and Apple fans’ engineered hatred of open source contributed to the rise of this article.
However, considering Apple will never ever ever allow user replaceable storage on a laptop, this might be worth it.
There's value in convenience. I wouldn't pay for a yearly license (that price seems more than fair for a "version lifetime" price to me?) but seeing as this tool will probably need constant maintenance as Apple tweaks and changes APFS over time, combined with the mandatory Apple taxes for publishing software like this, it's not too awful.
Which really means up until the dev gets bored, which can be as short as 18 months.
I wouldn't mind something like this versioned to OS. 20$ for the current OS, and ten dollars for every significant update.
That's why we see so many more subscription-based apps these days, application development is an ongoing process with ongoing costs, so it needs to have ongoing income. But the traditional buy-it-once app pricing doesn't enable that long-term development and support. The app store supports subscriptions though, so now we get way more subscription-based apps.
I really think Siracusa came up with a clever pricing scheme here, given his want to use the app store for distribution.
will produce a script which, if run, will hardlink duplicates
> clone: reflink-capable filesystems only. Try to clone both files with the FIDEDUPERANGE ioctl(3p) (or BTRFS_IOC_FILE_EXTENT_SAME on older kernels). This will free up duplicate extents while preserving the metadata of both. Needs at least kernel 4.2.
There's a place for alias file pointers, but lying to the user and pretending like an alias is a copy is bound to lead to unintended and confusing results
Update: whoops, missed it in your comment. Block (changed bytes) level.
It is really unfair to call it "software" it is more like "glued to recent version of OS ware", meanwhile I can still run .exe compiled in 2006, and with wine even on mac or linux.
Then again, this app was written with SwiftUI, which hasn't received some handy features before macOS 12 and is still way behind AppKit.
When I see an app that's not compatible with the second most recent macOS, I assume the dev either didn't know better or they were too lazy to write workarounds / shims for the latest-and-greatest shiny stuff.
For example if you have multiple node_modules, or app installs, or source photos/videos (ones you don't edit), or music archives, then hardlinks work just fine.
Anyway, congrats to Siracusa on the release, great idea, etc. etc.
The author is a "household" name in the macOS / Apple scene for a long time even before the podcast. If someone is spending all their life blogging about all things Apple on outlets like ArsTechnica and is consistently putting out new content on podcasts for decades they will naturally have a better distribution.
How many years did you spend on building up your marketing and distribution reach?