A few notes:
* By default it doesn't scan everything. It ignores all files but those in an allow list. The way the allow list is structured, it seems like Hyperspace needs to understand the content of a file. As an end user, I have no idea what the difference between a Text file and a Source Code file would be or how Hyperspace would know. Hyperscan only found 360MB to dedup. Allowing all files increased that to 842MB.
* It doesn't scan files smaller than 100 KB by default. Disabling the size limit along with allowing all files increased that to 1.1GB
* With all files and no size limit it scanned 67,309 of 68,874 files. `dedup` scans 67,426.
* It says 29,522 files are eligible. Eligible means they can be deduped. `dedup` only fines 29,447. There are 76 already deduped files, which is an off-by-one, so I'm not sure what the difference is.
* Scanning files in Hyperspace took around 50s vs `dedup` at 14s
* It seems to scan the file system, then do a duplication calculation, then do the deduplication. I'm not sure why the first shouldn't be done together. I chose to queue any filesystem metadata as it was scanned and in parallel start calculating duplicates. The vast majority of the time files can be mismatched by size, which is available from `fts_read` "for free" while traversing the directory.
* Hyperspace found 1.1GB to save, `dedup` finds 1.04GB and 882MB already saved (from previous deduping)
* I'm not going to buy Hyperspace at this time, so I don't know how long it takes to dedup or if it preserves metadata or deals with strange files. `dedup` took 31s to scan and deduplicate.
* After deduping with `dedup`, Hyperscan thinks there are still 2 files that can be deduped.
* Hyperspace seems to understand it can't dedup files with multiple hard links, empty files, and some of the other things `dedup` also checks for.
* I can't test ACLs or any other attribute preservation like that without paying. `strings` suggests those are handled. HFS Compression is a tricky edge case, but I haven't tested how Hyperspace's scan deals with those.
With a FOSS project this would have been expected, but with a ShareWare-style model? Idk..
Again, I've got no problem with people selling software or closed source models, but I've never understood using this justification. Maybe in this instance he's a well known public figure with published contact info that people will abuse?
One does not simply “not accept bug reports”.
https://github.com/sqlite/sqlite/blob/e8346d0a889c89ec8a78e6...
Who says you have to deal with support requests if you open source something?
> All his apps are personal itches he scratched and he sells them not to make a profit but to make the barrier of entry high enough to make user feedback manageable.
That makes no sense
Almost anyone who has ever maintained popular open-source software, even if dealing with them means putting up a notice that says "Don't ask support questions" and having to delete angrily posted issues.
My understanding from listening to his explanation is he wants to be able to support users and have an income stream to incentivize that.
As an open-source maintainer of a popular piece of software, I'm very empathetic.
So GitHub created a mess, and the whole of open source is considered to be GitHub now.
The solution is the same as being able to avoid tons of Windows-related headaches when you don't use Windows. Just don't use GitHub.
A tar or zip file with source code posted online (or bundled with the program, even) under an open source license is still open source.
There's a lot of merit to Open Source. But there's also a lot of spam, politics and drama that comes with opening up. That negativity is invisible to people who haven't encountered it, or are simply guilty of causing it.
Maintainer burnout is real; more power to John for choosing whatever keeps him focused on building good software.
I mean, that's very obviously a false statement. You don't have to post any notices or reply to or delete any issues.
> My understanding from listening to his explanation is he wants to be able to support users and have an income stream to incentivize that.
That's valid, but is basically the opposite of the reasoning provided in gp comment.
Just keeping everything closed is really missing the point of how trust in infra that handles critical data is built nowadays.
Apps like this can easily bit rot, and more users does often mean more work e.g. answering or filtering emails, finding more edge cases, etc.
From his perspective that means having a income to dedicate time to this. I don't think he's interested in being an "infra" app as you would think of it.
As someone who maintains critical open-source software, I can strongly empathize, even if it’s not an approach I would take.
Making software proprietary and for-pay, especially such a small tool, doesn't just significantly reduce the number of eyeballs this testing is crowdsourced to, but it also disincentivises issue reporting .. why should I spend the time to for free report sth to somebody who is making money off my testing and doesn't even bother to be transparent about how things work exactly (i.e. the source)?
If you really care about the quality of your work then maximizing the eye ball count and incentivise high qualith issue reporting.
Though if you want to maximize income instead, you keep it closed and ask for a subscription.
Quite obvious which option he chose.
You don’t have to. No one is saying you are compelled to report bugs in software you paid for. Most people don’t. The benefit to you as a customer is it can help get the bug fixed. That is clearly a mutual benefit.
> If you really care about the quality of your work then maximizing the eye ball count and incentivise high qualith issue reporting.
I think you’re vastly overestimating the value in the “higher quality” bug reports you’re getting from free users. You might get some higher quality reports but you’ll mostly get a lot more noise.
There are limits to how practical it is to allow for more and more feedback and that threshold for a solo developer is quite low. Restricting your user base by charging for your work means that there is less noise because the only people sending bug reports are paid users.
The quality of these reports are probably lower than if you had an open issue tracker, but you are substantially reducing the mental overhead and you know the people that are sending feedback are doing so with their own interests in mind.
An issue tracker, on the other hand, requires active engagement from the developer. Every issue, even low quality ones, require some form of processing, be that responding, closing, or categorizing. While tools can assist a person in these tasks, the developer is ultimately still responsible for it.
I'm not saying people should only create closed-sourced paid software, but I strongly disagree with the idea that it's negatively affecting the quality of the software because there's no open issue tracker for people to post to.
It's not just github. It's every single issue tracker where users can submit feedback, some of which are almost entirely opaque, like Apple's feedback system. Look at Mozilla's issue tracker, or look at the mailing lists for linux. It's a lot of effort which simply is not worth it for a lot of people in a lot of cases.
Nope.
Pretty strange thing to say in the context of a closed source app.
My intent was to express sympathy to making a closed source app instead of an open source one.
Something must have got lost in translation because it fits the context exactly.
It's equivalent to saying 'I worked on a closed source project and liked it, therefore open source model sucks.'
App store prices are localized. If the blog post said it costs “$10” or whatever, that doesn’t mean anything to millions of potential customers who live where they don’t use $, and is confusing for millions more that do use $ but don’t know if the price is in their local $ or USD
The local currency argument is wrong btw. I'm located in Europe and use a Spanish IP. The prices shown are in USD.
There are lots of apps called "Hyperspace" in the Apple app store, by the way.
https://apps.apple.com/us/app/hyperspace-lighting/id15371988...
https://apps.apple.com/us/app/hyperspace-gpt-chats-ai-art/id...
...
I ran it over my Postgres development directories that have almost identical files. It saved me about 1.7GB.
The project doesn't have any license associated with it. If you don't mind, can you please license this project with a license of your choice.
As a gesture of thanks, I have attempted to improve the installation step slightly and have created this pull request: https://github.com/ttkb-oss/dedup/pull/6
I see it is "pre-release" and sort of low GH stars (== usage?), so I'm curious about the stability since this type of tool is relatively scary if buggy.
Whenever using it on something sensitive that I can’t back up first for whatever reason, I make checksum files and compare them afterwards. I’ve done this many times on hundreds of GB and haven’t seen corruption. Caveat emptor.
There is one huge caveat I should add to the README - block corruption happens. Having a second copy of a file is a crude form of backup. Cloning causes all instances to use the same block, so if that one instance is corrupted, all clones are. That’s fine for software projects with generated files that can be rebuilt or checked out again, but introduces some risk for files that may not otherwise be replaceable. I keep multiple backups of all that stuff in hardware other than where I’m deduping, so I dedup with abandon.
I’m a nobody with no audience. Maybe some attention here will get some users.
I was also really impressed that `make` ran basically instantly.
I love the documentation from FreeBSD and OpenBSD. Only having target one platform and only system libraries makes building simple.