On reading this, I felt like I already sort of knew this, and this internet stranger validated my thoughts.
On reading this, I felt like I already sort of knew this, and this internet stranger validated my thoughts.
Then I smack it down, because that is a crap-load of effort to recall a link every year or two. And let's be honest, the marginal value of that link isn't all that great either... in the moment the need may seem large, but sitting here typing about this I couldn't tell you even a single such thing I've forgotten about, because that's how important they are... just more ephemera in the stream themselves.
My MP3 collection is a bit of a mess. I've cleaned up the worst instances of "Band, The" "The Band" "Band" "Band - The" sorts of duplication, but that's about it. My book collection is similarly messy. Heck, even my family photos are basically sorted only by year and not much else. So what? I can fix it. I can fix it all. But it's hard to even so much as recover the time I'd put into it once over, let alone in multiples.
(Much more important, especially for the family photos, is not losing them. So I've got a backup solution. But it's just a fire-at-directory solution, not all gloriously organized by type either.)
So I've learned to just sort of let the desire to have greater organization pass over me, Litany-against-Fear style. It's just a siren call.
I'm through at least three complete reorganizations where I even dusted off old backups and collected all photos (because I felt I was deleting photos too liberally last time), de-duplicated (and de-quadruplicated) them all, and built the new "forever" structure.
It sucks that I just know I'll do it again at most five years from now.
But reading your comment gave me a thought: filesystem hierarchies are indeed insufficient, but what about filesystem hierarchies with liberal use of hardlinks?
* A database for backlinks. (Links from file X to file Y would only be possible when X has an appropriate file format -- `.txt`, `.md`, `.org`, etc.)
* A search grammar with the following primitives:
* find children of (links from) query results
* find parents of (links into) query results
* take the disjunction (OR) of queries
* take the conjunction (AND) of queries
* group queries with parentheses
* The ability to pipe files found via ordinary shell commands into that grammar.
Given the size of most peoples' knowledge graphs, you wouldn't even need to keep a text index (ala Lucene) -- `find` and `grep` would be more than sufficient.I figure that a combination of wetware and software is the current sweet spot. My brain usually has enough associations and context to turn every photo search into a time or place filter - "I think it was downtown last year" or "some time in summer at home" or "it had my wife in it". The photo storage system need only provide search/filter on date and place to narrow it down to a few hundred thumbnails, plus machine-learning to tag people. Which is basically what iOS provides, no more, no less.
Any other up-front categorization or tagging is basically wasted effort.
For me, I wrote a bulk tool that renames my photo file names by reverse geocoding the GPS information via Open Street Maps. That way I can do text search for place, as well as 2d map search. It's at https://unto.me
Problems local to my machine, not Orwellian nightmares.
photos and pictures organization should be a solved problem.
(That is, let you search using words for things in the photo or themes like “winter”).
~/pictures/2022/202212/20221225
And I tag them with as many tags as I can be bothered with using XnView. XnView lets me find pictures by name or by tag.Tagging them with embedded IPTC tags is the way to go. DigiKam works (mostly) as a substitute for the late Picassa. (I'd use that, but the last version has a nasty bug in that it sometimes swaps faces in the recognition database, which then tends to corrupt it all).
The major problem I've had is that in the beginning I didn't really have enough free disk space to keep up. That is no longer an issue, nor is it likely to be again.
MusicBrainz Picard cleans and labels your music automatically using sound signatures, even when the file has no metadata. You can just give it your files and let it run, rarely have I felt the need to monitor it. It gets stuff right 99% of the time, the rest 1% is easily fixable whenever you come across it.
Well, that makes two of us.
I have many questions, about backup and disk space. I’m going to give it a try.
Thanks!
As an example, take physical paper. Receipts, bills, anything that ends up in one's mailbox and doesn't immediately get recycled. I used to oscillate between two approaches: over-elaborate filing systems and just ignoring the problem and letting the mail pile up in snowdrifts.
Eventually I realized that my love of elaborate systems was a giant fucking problem for my actual life. I thought about it like I was designing a production system. I very rarely needed to retrieve old documents; most of it was for "just in case" conditions. I needed to frequently file things, and if the cost of filing was too high, I wouldn't pay it. So I bought 8 filing boxes, each 3 or 4 inches high and big enough to comfortably hold legal-size paper. Each one is marked with a year, and almost everything for that year just gets tossed on top. A few exceptional kinds of paper then have their own separate file folders (e.g., tax documents, my current landlord, key retirement paperwork, key medical stuff). Once a year I throw out the contents of the oldest box and relabel it.
This works great. It turns out I almost never need anything from an old box. When I do, it's a quick rummage in one spot. With infrequent, hard-to-predict retrieval, storage-optimized organization is the best organization.
For paper, I don't have much trouble. Things go on the fridge if I will need them soon, or in one small sterilite plastic file box. It's nowhere near half full and I would not be surprised if it lasts 10+ years before I need any more storage for paper.
As an experiment I've been working on sorting things by category in a more general way, like the dewey decimal system rather than true categories, to remove the overhead of half full containers used to sort things. They're based on observation of what was already stored vaguely together rather than starting with an idea.
One common category is BAM, bulk artificial material. This includes paper towels, laundry soap, paint, water repellant spray, etc.
Another is TAM, tapes attachments and materials, containing tape, steel wire, foam, webbing, carabiners, key split rings, screws, and all similar things often having to do with either attaching things together or long things sold by the foot.
With wider categories I have fewer places to memorize, and organization within a category isn't that critical because they can be rummage-searched, without the overhead of a buch of individual drawers or boxes in some ever evolving system. It's just a formalization of random boxes of junk.
The main place where I'll depart from that is with higher access frequency. E.g., I have a box that is sort of "LRU tech stuff". In there my various cables are sorted by type into gallon ziplocs. Finding the right kind of USB cable is something I do too often to want to pull it out of a 50-wire snarl.
My first line of defense with tech is volume reduction on the stuff itself. Bluetooth over cables, software over hardware, USB-C over everything else, resisting any kind of random tinkering gadgets in favor of phone apps and zero-friction stuff I'll set up and never bother to upgrade, etc.
If you never buy random stuff just because it looks cool, then when you really do want something you can afford to future proof it and sometimes get one thing that replaces multiple separate things.
I have one box dedicated just to power(Which I may split up into separate cables and batteries boxes) that has all my USBs not in active use.
USB-C has been fantastic. If you don't do a lot of fancy stuff with high data rates, it's an amazing way to collapse down the number of unique objects.
This is what I do. Things organize themselves organically this way - the important stuff bubbles to the top, and the rest is already sorted by linear time.
Anything more than that is a recipe for burnout - and I've definitely been there.