Designing better file organization around tags, not hierarchies (2017)
nayuki.io
nayuki.io
I remember about 10-15 years ago that there were a lot of photo organizers that supported tagging. Tag the events as birthday or the people's names. I saw a lot of people who would then post "My photo organizer has been discontinued and I spent years organizing by tags. How do I migrate all the tags to another photo editor."
The hierarchical file system is kind of a lowest common denominator. I can copy a directory structure from ext4 to brtfs to zfs to ntfs to exfat to macos or whatever and have the same organization.
The universal standard is Dublin Core (IETF RFC 2413 / ISO 15836) `dc:subject` “Unordered array of Text” stored in an XMP (ISO 16684) sidecar file:
- https://en.wikipedia.org/wiki/Dublin_Core#Dublin_Core_Metada...
- https://en.wikipedia.org/wiki/Extensible_Metadata_Platform
I like to put the DC stuff directly in my filesystem xattrs, but using a sidecar file will get you compatibility with any old filesystem and with any tooling that might destroy or disregard an xattr.
If you were serious, then I would point out that even nerd file browsers like Nautilus do not support xattr metadata editing, searching, or even viewing and Linux metadata indexing systems e.g. Beagle are all long-dead projects. The whole idea has no traction.
I currently use a combination of my own filetype-detection-and-tagging software “CHECKING YOU OUT” (insert Dropbox meme here), Windows Explorer with “XMP IFilter Desktop Edition”, and voidtools' “Everything” 1.5:
https://www.ifiltershop.com/downloads/xmpfilter/readme.html#...
Which, in the form of folder and file names, can be your the first layer of tags. But familiarity with file systems is going out of the window as well, I've been made aware.
I found the article from 2022 about younger people not understanding them.
Gen Z Kids Don't Understand How File Systems Work
I tire of the either-or dichotomy. Yes, search is great. But a couple years ago, I adopted Johnny Decimal[0] for storing my archive files and I couldn’t be happier. Our 2023 tax filings are in 22.23. Physics articles are in 34.04. Her CME records are in 43.03. Go into one of those and you’ll find every file I have related to those things.
And they’re still searchable. You can still do tag-based queries, indexed text searches, and all that, but you don’t have to. When I want all my 2025 tax deduction receipts come tax filing time next year, I’ll send the contents of 22.25 to my accountant and be done with it, confident we found 100% of the related files.
I don’t do those for everything. Music files are in their own directories. Movies live with movies. I don’t try to shoehorn Obsidian into it. But all my long-term storage files go straight to their home in JD.
This has the huge benefit that I can instantly locate them when I’m accessing the archive from my phone, my wife can find things without learning special tooling or the vagaries of search, and there’s zero need for any additional apps or software beyond Finder.
I love using great search software. I used Devonthink daily for years. But I also adore not having to.
I had hoped you meant that not being "either-or" would be that you should be able to get "realized hierarchies" on things. Where you could basically ask the computer to build a hierarchy on the fly.
In my case, I got a lot of value in coming up with that on my own. The result may not be rigorously perfect. In fact, I know it isn’t: some folders have lots of files, and some have few. It maps really well to my brain, though. That makes it easier for me to use, because I’m not memorizing a hierarchy someone else invented and pushed upon me. I’m using the system that sounded about right to me in the first place.
That way, for a tax year `2024` containing expense receipts, the accountant or us can just look at the file name to confirm. If you still need to look, then opening the file works.
I use Hazel on my Mac to auto-rename bank statements etc like “$closing_date Visa statement.pdf” and move them to the right place automatically as soon as I save them to the ~/Downloads folder.
All my new files look exactly like you’re describing. I give my accountant a set of files like “2023-06-08 Goodwill donation receipt.pdf”
It’s slightly more effort to do this work up front, but pays off in spades when I need to actually retrieve something. Even without any fancy search software, just looking for files named like “2023receipt*” locates them.
The problem is acute in organizations, where the question of ontology ownership crops up. The last place I worked at had a dedicated engineer on the job. And marvel at the elaborate systems Wikipedia and StackOverflow have to manage their ontologies. It is best to avoid that stuff and improve search instead, which is likely why an ontology was sought in the first place.
Maybe there is room for both, though I am in agreement that it seems dubious a human should be curating.
tensorflow object detection is one such means. Hosted solutions would be Amazon Rekognition.
Only for information that has no real value.
Search misses things. Hierarchy or tag ensures that all you have marked can be found.
AI might be intelligent enough to fix that, but in the end it’d just be automating the curation of tags/ontologies (ideally in manner that’s user-visible and editable)
LLM powered search can.
>but in the end it’d just be automating the curation of tags/ontologies
Automating those is the whole point though.
The only mechanism that will ever be properly able to contextualize the relevance of my ingested information when it counts (as in while I am alive) is me; the only way to do this is by curation. A someone else who understood this once said: "The tool shapes the master; the master shapes the tool".
Manually assigning tags are an excellent way of doing this, as long as one doesn't overthink it to such an extent where the curation process meets the harsh realities of diminishing returns. Sadly, many people, constantly jumping from one fad to the next, already fail here.
A few nights ago I was looking through photos with my daughter, and she asked to see photos of her. I looked for them using the AI way and found some, but not a whole lot. That's because it missed the vast majority of them, which I then found when looking through files with the same date stamp as some of the AI discovered photos of her.
I understand we're on our way to automatically labeled content, but we are a long way off, and the time for manually organizing important information has not yet passed.
I wonder if there's a counter-argument to be made that humans organize their knowledge of the world in ontological hierarchies, so a hierarchical file system is intuitive.
I think the idea will go away, similar to how ontologies and Semantic Web, manual knowledge graphs etc gave way to processing unstructured data with LLMs. Instead of tags, we should have source/creation-context-based metadata and embeddings computed with language models. Then you can do natural fuzzy search.
Tags also fail because people then want to categorize their tags. Which tag is the author's name, versus the editor's? Publish year, versus year I read it? I suppose we could say that people want "slots" not "tags"?
I have my _really_ important files that are cumbersome to recreate like a <50kb jpg of my signature for those annoying forms, scans of my identity cards, etc, in ~/Documents/important
I have my to-read stuff in ~/Documents/books
All my other files are slammed together in ~/Downloads. I search through the folder using everything.exe (on windows) or fsearch(on linux).
That already covers the 90% case, and for the rest recoll maintains a search index that even stores file contents (including zip,pdf,xml-based stuff like office files). It works well. These days I hear everything.exe itself supports searching within files.
On the other hand, I'm really picky about properly hierarchically organizing code files in any SW project I do, and proper hierarchical organizing of the app folder of any app I make. As for tags, the only reason I see myself using them is for pictures. It can enable a deterministic version of Google photos search. But I'm happy with Google photos search for now.
Azure blobs support multiple queryable tags per file, which aligns with the idea of organizing by metadata instead of folders.
It’s not as deeply integrated or user-centric as the article suggests, but it offers a practical usage, especially for large-scale, developer-driven storage.
However, if you have multiple overlapping hierarchical organization, common operations such as "moving" and "deleting" become overly complicated.
IMO, there's no best way, only different pros/cons and compromises.
I know this is about an OS and not a CMS, but it seems storing your documents in a repository like this would be useful - not sure about the rest of the OS though.
Classic taxonomies are a thing of the past.
It's probably easier to just through everything into a clustering algorithm and autotag than going back and tagging tens of thousands of files.