TMSU: Command-line tool for applying tags and viewing virtual tagged filesystem
tmsu.org
tmsu.org
It almost feels like a personal categorization version of the "AI Bitter Lesson": people keep thinking that doing a bunch of manual taxonomy work is going to help them find files faster but eventually search catches up
This has always felt like one of the primary issues with tag-based lookup over hierarchical. By the time you're knee deep with enough stuff that you realise tags would help, you've already accumulated too much to practically deal with.
That and figuring out what the tags should be upfront and hoping you don't realise you need additional or different tags later on.
That's because they are already tagged via path, or in the file. I'm just going to wait for multimodal LLM tagging solutions to catch up, rather than just try to hack it with current models/tech.
This. I spent many days cataloging, tagging, deduping and organising my photo and data files, programs, bookmarks, etc.
And I've barely used any of those photos or data files since. The time invested totally wasn't well spent, and I should have just left everything called "DSC0000565.jpg".
Edit - Someone mentioned befs but deleted their comment, it seems like it might sorta be supported in modern linux, possibly just read only though:
On my to-try list, there's also supertag[3], a tag-based filesystem that's mounted via FUSE
[1] https://dogsheep.github.io/ [2] https://github.com/dogsheep/dogsheep-photos [3] https://amoffat.github.io/supertag/
I thought about photos and EXIF tags, too. Duplicating the data from the EXIF into another repository strikes me as a bad idea. That's why I was pining for BeFS.
(I have a lot of crazy ideas about filesystems (arguably more like digital asset management systems) and data ingestion and export. Ideas kind of like the failed WinFS. Nothing will ever come of it because I don't have the skills or the time, but sometimes in fever dreams I imagine this stuff.)
That doesn’t work when I want to use Capture One, Lightroom does not apply Phase One calibration profiles which makes it useless for them, or my own raw processor for Sinar digital backs.
Recommending the most common digital photo DAM/editor is not really a helpful comment either. The number of people who know what exif is and don’t know about Lightroom has to be…small.
The TSMU examples for mp3 files + VFS are similar to BeOS.
One of the BeOS advocates - Scott Hacker - created bash script for ripping CDs into MP3s called RipEnc. It would query the CDDB to get the metadata - track names/artists etc, so the files would be renamed from TRACK1 to e.g. "Dead Milkmen - Punk Rock Girl" for the CD. It would then convert the CD tracks to MP3 files. The metadata would be added both in the MP3 ID3 fields, as well as to the extended attributes of the files in BFS, and it would organize the music in folders by Artist or Album or something.
You could then have a query - a virtual folder/directory that lists files based on extended attributes - all mp3 files by ARTIST foo, and from ALBUM bar, that would stay updated if the file metadata changed. I can't remember if this virtual directory was available at the command line - or if it was only available in Tracker (the native BeOS/Haiku file manager).
The problem with this, and it's not just a BFS problem, is that the metadata in the file and about the file get can get un-synced, either when updating it, or transferring it to another system that doesn't support the extended attributes.
Additionally, Linux _does_ support tagging files right in the filesystem via the user.xdg.tags xattr. Although it looks like Dolphin is one of the few userspace tools that knows about it.
[0] https://archive.org/details/bitsavers_jpsoftware_65101374/pa...
[1] https://4dos.info/4tools.htm#02
[2] http://www.optimasc.com/products/fileid/4dos-descext.pdf
but this seems even better, this is why I am on hackernews
- "Designing better file organization around tags not hierarchies" [1]
- `tag` - a macOS version of `tmsu` that uses the system tags (xattr-based if I recall) [2]
[1] https://www.nayuki.io/page/designing-better-file-organizatio...
https://wiki.archlinux.org/title/Extended_attributes
That being said, it'd be cool to see a port of that CLI to Linux using user.xdg.tags. You can avoid deleting them if you're careful.
https://github.com/TagStudioDev/TagStudio/ https://www.youtube.com/watch?v=wTQeMkYRMcw
https://man7.org/linux/man-pages/man7/xattr.7.html
Since it is baked into the file system, it is pretty easy to create bash scripts to add keyword tags by parsing the directory tree (e.g. batch add tags to books, movies, videos, etc stored in hierarchical category directories).
I use rsync to back up my book, music, and video collections (which is where I use them) and the meta data is backed up with them - so if something ever happens I can always restore them. The xattr commands also have a backup and restore for just the xattr built in.
then use ln to add tags.
$ ls -l # How grep for files with tag "foo"?
$ find . -tag foo find . -samefile group/beatles/love_me_do
album/please please me/side 2/1
vocals/paul mcartney/love me do
vocals/john lennon/love me do
year/1963/love me do
group/beatles/love me do
You have to love ontology to go down this route. I do, and did this once... It is possible, but does not really provide any meaningful advantage.Imagine the opportunities if a folder structure could represent a "document" where each file represents a paragraph, or image, chunk of that document. We would be able to do 'block-based editors' (like content management systems, or Jupyter Notebooks) without having to have some large XML file holding everything.
Even if we had simple "ordinal" (ordered position) for files that would open up endless opportunities for innovation in the 'block-editor' space, but sadly File Systems development has been frozen in place for decades.
Or is this simply the Mandela Effect and xattrs didn't exist in the universe I've been living in, and I've jumped? haha.
E.g. `__o_car`, where `o` means object, or `__p_supercode`, where `p` = project, `__t_ml`, where `t` = topic, ml = machine learning, etc.
No dependencies, hardcoded into the files forever, and search is reasonably fast too (don't need it that often anyway).