TMSU: a tool born out of frustration with the hierarchical nature of filesystems
tmsu.org
tmsu.org
I've strung together something using bleve full-text search and some OCR libs to scratch my particular itch but it still doesn't quite get all there.
I actually like the hierarchical organisation, and I don't like the 10 year usability trend, particularly driven by Microsofts attempts to patch over their horrid structure with even worse workarounds.
The solution to learning my mother were to find her documents is not hiding the place 15 levels deep and having 10 symlinks to it.
Fast, unobtrusive indexing and a good structure is all that is needed. Humans organise and memorise things in hierarchies, and while categories is another helpful abstraction they tend to be confusing if they do not point to a fixed position in a hierarchy in my opinion.
You can also explore a hierarchy; Not so with a cloud of things connected to a bunch of tags.
I don't think I've seen tagging systems that let me tag tags out in the wild.
The tagged objects define the relationships between differing tags which allows you to define that arbitrary graph. The graph can be viewed as an effectively infinite depth tree if you wish to.
You can see this twin representation conceptually as Artist->Album->Song and also as Compilation->Artist->Song, the latter being a small folder with symlinks to the actual song directly.
It's really just two different trees.
I think that this is extremely important. You mention iTunes, which does its best to pretend that there's no hierarchical file system underlying it. There is, of course, but—ugh, one look at it will make you long for a totally flat file structure.
I think that it's important that any hierarchical-file-system 'killer' actually not replace (or render meaningless), but only supplement, the often-useful hierarchical structure; and it looks like this tool does that.
(In case that looks like an argument, let me say explicitly that I am agreeing with you.)
Big claim. Proof?
https://en.wikipedia.org/wiki/Prototype_theory#Basic_level_c...
https://en.wikipedia.org/wiki/Stereotype#Cognitive_functions
https://en.wikipedia.org/wiki/Concept_learning
Maybe less abstract: if you try to remember where your phone is, do you follow a pattern like "which room", "which coat/pants/desk", then "which pocket/drawer"?
Conceptual categorisation + language + ? is totally how we humans comprehend and analyze the world around us, granted. (The devil is in the details of course.) But I don't see how we can go from there to "it relies on a hierarchy of defining attributes"
I don't buy your example because the concepts in our heads seem to be organised in context-specific word bags, i.e. tags.
I think I would believe you if cognitive science discovered that tree-hierarchy navigation relied on built-in neuro-spatial machinery or something like that :)
I _want_ to believe you but I also want proper proof! :)
edit: thanks for the heads up on prototype theory, can't believe I'd forgotten about this. also, thanks for making me ponder more deeply about stereotypes.
I'm not a developmental psychologist btw, just have an interest in cognitive development. The proof you're looking for would be more the field of cognitive neuroscience.
There definitely are graph-like connections between concepts all over the brain, I didn't mean to suggest that our brains work hierarchical only, to the exclusion of all other relations. As an example: if our brains were strictly hierarchical, we would be unable to see the similarity in colour between a grey table and a grey hare. I just meant to explain my thinking on how abstract thought is tied to a attribute/property hierarchy, not proclaim that I know how the brain works.
(edit: snipped on-topic content, best reserved for a separate post)
File systems are arranged hierarchically. Fluke of development? Library catalogues are organised hierarchically. Lucky coincidence? Books themselves are laid out hierarchically. I see a pattern emerging, but is it a trick of the light?
Seems like we like to chunk and splay amorphous informational units into graph-like rooted structures. Maybe cuz of its flexibility?
For something we use all the time and something we do all the time language-use and conceptualisation are deeply mysterious.
I'd say it's fundamental, not any kind of mystery.
Btw, Frege, following Hume has pointed out that abstraction is an equivalence relation: http://logic.uconn.edu/2015/01/21/reference-and-invariance-i... I wouldn't lump simplification and abstraction together, and characterising them both as associations feels wrong.
What I'm trying to say is that your time would be better spent in not correcting every instance of mild hyperbole on HN :)
And mine too. :)
What I'm trying to say is that your time would be better spent in not correcting every instance of mild hyperbole
What was I correcting, exactly? I was giving a casual opinion. To call a casual, armchair hypothesis that differs from someone else's "correcting"... well let's just say you failed spectacularly at understanding human language patterns.
I think it is fundamental. A hierarchy forms naturally from iterative steps of aggregation and differentiation: first, group all similar objects together; then, for each high-level group, look for differences within each group, and split the groups into smaller subgroups. Rinse and repeat until you've reached an acceptable number of items per group.
Humans also track a lot of things based on spatial memory, which no OS since Mac Classic has even tried to make use of, which is a shame.
[0] https://en.wikipedia.org/wiki/Spatial_file_manager
[1] https://en.wikipedia.org/wiki/File_manager#Navigational_file...
http://arstechnica.com/apple/2003/04/finder/2/
Any software developer interested in designing usable software should read and digest this. Alas, nobody working at Apple did, so I'm no longer an Apple customer.
So you had the option of remembering "Excel is the icon below the Special menu" and you could be confident that the Excel icon would always be located below the Special menu, exactly where you put it.
One of the reasons I switched from OS X to Windows is that Apple threw that brilliant design in the garbage. Shame.
Well, sure, if things do have a hierarchical relationship, but a hierarchy is just one form of network. It fails for concurrent dependency problems, e.g.
All of the other files I work with on a day to day basis I don't really want to query by structure or tags. I just want to know where the %^&*@ they ARE! And a simple hierarchy is actually better than tags for that.
We inherit all the information from your existing file system structure, as well as the people files are shared with, and provide a powerful content-based search on top of that.
Check out our website if you'd like: https://www.meta.sc
That and the fact that it's available offline on my machine.
I know that I'm not always searching for files along the same hierarchy: sometimes I want to search along a timeline, because I know that I edited two files around the same time. Or I remember where I was when I wrote something, and I'd like to search on geolocation. Sometimes I know that I've implemented a similar feature for a different project, and want to copy/review what I did then.
All these searches still have hierarchical components in them: in the "date" case, I would expect to browse to a generic "N weeks ago" directory, look at a few specific items to get my bearings, and start browsing deeper to narrow my date range, or move forward or backward in time. In the geolocation case, I'd want to pull up a map, click on a country, then a city or area, and look at the files I accessed while there. As for the project example, I probably have the project files lying around in a hierarchical directory already. But I'd want to access only the relevant feature file, not browse the entire project structure again.
For me, the problem with non-hierarchic interfaces isn't just the lack of metadata: it's a lack of tools for visualizing non-hierarchic matches. Every search I described above requires its own GUI...
I have more than 5 todo.txts scattered in my system. I have cloned my webpage's git repository countless times in my local drive. I have countless LaTeX documents named letter.tex.
Of course, if I search for letter.txt I get a hit, but what other documents in my system did I create around the same date? were any other associated file changes? these questions are hard to answer in a hierarchical filesystem.
While we're at it, we should separate the task of document naming and classification from document saving. Right now, if I create a document in say, Microsoft Word and I want to make sure it gets saved, I have to pick not only a discoverable filename but also the right place for it in the hierarchical tree. I am just trying to type up a quick letter and get it printed, not grapple with the philosophical question of "where does this file belong and what should I name it?"
One comment mentioned BFS, which had some really cool stuff. There's an Ars Technica article that touches on some of it[0].
The secret to BFS, in my mind, is that applications use it. The Haiku Mail app, as noted in the article, used the filesystem as its email database by attaching its own attributes to messages. This is also an example used in the "Practical Filesystem Design with the Be Filesystem" book[1].
Unless the metadata becomes a first class citizen in the filesystem, any attempts to layer it on top will have problems. Either applications won't understand it or normal filesystem operations will cause the metadata database to become de-synced with the filesystem data.
[0] http://arstechnica.com/information-technology/2010/06/the-be...
[1] http://www.letterp.com/~dbg/practical-file-system-design.pdf...
Remember when it seemed like Mac OS might give us a modern era of rampant metadata (http://arstechnica.com/apple/2005/04/macosx-10-4/6)? Ah, those were the days.
MacOS does an OK job of helping the user find things with Spotlight, but it's not a full metadata system like BFS had.
Mail.app, for example, keeps each message in a separate file[0] (and probably has a cache or separate database of this to make displaying mailboxes quicker). This makes it easy for Spotlight to index, but all of the stuff that you'd think of as metadata is actually just regular data inside the .emlx file.
If Apple made a huge effort to start treating the metadata (assuming the infrastructure described in the Ars Technica article still exists) as a first class citizen and using it like BeOS did, maybe we can get there. This would be a drastic rethink though. It feels like files, in some ways, are becoming second-class citizens in the Mac world. Photos, for example, are managed in the Photos app - you do not go into the filesystem and organize your photos.
One big problem with filesystem metadata is how do you transfer it? The Ars article showed a sidecar file (._filename) being created when the file was copied to a non-HFS volume. Now the metadata is detached from the file and we're back to the same problem.
Also I like that they describe what data they actually change on your computer right on the homepage: "TMSU does not alter your files in any way: they remain unchanged on disk, or on the network, wherever you put them. TMSU maintains its own database and you simply gain an additional view, which you can mount, based upon the tags you set up."
Unfortunately building on a foundation of sand (meaning not TMSU's code, but Unix filesystems) has downsides:
https://github.com/oniony/TMSU/wiki/FAQ#why-does-tmsu-not-de...
" Why does TMSU not detect file moves and renames?
To detect file moves/renames would require a daemon process watching the file system for changes and support from the file system for these events. As some file systems cannot provide these events (e.g. remote file systems) a universal solution cannot be offered. Such a function may be added later for those file systems that do provide file move/modification events but adding support for this to TMSU is not a priority at this time.
The current solution is to periodically use the repair command which will detect moved/renamed files and also update fingerprints for modified files. (The limitation of this is that files that are both moved/renamed and modified cannot be detected.) "
Ouch.
they could have gone the other way, throw everything into
a DB, and then wrote a fuse plugin to access it all
through traditional file system
This is the Camlistore strategy! Of course, there are other problems with that approach
Could you elaborate more on these? I've never worked with FUSE.Specifically, I was referring to the different off the shelf database systems which could be used. Each will have it's own benefits and drawbacks to storing large chunks of data per-record. Benefits might include (relatively) easy sharding or replication. Drawbacks might include not being space efficient for removed files, not being as resilient to corruption due to crashes or corruption affecting more than the files in use, or overly aggressive use of memory to function efficiently.
If a custom database was developed, you could tailor to your exact needs, but then you have much more work to do, and a period of immaturity.
Off the top of my head, if I were designing a general purpose system for tagging files where people were expected to use it as a regular file system and some overhead from FUSE was acceptable, I think I would leverage the file system but in a different way. I would set up a specialized directory for the files themselves, and store then hashed within it, and have a BerkelyDB database relate filename to hash and tags, and use FUSE to do direct file access. But that's my 5 minute assessment, so I reserve the right to change it completely given someone pointing out the obvious problems. :)
Tagging is a really useful idea, it is also a naming thing and as such either it lives in the naming infrastructure (aka dirents) or it rots over time. A simple example I used to use in the 'object naming' [1] days was, imagine that instead of house numbers on the street you wrote down last names. That works fine until somebody moves and now not only did you show up at the wrong house, you don't even have a chance of knowing what the correct house is. [2]
Microsoft's LongHorn project was way out there but took a swing at the actual problem. Just make the file system an actual relational database. Then your home directory is simply 'select * from files where (owner = chuck);' It really does solve the problem at a more fundamental level, using naming by attribute rather than mapping. I got to observe that effort from the outside (I was at NetApp at the time) but I believe it died due to really horrible performance issues.
I find it pretty awesome that people can lose files, back when a "big" hard drive was 100MB it really wasn't all that hard to just look through all the files on it, but when its a couple or three terabytes, all bets are off!
[1] Object File systems were all the rage in the early 2000's, files themselves were object ids and the naming was a database that connected object ids to user recognizable names. -- https://en.wikipedia.org/wiki/Object_storage
[2] The typical solution is to add "tombstones" or redirects at the previous address. That then is a layer of additional meta data to maintain, and sometimes the file doesn't move, it just changes value (trivial example you have a file 'my-favorite-song.mp3' which is tagged 'jazz mp3' and then you discover techno and make something from Tiesto your favorite song and while the name and type are still valid, the tag 'jazz' is now invalid.
Then, it's OK if the original file gets renamed or moved, as long as it stays on the same FS. You still have your hardlink, and so your symlink still works.
Also for not tech savvy users, folders have the nice interaction pattern of question-response via menu selection. They see something, click, see something, click, without realizing that they are navigating a folder hierarchy.
I have a couple of points/concerns after a quick read.
1) So 'tag' is the verb that updates the records, but 'tags' is a read?
Poor taxonomy, IMHO. If I'm using "tmsu" I know I'm working with tags, so I'd think natural switches would be the way to go ( tmsu add, tmsu ls|list, &.) Or at least don't make the "create" verb and "read' verb differ only by the plural 'S.'
2) Is there a way to list all existing tags? —not just the ones already bound to a file but all available in the database (with regex filtering, of course).
That's what I'd need to pick from my 'tag pallette' before actually tagging so I could avoid creating synonyms accidentally that'd later require a merge.
Specialized files like music, movies and images are a solved problem. iTunes and other software do a great job at organizing this information and making it easy to use. There's also software like Quicken for dealing with the other common mountain of data people have.
What might be useful is a piece of software that is capable of extracting and managing metadata automatically. Think of a tool like iTunes that you feed it a collection of files and it uses some form of ML to extract and create a database build logical ontologies for this data. The big problem with this kind of tool is finding a large, complex dataset that an individual has, but that has not been organized by a specialized piece of software. I doubt these exists in numbers significant enough to justify creating a project.
tl;dr: Directory structures are low-effort, generic, and discoverable ways of dealing with files that are not managed by other applications. It will be hard to improve on them without sacrificing one of those three attributes.
With liberal use of tags, and the ability to browse them fluidly, we could ask those sort of questions.
At the top.
The way that one person manages their documents probably isn't going to be the same as another, but generally a person is consistent with all their files.
In this case, it's either a document that you have several types of, in which case you would have an existing folder structure, or it's a one-off that you dump in My Documents along with all your other one-offs. Even if you lose it, file systems all keep date-time information, so you can easily search for the all the files last modified yesterday.
The problem with liberal tagging is that it requires a bunch of up-front effort that you're never going to perform. In this case, if it were an email, you'd just use your email browser (specialized software!) to find the document again based on the things you remember about it (date, source, size, etc).
That 'my documents' thing - you could easily create a 'view' on tags that yielded that result. Without relying on Microsoft or whomever to do it for you.
mkdir -p ~/tags/{music,big-jazz,mp3}
ln -s /path/to/summer.mp3 ~/tags/music/
ln -s /path/to/summer.mp3 ~/tags/big-jazz/
ln -s /path/to/summer.mp3 ~/tags/mp3/
From there, you can use all normal filesystem tools to interact with your 'tags'. You could extend this with a simple script that handles duplicate file names in the same tag by sticking a hash of the file before the extension. Having a separate database for this information seems unnecessary.* Doesn't account for files with the same name
* No helper methods like merge when you typo
Just a couple of reasons off the top of my head
Edit: Misread. Yes it would work. Well, consider this a tool for doing it automatically, without filling your hd with bogus folders and links
What you want is to also have AND, OR and NOT available. How would you find 'music' (and) 'mp3' (and not) 'big-jazz'?
First, do a 'ls' directory dump of the 3 folders. Then 'cut' out the softlink destination from each file. This will end up with 3 files with soft links. From there, 'grep' the music and mp3 files together to get a list of soft links that are in both music and mp3. Then you can do an inverse 'grep' to remove softlinks that are in big jazz.
Even if the system creates hard links instead of softlinks, then the same could be done via inode numbers
One problem is that regular people don't know how to organize a hierarchy. General -> specific works really well, but that requires the ability to generalize.
The only use I can see for tags is if you want files to be a member of more than one directory. Other than music, everything I have is generally created for a specific purpose, so tags are not particularly helpful.
In fact, nearly every large app does something to avoid it. They create their own representations of a log, or a mail folder, or a document, or an image (and on and on) and manage the details themselves. Because 'file systems' are so lame and underpowered.
This tool begins to help. Creating flexible groupings (tags) resembles a relational database. That's a start. I'd like to replace the OS file system with something like that.
Instead of renaming files when you bring another copy onto your persistent storage, you could just add a version tag to them. Leave the names alone! I can tell my build system (or document store, or mail tool) what version I want to deal with e.g. tag='version' value='2.5'. No collisions any more. No requirement by the 'file system' to mash them into some file tree so they can still be found, but don't collide.
In fact this system can do everything the hierarchical file system can do, and more. Just add the tag 'parent directory name' and voila! You have a file tree (if you want).
i.e. `music mp3 folk` results in a different virtual file system than `folk music mp3`?
I often think of tags as an unordered set, rather than an ordered list.
My default view is a detail view sorted by "date accessed" (descending) which is what I need 99% of the time. Especially handy when uploading random images, quick edits etc from the browser.
btw I highly recommend https://pathcopycopy.codeplex.com/ for those that use a terminal and win explorer at the same time a lot.
I suspect there is an inherent scaling problem with the approach. While it demos well and can be useful in a lab setting, the full scale version ends up being so complex that mere mortals aren't able to comprehend it and end up just losing their files constantly in the morass of schemas and structured accesses.
The talk on the wiki about being able to write an editor that can handle any file type smells of pure insanity to me. An application that can present a usable editing interface for any filetype would be so bloated that it's almost inconceivable to think about a single company writing it. Can you imagine the kind of application that has an interface for manipulating photos, writing reports, tabulating data, programming (all languages), doing CAD, editing textures, composing music, etc...?
I think the more general-purpose a tool is, the more intelligence and/or general knowledge is needed to wield it:
"Here's a widget thromper. You push the big green button and it thromps the widgets."
vs
"Here's a car. It will get you from virtually anywhere to anywhere else. Don't kill anyone."
vs
"Here's a general purpose computer. Knock yourself out."
In my last "proper" job, I wrote a lot of code, and the vast majority was one-off; virtually none of it ended up in what you might call an "application".
Different name, exact same promises.
These:
$ tmsu tag summer.mp3 music big-jazz mp3
$ tmsu tag --tags "music mp3" foo.mp3 bar.mp3
$ tmsu tag spring.mp3 year=2003
Are confusing. Very non-Unix. They should be:
$ tmsu tag music,mp3,year=2003 summer.mp3
The biggest problem is that I have to figure out which tags are going to be useful to me in the future and where to add them. This is relatively easy for music (but even there can explode in complexity depending on how granular you want to be), but more difficult for things like photos or papers.
IMHO, fully general tagging systems never work because the complexity explodes as the number of potential tags increases. You need to narrow the scope down to a specific domain so your tags can be limited to human scale.
I've been working on ideas for managing filesystems and tools for thought for a while off and on, it's starting to coalesce into a set of design ideas as well as prototypes.
What I won't do is set up a filesystem: I don't think that adds value to me; it mostly sounds technically complex and hard to figure out. And, I think that driving the whole business through manual tagging is a lost cause. Manual tagging can be _useful_, but actual attempts to derive semantic knowledge will be _more_ useful. I have many thousands of documents - manual tagging ain't gonna happen.
I need _semantic_ search and _semantic_ cross-referencing; something like a Xanadu or a (much better) wiki/hypertext system.
So, depending on the need, I will either have two directories (files and tags) or a number of file directories and a tag directory. File directories can have whatever they want, and tag directories have either only tag directories (like workout, driving, etc) or soft links.
Tagging a file / directory is easy - just link to it. Untagging is just as easy. The links don't take up much space, especially next to the music. When I'm transferring the files, I either use a script to make a directory with the links replaced with their file counterparts, or I transfer with something like rsync that can do that itself.
I'm amazed how similar this project is to my own solution. It's nice to have a dedicated script for the whole thing, but the solution itself is very simple, and easy to script with.
This isn't born out of frustration - hierarchical filesystems are perfectly adequate for most tasks. But they have not been trees for a long time - we have links, which let us make any graph we want out of those trees.
Today, it's still all ready to use. I haven't touched it. I'm actually quite happy with the way my filesystem works, I just had this idea how great it would be to work with tag selections instead.
The only reason I might still use a tagging system is to tag some files I want to back up manually (if at all), like a 50GB disk image or some temporary big download, but in general I create one or two symlinks a year and I'm good. The hierarchy works fine.
Same here. I have a few things that tags would be nice for, but it's infrequent enough that current filesystems are fine. Couple of examples
* Tagging the source [CD/iTunes/Amazon etc.] of my music - tags would be nice (and possibly doable as IDTags etc.) but "/music/source/artist/album" or "/music/artist/album [source]" works fine
* Multiple paths to the same file - I have various media (movies, music, books, TV shows etc.) related to a single series in one folder. Those should also be in my main music etc. folders. Again tagging with the series name to make a virtual "series" folder would be nice, but symlinks solve that and it happens infrequently enough that it isn't an issue.
Other than those edge cases, I'd say most of my data fits pretty well into a hierarchical structure.
There is a frequent need to produce documents for discovery purposes. However, you generally have to review the documents or send the to outside counsel first for review before they are produced. This can take hundreds of man hours and cost thousands of dollars.
Building something on top of TMSU could be a great solution for this task.
I'm sure FUSE can do this.
Ahhhhhhh!
You're tackling a laudable goal, but it would help to link to a page with more technical details.
On top of that we try to get further accuracy by using user information to help us find the context. We call it a Personal System, a repo of knowledge and learned preferences heavily tailored around each individual user. We're just in beta now but definitely put your email down if you want to try it out eventually!
My personal feeling is that (A) you're totally right that we need better _personal_ organization systems, but (B) the bottom layer (tagging, schemaing, relationships) should not involve fuzzy processes but should be totally understandable by the average user.
I know (B) is a weird opinion though so I look forward to seeing how far you can get with (A) plus machine learning & whatever other tricks you guys plan:)
So for us, we very clearly distinguish "a user labelled this item" vs "we guessed it was this label".
I'll keep you posted!
I'm not so sure this is a great design decision. Now your tags are only in your sqlite file, and you'll have to work extra hard to get a copy of the relevant tags when you backup/copy etc.
I think storing tags in extended attributes[x], and possibly a separate utility that maintains and index (hopefully shouldn't be needed just for the tags, but might help with a) exposing file-level tags (like ID3, exif, file-type (magic number) etc), and b) allow for automatic organization based on full text and other content-based indexing.
It appears, on a * nix system, the only major reason to stay away from extended attributes (apart from the limit on size of tag data) is NFS. But samba should (AFAIK) work fine with extended attributes.
As far as I can gather, Gnome Beagle is dead, and Gnome Tracker[t] has taken its place. But it's not crystal clear if Tracker will index tags placed in files' extended attributes or not. If I understand correctly, Tracker's own tagging utility, will only place/edit tags in the Tracker database/index. But the indexers will certainly honour file-level tags for some files.
I don't really use full Desktop environments, but some kind of system with inotify support, and a Xapian or similar back-end (like Tracker), does seem like a good idea. It would certainly be nice to see such a system implemented in Go, but I think an architecture along the lines of Tracker is probably worth keeping: A database daemon, an indexer and a set of query/view tools (I'm not a fan of the centralized tag database, though).
Another alternative to Tracker would be Recoll:
http://www.lesbonscomptes.com/recoll/index.html
[t] https://github.com/GNOME/tracker
[x] http://www.lesbonscomptes.com/pages/extattrs.html
Btw, for editing/automating ID3 tags, I recommend "Ex Falso", the tag-editor for Quod Libet (which is an audio player): http://quodlibet.readthedocs.io/en/latest/
Here's an article showing how they work for Linux, OSX, and FreeBSD: http://www.lesbonscomptes.com/pages/extattrs.html