Tag Systems
buttondown.email
buttondown.email
The internal tool at G has a strict hierarchy called "components", and fbtasks is tags alone.
This means if you want to file a bug against anything at G, you could navigate this large hierarchy and find a rough spot for the bug. Regardless of that team's documentation, tooling, bug tracking helper forms, etc you could always file a bug that would get pretty close. Then someone on that team would move it into the proper component if you got it wrong.
At FB... you go to file a bug, but you need to know which tags! If you miss the right tag, that team may never see it. All teams implement their tagging differently, and many implement almost nothing at all. This devolves to pretty much all bugs being tracked through internal group posts, because at least internal groups have names and membership lists that are easily searchable. Then teams implement bots & programs to take every post filed in a group, and auto generate a task with the right tags! And pray that they append the `donotreap` tag...
Bugs need to be filed against specific systems, and those systems need to be owned by teams. Those teams are in orgs, those orgs are in companies. I cringe whenever I see a bug tracker for a large company that is not in a hierarchy.
Nowhere is this more prominent than in traditional file systems where the folder hierarchy is often used as a kind of metadata tagging system. You take a picture while on vacation and copy it to your /pictures/2020/vacations/Hawaii/beach folder. But you could also copy it to your /2020/vacations/Hawaii/beach/pictures folder.
Your pictures, documents, music, and source code gets spread out using whatever hierarchy seems right to you on the day you copied or created it. Heaven forbid you decide to change the folder structure using anything but symbolic links so it breaks any stored references (i.e. full path) to the file(s).
for finding untagged things, i had "system:untagged" automatically added if the items had no tags because mixing inner and outer joins was a nightmare in mysql.
I have been working on a new kind of data management system that is designed to replace traditional file systems while also doing a bunch of database operations. It is an object store that makes extensive use of tagging. Nearly all file metadata comes in the form of tags attached to each file. Even its name is a tag, so every file can have multiple names (each object's unique ID is a 64 bit number).
The system is currently in beta and still needs work, but it can do some amazing things so far. It can create 100M+ files and attach dozen of tags to each one (there is a limit of 255 tags per object). Searching for all photos with certain tags, for example, can give you the results extremely fast. Even with containers with 100 million objects can give you the 100,000 that have a specific tag in just a second or two.
Tags have context. Instead of 'John' as a tag, you attach Person.FirstName = 'John' to photos of John or documents written by John. The whole system is like a sparse relational database table where each row/column intersection can have multiple values.
We are looking for more beta testers and anyone can download and try the software for free. Here are a couple short videos showing a few things it can do.
https://www.youtube.com/watch?v=yLPNLHm9fIk https://www.youtube.com/watch?v=dWIo6sia_hw
I’m curious how you implemented them. I’ve been mulling using a tag based system for a completely unrelated project. I figure one might get a lot of way there using hash tables and sets from any of the high level languages. But it seems you maybe used a different approach?
All tags within the same context are stored together in a highly optimized Key-Value store (also stored within specialized Didgets) like how a columnar store database stores all the data for a single column separately. So if you attach a tag like Person.FirstName = 'John' to a photo of John that is stored in a Didgets with ID = 123; then within the tag the value of 'John' is mapped to the key 123 within the Person.FirstName tag store.
This makes it extremely efficient to find all the Didgets where 'John' was attached as a person's first name. It can also find everything where the first name starts with 'J' or a vowel (using regex). The values within each tag store are de-duplicated and reference counted. In the example, the value 'John' is stored once. If 100 things have 'John' attached, then its reference count is 100. This makes it extremely fast to find everything with the 10 most popular names.
The tags were so efficient and fast that I built relational tables using them that are often faster to query than traditional database systems. In this case the 'Key' each value is mapped to is the rowID within the table.
Here is a video showing how that compares to SQLite: https://www.youtube.com/watch?v=Va5ZqfwQXWI
If you want beta testers for a Linux version... I'm more interested in the Manager with an API, not the GUI Browser.
I broadly like this post, but I disagree with "The lesson: don’t let users encode logical paradoxes.".
If a user says they want (A && !A), just hand them back nothing. Let the expressiveness exist regardless. (Ideally, warn the user, though)
I want expressive tags. It makes everything so much easier. If I can't do (A && B) || !(C or D) then as a user I'm going to get frustrated if there's a lot of content to sift through. Grouping and negation are critical, IMO.
The post touches on this, but I'd say by far the largest issue is the moderation of tags. If only the author is allowed to tag, and there's no moderation or oversight from other users or official moderators, the tag system is effectively useless. Some users will be lazy and not tag effectively. Others will tag incorrectly by mistake or intention. You want the system to support new tags for a variety of reasons, but this allows typos to create new tags.
Generally, when I see that a tag system sucks, it's because it's been mis-used, not because the underlying logic is bad.
A program that runs in the background on your machine and watches processes and the file handles they open. Then, keeps a database of associations between files and programs. E.g. you open a pdf with adobe acrobat, and now it adds an entry with a relation between that specific file and acrobat, and between pdf file extensions and acrobat.
Then, this information is used to provide an intuitive interface into the file system. Searching for that .docx file you made last year and cant remember for the life of you where you put it? Just search for .docx, or files associated with "Word"
This to me is the holy grail of filesystem search, and it doesn't seem that hard to implement, but I've never found anything like it.
Make a userspace daemon process that adds eBPF tracepoints[0] to open{,_at} etc syscalls which match files of your user directories with specific extensions (e.g. .docx).
Associate PIDs that open those files with their .desktop entries[1]
Store results in some database like sqlite3.[2]
Search this database with your favorite interface, like a CLI script or a GNOME shell search provider[3].
I have seen this Rust project on HN which does something similar but with file attribute syscalls, you can use it as reference: https://github.com/javierhonduco/sweeper
[0]: https://github.com/iovisor/bpftrace [1]: https://www.freedesktop.org/wiki/Specifications/desktop-entr... [2]: https://sqlite.org/ [3]: https://developer.gnome.org/documentation/tutorials/search-p...
Originally wanted to do it for windows but I eventually concluded it would take some prodigious hackery to get it do work
Anywho, I'm happy for the info on linux. I was looking for a project that would require me to up my bash skills anyway. This sounds perfect
So you'd have to write a driver that hooked itself at appropriate step in IO Request processing, then it could inspect and gather data just like eBPF.
Procmon uses diagnostic interfaces from outside kernel, might require debug permissions.
Spotlight uses the LaunchServices database of file type associations to give friendly names to file types but has a hierarchy of type identifiers[1].
[0] https://support.apple.com/guide/mac-help/narrow-search-resul...
[1] https://developer.apple.com/library/archive/documentation/Fi...
This is a problem that can be solved simply by better organizing your files. No security nightmare Rube Goldberg contraption needed. Ive been in that boat before and it just boils down to organizing your stuff better.
It's basically an extension to the VFS that would maintain atime and some additional metadata to it, what's so horrible about it?
I just think windows tools for navigating the file system are straight garbage and think it could be done a million times better by just indexing interactions between processes and files. But who knows maybe im barkin up the wrong tree.
Also- organize your files better : first off easier said than done and second off, my shit is already a total mess and I am not willing to take off work for 3 weeks to organize everythibng. I just want a file viewer UI that is actually useful in allowing my to find files.
Imo the folder hierarchy paradigm shouldnt have ever been exposed to users as the primary mechanism for accessing files. It makes no sense for 99 percent of use cases for the average user.
But also in the case that I do not know the extension of a file, or want to find all files involed in configuring a program, etc, it would allow you to find that info quickly.
Search is best done via Everything by Voidtools, which is instant, though it doesn't connect to this database, but instead uses manually entered lists of extensions
But you could write integration that extracts the list from the OS database and dumps it in Everything's lists
Or you could integrate Everything in you file manager and then add the OS file associations integration there
I never implemented this before but the tags can also hold references to the tags they are most frequently tagged with to further add content to the depleted search results. Never thought how the implementation would look like though.
My hierarchy of tag systems goes something like
key, key=value, hierarchical key(the unix filesystem), hierarchical key=value(oh shit you found ldap, abort, you went to far)
a file is a generic unformed array of bytes, the application can store anything it wants in this array. files are selected via tags, a file can have any number of tags, however there is a limitation that each tag can only reference one file. tag prefixes are use to group and match common tags.
Having a system that can point out unknowns is useful both for search (by including them in results) and for filling in that data (by suggesting those stories for review).
This is a very interesting point. My initial thought was: there definitely IS a correct choice, which is "no". If the choice were "yes", that would imply there's a difference between the tags, but they're aliases, so there shouldn't be a difference. An alias should be exactly interchangeable.
Then I thought about how practically impossible that is to enforce, given the nature of language. The usage choice itself implies a difference; there are no true aliases in language.
If someone wants to type in dumbbell and someone wants to type in weight, and it goes to the same tiny picture, is that really so broken?
I think we're ultimately in agreement — aliases should be 'transparent' and not imbued with any specific meaning. Of course, it's much easier for a very limited system like emoji to be able to deal with these things than an open-ended tagging system, especially one that caters for the entirety of Unicode.
The same person will choose different sets of tags for the same file, when asked on different days; they will end up with synonymous tags that should be coalesced, but nobody wants to do that drudgework.
The semantics of each tag will alter over time and by person.
An optimal tag system is inferior to a full text search system with an indexer that recommends keywords and has a reasonable syntax for includes, excludes, and boolean logic.
however, your suggested search system doesn't work - many people built these and they weren't right - because they don't encode whether or not the object had attention.
One of the things I'm curious to try out with GPT is auto tagging content and patching documents that have missing tags. You train GPT with your corpus and existing (and incomplete) tagging system, and have GPT fill in missing tags and revise tags that are synonymous or spelling variants.
Maybe there are already good non-AI system that do this, I haven't really looked deeply into other solutions yet.