Taxonomy Is Hard
autodidacts.io
autodidacts.io
A real world example of how tagging can be both super useful and get out of hand is open street maps. The only meta data allowed in there are tags. The OSM community depends on people using tags correctly. But of course the tagging is incomplete, inconsistent, and subject to regional and data source specific variations. Which makes interpreting the tags a bit of a dark art.
But it still adds up to a very complete and rapidly evolving world map. So, there's that.
I've had the pleasure of seeing the documentation for Navteq's internal meta datas schema that they used for their maps while I still worked for Nokia Maps, which at the time owned Navteq. These days the whole thing is known as Here Maps. This was a PDF of around 4K pages. Thousands of attributes. Lots of weird little details related to traffic lights, subway entrances and exits, and other features you have on maps. This stuff gets complicated quickly.
Two very different approaches to the same problem. I think I like the OSM way a bit better. Neither is easy.
My job at the time was trying to align some of that data with data we got from external data sources like TripAdvisor, Qype, HRS, and a few others. Way too much time got invested in dumbing down categories, mapping one to the other, and trying to make sense of stuff. We had lots of issues with duplicate POIs because of all sorts of subtle differences in how different data sources were annotated with categories, tags, etc. Some of the data just isn't that good, complete, or consistent and you have to deal with that.
My conclusion is that composition is a better method. Has a. And composition by capability is even better. Not great, but better. I was part of several efforts by standards groups to come up with SNMP standards for storage devices. What is a disk? And it has a controller right. Wait. Now the controller is in the disk. Wait now there are RAID systems. Wait...
Trying to build a system of composable tags for functionality seems like it might have a chance. Can it store something? Can it retrieve it. Does it have a retrieval time? Etc.
Another annoying problem is the categories for Craiglist. Where do cooking pots and pans go. For sale. A locality. Household items. Really? Nothing under cooking or even kitchen? I'm not saying this is wrong, I'm just thinking about how do I go about finding stuff other than just "search title".
Both library science and cognitive linguistics have solutions for/discuss this problem.
That said, it's not obvious that the underlying data model needs or should map to a theoretical model for how human cognition works. Maybe it should and it's worth considering the theoretical landscape before setting out, but there are other features that come into play.
I know very little about library science, but I assume they take a more practical (for this task) approach to this that's worth looking into.
That is, embrace the restrictions that come from a taxonomy. It is a bit of a lie, but a lot of it is defined on contrasts and these contrasts help build things.
What distinguishes a car from a truck? One is more faceted to holding people than to holding cargo. But if you don't know either a truck or a car, you are unlikely to know from just that description. Indeed, you could wind up with a cargo van. Which is different from a passenger van, in much the same way.
To that end, we present the taxonomy as a heirarchy of descriptive properties. When, in many cases, it is a heirarchy of representative samples. In programming, we distinguish between class based and prototype based object oriented representation. In reality, most taxonomies are presenting class style based relationships, but using prototype based representation from a population.
Composable and namespaced tags is exactly what we do in my company. It gives us a lot of flexibility and it mirrors what OSM does. You can do a lot of duck typing against such a system and it will work as good as your data is without breaking completely when you get bad data. You just loose some features.
when you can use it as a truck (ie. when it can do what the car isn't adequate for).
The fuzzy line between truck and car isn't a problem when you approach categorization with this mindset. If you see a bank robber fleeing in an El Camino, you can tell the cops they fled in a "truckish car" or a "carish truck"; you don't have to neatly categorize an El Camino as one or the other to get your point across, but the arbitrary categories still help you communicate the idea.
If you're tagging pictures and come across an El Camino, and you don't have a "coupe utility vehicle" tag, you can simply tag the El Camino as both a truck and a car.
It seems what matters is not how files are stored/organized, but how one can find the files one is looking for. Taxonomy is mostly a search problem.
Yet, although one is usually capable of remembering specific or unique details about a file, it's still incredibly hard to search for a file or its contents effectively.
Dropbox, to pick just one example, ditched all its advanced search functionalities a couple of years ago and only allows to search for any keyword (not all of them), with forced stemming, and no filtering of the results.
I seem to remember there were startups a decade ago trying to address the search problem on the desktop but they all got acquired or folded; I don't understand why. The problem is real and seems quite solvable; yet it seems there's no actual market for it. It's a bit of a mystery.
Work migrated to the cloud.
Important for technical users, not so much regular people.
Underappreciated aspect of this is trust. As much as "taxonomy is mostly a search problem", search is in large part a trust problem - trust that the search was done exhaustively, and if it returned no results, it means there aren't any.
This is a problem that IMO most offerings are blind to. One big offender for me is Windows Explorer. It has a search tool that can search both in names, metadata and file content. Yet in the past, it frequently failed to find files I knew were there. Truth is, it may not have been even looking in the right place - it depends on what's been indexed and how, which is information you can find somewhere in the system, but notably not in the search interface itself. I've stopped using it long ago, as I don't trust it at emotional level. I prefer to literally walk the filesystem structure by hand.
And, on Windows, I at least get that option. This is the fallback, the baseline: even if I don't trust a search engine to be exhaustive, I can cope with it as long as I can perform an exhaustive search manually. I.e. as long as I have a way to list everything. But this, then, is almost universally missing from cloud offerings.
(There's a whole rant to be written about the incredibly dumb idea of hiding the filesystem / database from the end-user, but I'll skip it today.)
> there were startups a decade ago trying to address the search problem on the desktop but they all got acquired or folded; I don't understand why. The problem is real and seems quite solvable; yet it seems there's no actual market for it. It's a bit of a mystery.
Nobody really wants to solve this problem anymore, because it conflicts with the major thing vendors want: for you to move your data into their cloud. The desktop, and the ability to own your data, survives only because most cloud systems are still shit. This might eventually change, and in the meantime, the market definitely isn't interested in helping you own your data.
Maybe, but that doesn't explain why there are no startups on the case, that wouldn't have this conflict of interest.
The trust issue you mention may be a clue.
Most people have Word Documents, PDF files, and other things that need a more complex indexing strategy. Also a lot of people have lots of image and audio files which pose their own challenges, namely indexing textual metadata and possibly some indexing of the content.
Computing is a supply-driven market. The demand is there, it's just partially latent, partially ignored. Vast majority of technology users have no choice but to choose from what's being offered, and the minority of tech-savvy users with opinions are increasingly too small a niche to support "power user" tools.
(1) Performance. Back in the day people frequently turned full text indexing on Windows off because it would slow down their computer too much. People won't be happy with the overheard of a search engine that is always scanning 10's of GBs of document.
(2) Search quality. People are familiar with Google being an effective search engine but they've certainly tried a number with terrible relevance scoring and probably learnt that it is not worth trying the search on a web site or the search on the help of an application. In the case of Elasticsearch the default similarity is BM25 but that has two tunable parameters. There are other similarities you could use but most of them have tunable parameters. It makes a real difference what you choice and there is a methodology for tuning them that is now built into ElasticSearch.
https://www.elastic.co/guide/en/elasticsearch/reference/curr...
I talked to about 20 vendors of full text search products and found that only 2 out of the list regularly evaluated the quality of the results, 1 of them just did it so they could get some advertising by being on the TREC leaderboard. They told over and over again that customers didn't care about search quality, they just wanted to see a list of 350+ data sources that the product could index.
This seems to describe corporate buyers ticking boxes on a form, more than actual users.
Desktop-focused databases are of course widely available (Paradox, Access, LibreOffice Base etc.) but the volume of data you could manage with them is of course limited. I.e. non-toy datasets will naturally be hosted on some sort of "cloud".
What's the difference?
> Heuristic search is going to be non-exhaustive by definition.
To me, a non-exhaustive search is broken by design, because it will miss results. Making data accessible only through such search means there will be a hole in the system into which some of the data falls, never to be seen again.
https://www.elastic.co/kibana/
At work we have a search interface powered by Postgres that uses the full text index but also uses GIN indexes on arrays to index things based on categories/tags.
In the scenario where you're searching for a file in a group of low-hundreds of files, it's nothing but a bad cop-out excuse.
And yes, any taxonomy is by definition always a subjective, biased construct. But this is often useful in helping to reveal the preferences and motivations behind the creation and curation of the corpus itself.
Tags are better but can turn out to be even harder (for similar reasons, amplified combinatorially).
Labels are the most practical. The GMail inventor was genius. Folders must begone (except for system files).
E.g. the "javascript" tag implies the "programming" tag. I can even speculate there probably is a trait (linguists may suggest) which applies to some programming languages as well as to some spoken languages. This way tagging arguably can become even harder than taxonomy.
While a label is just a label.
"Tag" — if I understand GP's point — implies some structure. In the above examples, it would rather be "ByLocation::Planet Earth::Europe::Netherlands::Amsterdam", "BySubject::Animals::Dogs::Amsterdam(MyNeighborsDog)", and "MyProjects::ByYear::2021::ProjectOnAmsterdamStreet" — or similar. In this case, if you're interested in the neighbor's dog, you're searching by its specific tag, or — if you don't remember its name — a search for "BySubject::Animals::Dogs" might help you. Any such search will also keep the other "Amsterdam" results away.
The problem with tags is how exactly do you implement them. Your project on Amsterdam street may have begun in 2021, but still ongoing in 2022 with no end in sight; or you may have forgotten its dates at all. Additionally, "MyProjects::ByLocation::MyCity::AmsterdamStreet" is not an invalid way to tag that project.
For research papers (in PDF), I have a half-baked python solution I wrote myself that cobbles together the cermine pdf parser/content extractor, the whoosh full text search engine, and an ncurses-based interface.
For personal images, I use the elodie CLI tool, but I'd like to move away from it as I don't like how it modifies files by embedding metadata in them. For research/computer vision data, I use custom tooling based on sidecar files and a pg database kept in sync with the sidecars . For audio samples, I just use a commercial solution, sononym, that uses an sqlite database.
For other miscellaneous use cases, I've also used TMSU. Pretty nice as a more general purpose solution, except for the inherent issues mentioned by the article.
So yeah, I agree it's a hard problem.
But yeah, I can't argue that tagging solves the problem at hand.
Your options for finding out are:
1. Step back and do some research, look at all the tags and find ones that someone might have used to describe something similar, and then check the original item to make sure.
2. Ask your team if anybody has tagged something like this before.
3. Institute a sort of merge request process for new item tags.
4. Make up a new tag and move on.
Guess which option people usually go with?
Specific problems with tagging: - Need to tag every file (whereas in folders, you just navigate to the folder and everything you store there is in that folder) - Takes too long - Too much thinking overhead (at the time of storing) - To be effective have to enter the name for all tag entries (e.g. project, type, etc.). If anything is missed for a file, that file will never be found. - You have to remember what tag categories (e.g. project, type, etc.) you have used. If you don't use in 3 months, now you have forgotten. - You have to remember the enumeration you are using for some tag categories (e.g. for type you might decide to use only photo, video and music. Now you have to remember that. You also have to remember its "photo" not "image") - If tags were the solution, they would have already been used everywhere. The tagging system SEEMS like a good solution, but once you go deeper, it just doesn't work.
For an example of where such a system is used, see:
- WikiData[6]: They use RDF but they don't use OWL, they have a similar though less formal way of defining types, and then entailment has to be encoded in SPARQL queries.
- schema.org[7]: Uses RDF and also provide OWL specifications.
[1]: https://www.w3.org/TR/rdf11-primer/
[2]: https://www.w3.org/TR/owl2-primer/
[3]: https://www.w3.org/TR/rdf11-mt/#entailment-rules-informative
[4]: https://www.wikidata.org/wiki/Wikidata:SPARQL_query_service/...
[5]: https://www.wikidata.org/wiki/Wikidata:SPARQL_query_service/...
[6]: https://www.wikidata.org/wiki/Wikidata:Main_Page
[7]: https://schema.org/
https://ropercenter.cornell.edu/ipoll/
It is hidden in plain site, for instance if you read this
https://www.w3.org/TR/owl-ref/#Restriction
you might not realized that there are some other axioms documented elsewhere in the standard that let you query on "less than", "greater than", "matches a regex", etc. Instead of fighting with Protege or writing turtle you should be able to compose queries out of restrictions plus complements, unions and intersections the way you snap together a program in Scratch, but somehow the idea that OWL works as a query language (logically define a class in terms of its properties) had eluded people completely.
Of course, it all have to fit into a single disk drive, and actually deleting a file is difficult but hey, that must be easily solvable, details are left as an exercise for the reader.
- new textual entries (headings) goes into monthly notes ($org-roam-directory/timeline/year/$month-name.org), they might attach files or not of course. Attached files are generally directly linking in the heading/inside the textual content of the heading for a single-click quick access and glance view. Doing so allow to have not too many too small files, not too big ones who operate slowly;
- another subdir of org-roam-directory is for "topics", a note per topic, linking or org-transcluding (slow and a bit limited but still useful) the collected entries in timeline style;
- another is workdir where I craft my catalogue (using org-mode drawers created with templates to allow easy org-ql queries) and queries to explore my notes in different view. It's not as easy as TiddlyWiki transparent transclusion but allow a certain degree of practical usability, fine grain selection and ease composition.
MOST of my files and config live or as org-attachments or tangled from org-mode. So yes, taxonomy is hard, but we have tools to master them IF we decide to discover them and invest time in improving our digital garden for real instead of leaving classic mess of files hoping for some miracle "application" that solve automagically all issues. Unfortunately due to the lack of interests by most leave such systems too little developed to be as effective as they can...
My personal experience is:
- we need taxonomy anyway, just mere full-text searching with extras à-la-google do suffice for a certain percentage but fails more than that;
- we need taxonomy that are a bit flexible in storage terms and can change at a slow peace;
- we need integration, witch is NOT possible in ALL modern software, we need for that classic desktops where the OS was a framework/live image and anything is just a module, a bit of code, of it. With end-user programming concepts because no UI can be effective enough in "no code" style and no "modern programming" styles are usable for user programming.
A bottomline: people should learn a bit about information management at school, from how a library or a pharmacy organize books/meds on their shelves to book's indices and personal information archives. Nothing exaggerated but the bare minimum to understand how to manage data, digital and physical in various forms for a lifetime...
Sometimes you want to locate a specific item - in which case you need a good way of searching - and sometimes you want to browse through related information so you want to see a hierarchical structure.
Google drive was originally designed on the principle that search was all you needed so it was all tag-based. And it was terrible as soon as you had a lot of data. So google was forced to introduce the ability to create a folder structure.
[0] https://github.com/slowernews/hamster-system#hamster-folder-...
I tried for a few years to store things in a more hierarchical way or use tags, but it was too much mental effort to think about it every time. I also tried something similar to symlinks when google drive used to support putting files in multiple locations, and it was surprising how confusing it was to manage. Unless you apply symlinks or tags in a very consistent way, it just ends up being frustrating finding something that you think tagged or linked in one way but actually, you tagged it in another way or forgot to tag it at all.
I think spatial datasets/spatial aptitude either makes them relevant or just an untapped avenue for exploration.
Interesting article and most likely part of 21st century technology on many.... levels.
- From computer science, we know graphs give us expressive modeling capabilities. I sometimes use mermaid ER diagrams as a concept map to capture complex relationships between files and concepts.
- From library science, faceted classification works well for extensive collections because inserting a new entry does not require thinking about existing entries. I maintain entries in a spreadsheet for extensive collections that matter to me. Note: Facets are meant for unchanging or infrequently changing properties. Creating a concept map and maintaining a faceted classification system take work, so I only use them for things that are very important to me.
90% of files I only care about for a short amount of time. I use the file system to co-locate the files I'm currently working on (so a project) but then archive all of it when I move on to something else.
The trade-off is that I give up on sharing files between projects. I don't want to deal with references. I copy from the archive when I need to. On the rare occasion when I need to reconcile the same file between projects, I do it manually. What helps is working on only a few projects at the same time.
TL;DR: Archive more. Use high-investment techniques only for the small percentage of files that really matter.
KDE has decent support of tags in extended attributes. Tags are shown and can be edited in Dolphin (file manager), gets indexed by baloo. This is far from stone age as author claims!
Biology. Everything is a fish or nothing is a fish. Trees don't exist. Tomatoes are fruits as are cucumbers, pumpkins, bell peppers, and most things we don't consider fruits. But all fruits are also vegetables. Strawberries are neither straw nor berries. Etc.
When they said the two hardest things in computer science was naming things and cache invalidation, it's partly because naming things is a hard problem in every discipline.
The trouble with that though is that people have different perspectives on documents.
The librarian in me wants to ingest a document and never modify it, such that content addressable storage is what I want. I want to attach metadata in an external database.
There's another culture though where people edit documents, most notable in Adobe's tools which will try to save a JPEG even if all you did was print it! Adobe developed
https://en.wikipedia.org/wiki/Extensible_Metadata_Platform
which embeds metadata in the files which fits that point of view.
In the health data exchange format FHIR you have identifiers and codings have a system and a value/code. Usually you can specify multiple of them.
"Those that belong to the emperor" is an alright category as any animal could belong or not belong to the emperor. You get in trouble working in a system where that has to be disjoint from "bird".
"Those that look like flies from a long way away" is a category that contains all of them.
A while back I was interested in databases like DBpedia and Freebase where you don't really state that anything is disjoint from anything else as most of the categories overlap with other categories. For instance, a person can be an Emperor and an Actor and a Wrestler... See Nero Claudius. Wikipedia doesn't necessarily distinguish between a video game, a manga and an anime that have the same title, so unless you are going to split the topic you're going to have something that is all three even though somebody might think those classes are disjoint.
Lately I have been involved in OWL modelling of financial messages and there, as in some other domains, most classes are disjoint -- in some cases that's a deliberate decision of the modellers, in other cases it is fundamental to the platform I am sucking data out of.
Multidimensional latent spaces based on content and other characteristics of the object is the real solution here.
> Symlinks are brittle.
On the Mac OS, aliases are a lot less brittle. You can move aliased files around, and they usually won't lose their connection.
> Well, since you ask, here's Hans Reiser's old stuff:
https://reiser4.wiki.kernel.org/index.php/Future_Vision
https://reiser4.wiki.kernel.org/index.php/V4
(and http://lwn.net/2001/1108/a/reiser4-transaction.php3 )
. And here's some emails etc. I wrote in response:
https://web.archive.org/web/20040728044342/http://www.st-and...
https://marc.info/?l=linux-kernel&m=111624697710426
https://www.mail-archive.com/reiserfs-list@namesys.com/msg09...
https://www.mail-archive.com/reiserfs-list@namesys.com/msg20...
https://www.mail-archive.com/reiserfs-list@namesys.com/msg20...
https://www.mail-archive.com/reiserfs-list@namesys.com/msg20...
https://www.mail-archive.com/reiserfs-list@namesys.com/msg20...
, plus some of the discussion threaded from those posts. (Sorry, my stuff needs rewriting and updating but I'm not in the position to do it at present. If there's anything you would like to ask about please do. https://news.ycombinator.com/item?id=9809041 and https://news.ycombinator.com/item?id=10548477 touch on things that are a bit further down the line, but related—in particular, to the handling of "internal metadata" and files with a compound internal structure.)
https://web.archive.org/web/20040728044342/http://www.st-and...
https://marc.info/?l=linux-kernel&m=111624697710426
https://www.mail-archive.com/reiserfs-list@namesys.com/msg09...
https://www.mail-archive.com/reiserfs-list@namesys.com/msg20...
https://www.mail-archive.com/reiserfs-list@namesys.com/msg20...
https://www.mail-archive.com/reiserfs-list@namesys.com/msg20...
https://www.mail-archive.com/reiserfs-list@namesys.com/msg20...
Sorry about the breakage!
* Partitioning files into folders is most likely wrong: Most things need to be in multiple folders.
* (Sidenote: Symlinks are not a good solution)
* Tagging would be best but no good support from the OS for metadata
* We're working on something
This is a topic that is at the top of my mind, as I grapple to organise my growing gemini/gopher site. Is it better to index, list a table of contents, search, or try to classify it with the (DDC) Dewey Decimal Classification.
The DDC. It has come under criticism, and librarians have justified a lot of their efforts in moving away from it. I doubt that the effort was justified. It boils down to this: you have to put a book in a library somewhere. And that somewhere has to boil down to a taxonomy.
To illustrate the problem, is a book about programming microcontrollers a book about programming, or is it primarily about microcontrollers? The Arduino Cookbook is in DDC 621.3810285536 (yes, really. That's obviously extreme, though). That's part of the electronics section, which seems fair enough to me. So far so good, But "Beginning MicroPython with the Raspberry Pi Pico: Build Electronics ..." is in section 005.13, which is programming. A completely different place. "Programming with STM32: Getting started with the nucleo board" is in 005.262, which is also programming. But why 005.262 rather than 005.13? It almost seems that whoever is classifying these books has no idea what they're doing ;)
I could go on at length about the confusions I have in trying to place my content. In the end, you have to make a somewhat arbitrary decision and just go with it.
Tables of contents work reasonably well within a book. Subjects are often non-intersecting, so they can be treated separately. For the most part, anyway.
A solution which is fairly reasonable is to index your site. Indices are useful because they allow you to take multiple views on something, thereby eliminating the taxonomy problem.
I'm not a great fan of tagging. It is too much of a scattergun approach to my liking. Perhaps some merit, though.
Then there's textual searching. In fact, that's how I relocated some of my notes. So, text search it is, then? Well, not quite. It seemed like a good system for my site which is focussed. It has problems scaling. I don't want millions of results, a la Google, I want a few relevant ones.
This is even a problem with search engines for the gemini and gopher protocols, where nobody is even trying to game the system. You often end up with a lot of similar stuff at the top which I am not interested in.
Oddly, for gemini, I prefer the "Collaborative Directory of Geminispace" over at gemini://cdg.thegonz.net/ , which is a taxonomy of categories, the very thing that I has doubts about.
So, in summary, it's not easy.