Every normal user can understand hierarchy -- shoeboxes in closets in rooms in houses in neighborhoods in cities. The human mind orders things in hierarchies effortlessly. It's just incorrect to claim otherwise.
The only thing users sometimes can't understand are badly designed user interfaces for exposing and navigating a hierarchy.
Regular users don't have context for technology, they scoop up whatever companies offer on the market. If everyone gravitates towards the same crappy solutions, most users won't complain, because they can't or don't want to imagine how things could be made better.
I remember wondering at the time: do windows and mac not do enough to help users organize their files, or are the cute little My Documents and Pictures folders that obscure the underlying filesystem already too much?
There is a fundamental impedance mismatch between what hierarchies facilitate and how users think of their files. Users think of their files as having concepts associated with them. e.g. "this file is a photo, about my vacation, and is about scuba" or "this file is a book, about scuba"
The user may attempt to express the numerous concepts they associate with their files in terms of a hierarchy, but chances are good they will fail miserably because hierarchies simply aren't up for the task of organizing even moderately diverse data. Is 'vacation' a subset of 'photos', with 'scuba' being a subset of 'vacation'? Where then does their scuba book go? Is the 'scuba' sub-directory something they must create twice, once under 'vacation' and once under 'books'? Inevitably some concepts get duplicated sub-directories across the system, or the user settles on under-organization. This is because hierarchy isn't up to the task of organizing files in the way humans think about files.
The answer of course is some form of file tagging, so that 'scuba' can be a single distinct group that overlaps with both 'books' and 'vacation' and 'photos', but isn't a strict subset of any of those.
But that is separate from what file paths are for, which is essentially a semantically meaningful unique identifier. If you want to take a photo with the camera app, crop it with the image editor app, then attach it to an email with the email app, you shouldn't have to search through your photos again the second two times to re-find the one you're already working on. There ought to be something equivalent to copying the file path to the clipboard so that you can reference the file you're already working with in another app.
This is close to my core point. File paths and/or file hierarchies are not good for file organization. They have been pressed into this role for historic reasons and out of inertia.
Papering over the problems of file hierarchies by implementing a good copy-paste system is not a good solution to the fundamental problem. It alleviates some problems, but not all. You can conceivably copy-paste a file you've been working on in one app into another, but what if you want to retrieve a file you last interacted with a year ago? How will you find that file?
Under the file hierarchy scheme, in order to retrieve an old file the user must remember the idiosyncrasies of the hierarchy they created ages ago, with many concepts invariably being split into gods know how many subdirectories, OR nearly all of their files sitting in a small handful of directories, unorganized, with little more than the file name to help retrieve the desired file. These are both poor states of affairs. They require the user to store too much extraneous state about the system in their head in order to have any chance of efficiently finding the files they're looking for.
Modern systems have begun to improve on this state of affairs by implementing file searching capabilities that function more like internet search engines than traditional hierarchical directory explorers. In theory the user may be able to simply search "scuba" and the black box file search system may use a variety of methods for pulling up a list of files from numerous different subdirectories it believes to be scuba related. Perhaps it pulls in both /vacations/2017/scuba/neat-fish.jpg and /books/scuba/how-to-scuba.epub because both have the string "scuba" in their path. Or perhaps there is some exotic shit going on behind the scenes and an image classifier was previously used to determine the presence of scuba gear in /Pictures/1002389943-984983-198395.jpg This sort of file search engine is fundamentally an imperfect file tagging system. The tagging in this case is being done automatically in a variety of ways the user may or may not be aware of and understand, and may or may not be able to consciously influence and exploit.
Let's take a step back for a moment and define the problem:
In the mind of the user, there exists a many-to-many mapping between files and concepts. When a user is thinking of a particular file, they have in mind a set of traits about that file. (For example, the user might be thinking of "that picture I took of a strange fish while scuba diving on vacation." The concepts associated with that file might be picture, fish, scuba, vacation.)
How do we get that mapping out of the user's head and into the computer, such that it can be used by the computer to retrieve the files a user desires when thinking about a particular set of concepts?
That's not yet a question with a definite answer. Automatic tagging of files with concepts the user is anticipated to have should certainly play a role. Image classification is heavy, and therefore has privacy implications since that sort of thing is virtually always offloaded to some corporate server, but when it works well it's great. There are easy sources for automatic tags as well, such as extracting terms out of the file path, or extracting date ranges or even geolocation out of the file metadata. These are great, but typically prove insufficient on their own. Asking the user to manually tag their files with the concepts they associate each file with gets you a very high quality many-to-many mapping between files and concepts, but that is manually intensive. Notoriously so. One solution is to amortize it; during file creation ask the user to tag the file instead of naming the file. That's okay, but still not great. The user might be snapping 10 photos a second and not care to take the time to tag or even name each individual file. How do we meet the needs of that sort of user?
A solution I've been exploring is the creation of statistical models that assist the user in manual tagging, to blur the line between automatic and manual tagging. Here is what I mean by that: suppose the user imports 500 new jpg images, with metadata suggesting they were all taken in within a half hour of each other. The user tags ONE of those pictures with "vacation". What is the probability that the remaining pictures in that set are also vacation pictures? Pretty high. (Of course that's a trivial example, the user could just as well import the images all at once in a directory called 'vacation', but this goes back to my earlier point of extracting the many-to-many mapping that exists in the user's head. 'vacation' is merely one tag. By constructing the right statistical models, we can support an arbitrary number of tags.
With searchable tags, or whatever other thing you can come up with which is better than trying to remember a file path. But that doesn't mean you should get rid of file paths, because they're each good for different things and there is no reason you can't have both.
Here's another example. On a desktop you can open a static web page from a local .html file along with all the associated content. You'll have a directory containing an HTML file, maybe some associated .png, .gif, .webm or .js files and so on. A complicated one may have its own hierarchy because it's useful within the context of that document. Then that HTML file is going to have the names of the other files in it, including their relative paths if they're in subdirectories. The same thing can happen with InDesign documents or any other composite document like that.
Now how does the user find that HTML file a year later? Using some kind of tags or search function. But the way the HTML file then finds its associated pictures/scripts/videos is using something like filesystem paths. The alternative would be to combine them together into some kind of tarball, which is what some things do (like PDF), but then that's worse -- how do you edit a photo which is embedded in a PDF? Your average image editor will either not support it or will be inclined to do something horrible like rasterize the whole PDF first.
You're touching on a broader subject than file organization though, and that is humanist software design. Your static web page example is a good example. That's not humanist software design, it's design that serves the computer. When you save a web page you get an html file and a directory next to it that contains the images, css, etc of that webpage. The user is now burdened with keeping that html file and resource directory next to each other for the rest of the lifespan of that [meta-]file.
Consider instead the epub format. It too is one or more html files with images, css, etc. However unlike the traditional local static copy of a web page, all of those files are zipped together and stored as a single file (an epub is a glorified zip file with a special file extension. You can extract them with `unzip`.) This unburdens the user; they no longer have to worry about keeping those files next to each other when copying it, sending it to other people, etc. Epubs are humanist design (and not only for this reason, but I'll save that for another discussion.)
Of course inside that epub is a hierarchy of files. And that's fine. A hierarchy of files serves the ebook reader well. While strict hierarchies are poor for user interaction, they're great for software. The details of a file format need not be humanist because the user is not meant to pick it apart. Of course there may be perfectly valid and common use-cases for the user to pick it apart, to pull images out of it for instance. Software that facilitates pulling data out of documents, into a format that is generally useful to the user, is humanist. Consider your common desktop web browser for instance. It allows the user to easily pull an image out of a webpage by right clicking on it and saving it, or perhaps dragging the image out of the browser onto their desktop, or similar modes of interaction. The user needn't save the entire webpage, then poke through a directory filled with numerous files looking for the image they wanted. (At least, not unless the website and browser are conspiring to behave in a manner that is hostile to the user, for example by blocking right clicking....)
PDFs and PDF readers are shoddy in numerous ways of course, and strict hierarchical file management has done little to alleviate that.
But that's the issue. You can make it easier to use tags, you can make it so that they are what most people use most of the time, but you still want everything to be able to use paths, in a way which is still as universal and intuitive as possible, so that it still works properly everywhere when you do have a good reason to use them instead.
> Consider instead the epub format. It too is one or more html files with images, css, etc. However unlike the traditional local static copy of a web page, all of those files are zipped together and stored as a single file (an epub is a glorified zip file with a special file extension. You can extract them with `unzip`.)
Here's something even better -- what macOS does with applications. You look in /Applications and see a bunch of programs. Those are really directories with the application's files in them, whose names are the names you see but with a .app extension. If you click on one, you don't see what's in the directory, you cause the application to open. But you still can see what's in there, by using the actual path with anything that accepts file paths, and that is still useful. Because sometimes there is a .plist file in there or something else you want to modify, or you just want to see what's really there.
This is better than a zip file because it's still native. You get the behavior you want, but if you want to see or modify what's in there, you don't need to use a special program or go through the extra steps to unzip the file and zip it again, you just access it using the real path. Windows at one point did somewhat the opposite and exposed the contents of .zip files as folders, but that doesn't seem to work as well, e.g. modifications are not permitted or are inefficient or have other issues because you have to repack the whole file with every write.
It's a very compelling idea, but as I said, orthogonal to the matter of hierarchies or tagging. Say for instance you want to edit the .plist file of Blender.app. Where is Blender.app? Probably at /Applications/Blender.app but maybe it's in /Users/zrm/Applications/Blender.app To dive into that file/directory, you must first find it. How the user is meant to find it is the question that needs to be addressed. That's a simple example of course because there are only two likely locations for it (and the later is unlikely), the file/directory you're looking for has a name you remember off the top of your head, and because either of those directories probably has no more than a few hundred entries (probably a fraction of that.) In this case the user of a hierarchical system doesn't run into much trouble because the situation is virtually ideal, but even so a tagging system would have simplified the query. The further you get from ideal circumstances, the larger the gap between tag systems and hierarchical systems grows.
Additional thoughts on record oriented file systems: one example is Files-11 from VMS. An interesting property of Files-11 is that it allows files and directories to have multiple parent directories, which is similar to Unix's hard links except that directories, not only files, are allowed to have multiple parents. This gives the system the same expressive power of tags (consider each full path a file can be found at to be a tag on that file.) This allows all of your books to simultaneously be in /Books/ while still allowing you to have them in other locations like /School/cs/cs101/textbook.pdf However a system capable of expressing these relationships is the easy part. The important part is giving the human users a clean user experience, allowing them to take full advantage of the power of such a system without requiring them to manually create and maintain a web of nested links.
I do have other physical folders, just stored in three plastic bins in the closet. The bins are unlabeled, but I know what each one contains just by glancing at the contents and from there, knowing that my old computer science articles I want to keep folder is in the middle of this particular bin. Again, most of the folders are unlabeled/mislabeled but that doesn't matter to me because I know.
A filesystem can't quite replicate this spatial or looks feature of physical folders. There are also restrictions on what you can name a folder or file. For instance, I might want to name a file "Foo Inc. Invoice, Project: Snafu 3/19" but I can't---on Unix and Windows, the '/' is forbidden, on Classic Mac OS, the ':' is forbidden. Yes, I could place it in the "invoice" folder, or maybe the "projects" folder, or maybe the "foo" folder (or all three with hard/soft links) but now there's some cognitive overhead of just remembering where (unless I use links, but that's a weird thing with no real world counterpart, a physical folder can't exist in three different locations).
The spatial thing is important (but it could be just important for me). I've been using a virtual desktop on Linux for nearly 20 years now, in the same 3x3 configuration. The upper left corner is email; the middle right is the browser. Bottom left is a realtime view into system logs. The middle left and middle middle is where I do most of my development. The remaining spots are for overflow or temporary areas. When I got my Mac at work, I set it up with a similar spatial setup [1]. It helps to keep me oriented with what I'm doing.
And yes, I do know and understand the hierarchical file system, but that I still lose track of files from time to time ("I know I did this at one point---where did I put those files?"). I have 1.3 million files in nearly 200,000 folders in my home directory. It's amazing I can find anything at all [2][3].
[1] It's a linear series of 9 screens, but I still treat it as a 3x3 matrix of screens.
[2] A few weeks ago, I wanted to find the original image I used to make the favicon for my website about twenty years ago. I failed, and started a search with Google Image to see if I could find it (since I originally found it on the Internet). No go. A few days ago I found a folder I forgot about and lo', there was the original image.
[3] I also have 3,000 symlinks, three sockets, one named pipe, and 6 "others" (whatever those are).
That hierarchies can be expressed by tags is easy to demonstrate. Simply tag each file with all of it's parent directories up to root. Implementing hierarchies in terms of tags is not as computationally efficient, so you wouldn't actually write a filesystem in that way. The point is to create a new interface for humans, not for software. When a human has a dataset that is genuinely best organized as a hierarchy, they can express that hierarchy in terms of tags.
Of course traditional file hierarchies implemented in the traditional way, not in terms of tags, must still exist for some purposes. Source code, which requires a deterministic and efficient way to refer to specific other files, is one example. (However the user could still benefit greatly from augmenting the source code tree with tags.)
What I'm suggesting is not that hierarchical file systems be removed from operating systems, but rather that they be removed from nearly all user interaction with computers. The "Save as" dialog should not be exposing the hierarchical file system, at least not by default. Ditto the "Open" dialog. Hierarchical file managers like Finder, Explorer, Dolphin, Nautilus, etc should all be revised to be foremost file tagging systems designed to aid the user in managing and querying their tags. The "Desktop" should not be a directory on the filesystem, but rather a tag (or even a complex dynamic query.)
Why would someone think of files like that? I don't organize physical objects that way. My books are on my bookshelf and my pictures are in albums or frames. Why would a computer be different. I wouldn't store a book about scuba diving with or relate it to my vacation pictures of me scuba diving in any way in real life, why would I do this on a computer?
If I had a collection of books on my computer and pictures from my vacations it would be stored the same way I would do it in real life, keep the books together organized however and keep the pictures together organized however. The topics of my vacation pictures and my books have really nothing to do with eachother, even if they happen to be related.
What you describe sounds like the very opposite of how i'd want my files organized.
You don't put every single physical item you have relating to scuba into one big bucket because in the physical world, putting something into one bucket precludes it from being placed into another bucket. When you are organizing things you have to anticipate your most likely queries in advanced. So maybe you have one shelf for science fiction, and another for marine wildlife, and another for all your programming language books. For a while, this works well enough. One book can't go on two shelves, but with few enough books that's not a deal breaker.
However, when your personal library grows large enough, this system rapidly becomes impractical. At this point, the librarian resorts to creating a card catalogue, categorizing books by author, subject, etc. Books which have numerous topics have numerous cards in the card catalogue.
Since this is the 21st century, card catalogues are now digitized. The digitization of card catalogues saves space, but more importantly it facilitates more powerful queries. You can grab two different stacks of cards, relating to two separate topics, and instantly find the intersection between them. That is, the list of books which appear in both stacks. This greatly reduces the number of cards you must look through to find a book that you're quite certain is in both of the stacks. For instance each stack might have a thousand cards, but the intersection of those two stacks might only be 50 cards. That's 50 cards you have to search through to find your book, instead of a thousand.
A file tagging system is a digital card catalogue for your files. And as librarian of your library, you can decide which subjects warrant their own tags and which don't. What you don't have to do is decide ahead of time what sort of queries you want to facilitate. And it's no longer important to keep your collection of books and photographs separated. If you want a book, you intersect whatever you're looking for with the card stack of all books. If you want a photograph, you intersect it with the card stack of all photographs. If you have a book of photographs, it could even be in both, should you want it to be.
When i look for things on my computer, i don't have a vague concept of an idea of a thing i'm looking for, i look for specific things for specific purposes. Why would i need everything on my computer related to cats showing up if i'm looking for say, a cat sprite in a game i'm working on vs a picture of my cat or say a cat meme? I have no need for the other two things if i'm looking for one of those things. Yet in your system all those things and more would show up should i search for cat. That's pointless. I'd rather navigate to my programming projects and assets folder for my cat sprite, and go to my pictures folder for my own cat or meme cats. There's not many situations where i'd ever need all three at once.
For a library or massive database, sure, a file system may be insufficient. But as system for organizing things i need and use for specific purposes, it's not ideal.
What you're describing is that this sort of system excels at. By describing the properties of the specific file you're looking for, the tagging system can locate it very rapidly. Without a tagging system, knowing the file you're looking for is like knowing the bookshelf your book is on. If the book shelf is small, you can look at each book in sequence and find the one you're looking for. When your bookshelf is very long, it pays to reduce the set of files you're looking for.
>"Why would i need everything on my computer related to cats showing up if i'm looking for say, a cat sprite in a game i'm working on vs a picture of my cat or say a cat meme? I have no need for the other two things if i'm looking for one of those things."
If you have a single cat sprite, and search the intersection of 'cat' and 'sprite', then that single cat sprite would be returned to you.
Consider instead you have in mind a particular photograph of a cat and a dog, on a car. Perhaps you know you put that image in your cats directory, which has a thousand images because you're an avid fan of cats. To find that particular cat image, you'd need to search through a thousand images sequentially, unless you happen to remember the precise filename of that single image out of a thousand. Thumbnails will help, but that search is still going to take way longer than it should. If your collection of cat images is properly tagged, then you can instantly get a list of the intersection between cat, dog, and car. That list is going to be much shorter than 1000 images. It will take you seconds at most to glance through it and find the precise image you were looking for.
With terabyte harddrives being dirt cheap commodities, there is no real reason the average person couldn't or shouldn't have hundreds of thousands to millions of files at least. The reason most people don't is because with the organizational software currently available to them, (e.g. windows explorer) there is no feasible way for them to keep that many files organized. There is no good reason that should be the case.