A tree-based file-system is optimised for doing a search from the users perspective, finding a file takes log(files) steps and finding related files is trivially cheap. It is likely hard to outdo that with a relational model.
A tree-based file-system is optimised for doing a search from the users perspective, finding a file takes log(files) steps and finding related files is trivially cheap. It is likely hard to outdo that with a relational model.
I would really enjoy a filesystem that's a loose tag-based hierarchy rather than a strict single folder tree.
In SQL terms, something like:
create table files (
file_id bigint primary key
, file_name text
)
create table tags (
tag_id bigint primary key
, tag_name text not null
, parent_tag bigint null
)
create table file_tags (
file_id bigint
, tag_id bigint
)
So tags are organized in hierarchies, but there can be multiple parallel hierarchies, and a file can belong simultaneously to multiple hierarchies. Say one is a flat list of tags by user, another is a flat list of tags by apps, and another represents replication or backup strategies, plus the usual directory organization.(I'm not sure if a file should be allowed to belong to multiple tags within the same hierarchy; my gut says no.)
Let's say that as a convention we separate the root tag in each hierarchy with a colon, and other tags with a slash. Your file 'cool_code.py' may be found under the tags 'storage:sda', 'users:roenxi', 'apps:pycharm', 'pycharm:projects/cool_app/src/utils', and 'rclone:gdrive/roenxi@gmail.com/202201'.
Gmail is a good example. Has very flexible tagging schema. Every tag exists in a hierarchy. Emails can have multiple tags.
SQL is not good for hierarchies though. You need to use recursive CTEs or denormalize relations and things.
Doing it more as a DB enables the OS to use the knowledge from RDBMS for efficiency, which I'm sure rivals the best file systems and it's possible to create multiple indexes and views for other use cases.
Our current view on file systems and the knowledge we have is heavily influenced by slow spinning disks, while RDBMS have leveraged RAM a lot more. With todays fast SSDs the file system operates in a reality that is more like RAM than a slow spinning disk.
Yeah, a file system stores general data so it is very easy to map it to an RDBMs that also stores general data.
But what this is identifying that once the relational nature of the data isn't a factor, the best lookup structure is a tree.
> Give me all image files > 100 KiB from 31. December 2021 to 1. January 2022
While in current file systems you need to scan the entire content of file metadata to get that information. It will take a long time, especially if you have a lot of small files (think Windows C: drive).
That might not be something the average user would do by themselves, but developers of, say, image processing apps certainly would.
But all that still isn't really leveraging the power of the relational model. That is simply indexing files on a lot of different attributes, ie, leveraging an RDBMS implementation detail where they use trees. The point of the relational model is relational algebra (SELECT, JOIN & WHERE in SQL terms). And WHERE isn't the interesting one out of those 3, it is JOIN.
If the use case for a relational filesystem is interesting filters then it sounds a lot like a false start.
I can see where you're coming from with the false start if looking at it from that isolated point of view. I see it more as the first step to storing data, in general, in an RDBMS and once applications start to utilize that, new use cases will start to emerge. Linux in particular with it's "everything is a file" philosophy seems to suited to use this model. A table for processes, files, network connections, ect.
A tree structured file system is effectively conflating an index with relations. Another way to look at it is that most FSes are like databases which only allow one index, and expose the index to users.
From an end-user perspective, files are documents that I create and I want to be able to find them in different ways.
From the perspective of a typical app developer, the filesystem is a hierarchical key-value store.
The perspective of a database developer, backup software developer, system administrator, etc. is going to be completely different yet.
I think we have really, fundamentally, failed to communicate here. I’m 100% sure we are talking about different things.
Windows did this for a short while, but I believe it was removed again. I don't know the exact reason.