A database built like an operating system: the architecture of FaunaDB
fauna.com
fauna.com
It is why they are both so complex to implement and so much fun to build.
MySQL & Postgres carry some of those same attributes. Even though they are user processes and the OS controls the lower levels, you would typically deploy them on their own servers, and give them all the memory & disk I/O you could, because they were their own operating systems essentially.
It's not clear to me how building a database "like an operating system" is something new. What I did see new in there was a consistency checking tool. That is super important. Because in a distributed environment you're bound to have drift.
Also I'd like to see how they managed to get full ACID Compliance. Fast network interfaces are great but there is still some latency which means hiccups when multiple nodes try to update the same row.
To answer your question about consistency, we have implemented a transaction resolution mechanism inspired by Calvin. The whitepaper contains a lot more detail, but the short answer is that we use a consistent, distributed transaction log to provide a global order of all read-write transactions, which are executed deterministically on the responsible data partitions. This allows FaunaDB to provide a guarantee of strict-serializability for read-write transactions. Read-only transactions are by default a bit more relaxed and provide a hard guarantee of serializability in order to avoid global coordination overhead.
https://en.wikipedia.org/wiki/Pick_operating_system
(actually, I don't know much about this new system to be commenting on it - but whenever I see something with keywords like "database operating system" - my mind instantly drifts to PICK, because that was the first system I used when I started my software development career)
I felt the same way when "NoSQL", document dbs, etc. came out. Ok, yeah, now we are distributed and massive scale with better tooling... but.. same/similar concepts.
Hate the arbitrary ascii-file nonsense. If the state of the operating system is in a proper ACID database it would be so nice. Tooling could make it just as easy as ascii files but under the hood it would be better.
OS/400 was cool.
But one does not get transactions. I would love a filesystem with transactional semantics. Surprisingly often, applications use some form of relational database not because their data is such a good fit for that model, but because they want/need transactions. In theory, one could roll their own transaction layer, after all, that's what RDBMSs do. But I have a hunch that getting transactions to work correctly without totally killing performance is not exactly trivial.
I think I remember the API being marked as deprecated, though. Or is my memory playing tricks on me again?
I think, if they had made transactions work with the existing file API–such as by storing the hTransaction in thread-local storage, with an API to get/set the transaction association of the current thread, and having the existing non-transactional APIs behave transactionally in that case–it might have saw more adoption by developers.
I'd much rather configure Apache by inserting fields in a database (and tooling can be developed to help) than editing ascii files in its custom configuration language.
You could configure the entire system with a bunch of queries. You could view the state of the system with queries so easily in a universal manner.
"select server_name from nginx.sites where enabled = true;"
This is fundamentally different from the model that we have today. Right now there's no good safe way of querying for the true state of things. You have to parse a bunch of crap and even then you wont know for sure if the value in the file is actually the present value or not unless the specific technology allows for that to be queried.Tons of different methods of including and overriding and so on. Knowing what's the "effective state" is difficult.
"select pid, owner, memory_usage from os.processes order by memory_usage desc limit 10;"
Those are just some basic examples and I know it's not that easy to pull off but I think it's certainly worth thinking about and challenging the popular Unix-isms of arbitrary information stuffed in arbitrary formats in arbitrary places held together by a thing string of conventions.It adds so much cost and room for bugs and errors because of the countless boundaries that it creates between each little component (constant parse, spit out write read) meaning and context and structure is lost in each step.
An operating system with a structured state backend could bring harmony to all the components.
Microsoft has really wanted something like that too. For example, WinFS.
Although it doesn't handle unbounded concurrency, there are techniques (like manual locking, or pgbouncer) to deal with this.
PostgreSQL’s process per user model makes way more sense but as you say still has upper limits on concurrency.
The paper is interesting, it leaves me with more questions than answers.
The biggest question is: -is this a new beast from the ground up or more efficient packaging around existing distributed computing libraries?
--For example when I hear a company say something like:
> based on log-structured merge trees (similar to Google Bigtable
I want to understand if they are just rebranding HBase for that part of the tool
Some open source is used internally, especially Netty, but our goal has always been to build a tightly coupled system for performance reasons. For example, the entire replication pipeline including the optimized Raft implementation is from scratch, as is the scheduler.
LSM trees perform well especially on SSDs, but we probably will migrate away from them eventually to something closer to LMDB in order to get a zero-copy read path.
Just like Windows 95! I don't think I've ever seen a database do this. Apparently it operates per-query or at least per application. Operating system seems right; this is a long paper.
Since our focus has been getting OLTP right, ETLs are currently handled by chunking record imports across multiple transactions. (We have internal tooling for this that we've used to help customers move their data into FaunaDB; we are planning on open-sourcing it soon.)
However, first-class support for OLAP use-cases is on our roadmap and something we are excited about providing, so expect more from us about this in the coming months.
Read benchmarks are coming soon.
I heard older IBM systems had something like this. BeOS had BeFS with inbuilt metadata index. NTFS + WinFS in user land vision was a bit like this. Office server aka MS Sharepoint (which implemented WinFS vision for intranet) stores all files in the database.
Most (non-Unix) mainframe and minicomputer operating systems – not just IBM's – have support for record-oriented files and indexed files in the filesystem. The functionality is roughly equivalent to that of non-relational "flat file" databases or key value stores. Examples include ISAM and VSAM under z/OS (aka MVS), and RMS under OpenVMS. This is in contrast to Unix and Windows, where the filesystem itself only supports files as unstructured bytestreams, and any record structure or index structure is imposed by higher levels such as applications or shared libraries.
You might also be thinking of the IBM minicomputer operating system OS/400 (nowadays called "IBM i"), which embeds a relational database (a variant of DB2) into the operating system. (I have often wondered how deep the integration actually is–I believe it is something deeper than just bundling a relational database with an OS, like how many Linux distributions include MySQL or Postgres, but I don't really know.)