Skipping the boring parts of building a database using FoundationDB
blog.tigrisdata.com
blog.tigrisdata.com
The nice thing about FDB is that after 3 plus nodes, you can simply add nodes using your cloud provider of choice and it scales pretty nicely while still giving your high availability and fault tolerance.
It's pretty funny to me though to see this - I've been spending a few days building a simple database on top of FDB that supports indexes, secondary indexes and schema migrations backed by json-schema (very, very similar to this, totally independently!)
To get into a little bit, it's not super difficult if you use FDB. FDB is a very bare key value store. It's incredibly low level. You don't even get a notion of collections. You have to implement everything yourself. what it does give you, however, is a giant hash map that will guarantee that items are in sorted order.
so to build what I was describing it's easy:
a collection can be a tuple to map:
(your-app, your-collection, _id, your-model-id-number) => json
e.g. (hn-app, users, _id, 1) => { _id: 1, username: endisneigh }
(hn-app, users, _id, 2) => { _id: 1, username: reader }
an index can be something like: (your-app, your-collection, your-field-to-index, index-value, _id, your-model-id-number) => json
e.g. (hn-app, users, username, endisneigh _id, 1) => { _id: 1, username: endisneigh }
(hn-app, users, username, reader, _id, 1) => { _id: 1, username: reader}
Because FDB gives you transactions, you maintain the index by populating the keys according to the pattern above on your create* and update* operations.To do something like a schema migration, FDB gives you a get_range operation that you can use to find all keys that have a prefix. So what you'd do is store a value indicating that you're doing a migration into the database, iterate through the keys in a batch (so it's all a single transaction), update the value in the db saying what the last key you've migrated is, and continue until you've done all of the keys.
A lot of stuff is pretty trivial once you assume the underlying semantics are solved. I've seen some interesting projects involving things like using FDB as a virtual file system for SQLite, but the problem with that is FDBs primitives are actually flexible, and so there are optimizations you can make if you built it using those primitives from the beginning, as opposed to using FDB as simple a key value store without taking advantage of the transactions.
-------
On another note, one idea I've had (feel free to steal) is to reimplement IndexedDB using FoundationDB. IndexedDB is also a key value store which supports transactions, like FDB. Obviously IDB is not networked.
The idea is that if you can semantically map IDB with FDB, then you could use FDB as a store for IDB (scoped to the user, of course). And then any app that uses IDB for its storage (like an offline app) could use FDB as the backing without having to use a different set of data structures to actually represent the storage.
However, if you are an application developer looking for a ready-made solution that you can plug-in as your application's backend, then FDB does require heavy lifting. For example, you would have to implement auth mechanism, query layer, schema management, and indexing. This is where Tigris comes into play.
You have a very interesting idea about backing IndexedDB APIs with FDB.
Once you get all of your core functionality completed, you should definitely look at the IndexedDB APIs with FDB. I see you're considering FDB as a service. You could definitely compete with Firebase if you had some admin primitives around ACLs and you reimplemented the IDB APIs with FDB.
For instance, you and I both are on two computers obviously. We could each have a Tigris instance. If your app is down, we fallback to the regular IDB api and everything is saved. You could save entire transactions that aren't persisted to FDB and replay them when FDB comes back up.
More interestingly, as the admin, you could use all of the IDB tooling like LevelDB, PouchDB, absurdsql, etc and only concern yourself about the user (you and I) and things like how many keys they can save on the free plan, premium, etc.
Assuming you have declared your schema as shown here https://docs.tigrisdata.com/typescript/getting-started You can evolve it by updating your type definitions, deploy the new version of application. Once `createOrUpdateCollection` is called, it will update the schema.
--
The IDB idea sounds very cool, let me dig into it more.
Unlike Postgres or RDBMS, being a NoSQL store FDB has some advantages with long running things like migrations. In particular, FDB can store things in an arbitrary manner.
I created a notion of a "preempt" for migrations. The way a preempt works is that when you define a migration, you also define a preempt which represents how the old value changes to the new value after the migration.
For example, if you have:
{ username: endisneigh, _id: 1 }
and you want: { username: endisneigh, _id: 1, lengthOfUsername: 10 }
You'd obviously run some code to modify everything. Lot's of ways to do this, map reduce, batch job, etc. The problem is, if you happened to have 100 million of these rows, it will take you a long time to modify all of them. There are a lot of ways to solve this - locking being a popular one.I created a notion of a preempt so you can define the change in the migration and immediately have access to the change if you access the particular record prior to the migration job getting to it.
So in the above example, you could have a migration that looks like the following:
class Migration {
@up
function migrate(oldRecord) {
oldRecord.lengthOfUsername = oldRecord.username.length;
return oldRecord
}
@down
function migrate(currentRecord) {
delete currentRecord.lengthOfUsername;
return currentRecord;
}
}
What's nice about this is that if you use "preempts", you don't have to have any conditions around the long running jobs in your application code. You can treat the long job as already being completed as soon as you run it, regardless of the amount of records. You can call it a just-in-time migration for new records, to be run as you access records. The reason I felt this to be necessary is to maintain the transactional (completed or not completed) semantics FDB gives you because it made the code easier to work with if you can assume things are done, or not. Eventual consistency is a huge pain and creates too many bugs imho. The other reason I like preempts with FDB is because it's literally something you can't do with RDBMS (you couldn't treat a column as another type until the transaction has actually completed for a alter table, for instance).I would also not get too invested in your architecture such that you cannot change data types during schema migrations. I'd generalize it so it's always possible. Like if you have an integer field, and an index on it and you use $gte, it does what you expect. If you change it to a string, $gte still works, and uses the lexicographic ordering instead of the number ordering. You can imagine the equivalent for all of the operators.
Only caveat is that you'd either need to ensure all preempt code is idempotent since your long running job might run the code twice (once just in time, and another during the migration), or you'd need to save which records have already been processed via the just-in-time migration and skip those as necessary. This leads to issues since you would need more storage space, and then you'd have to clean it up, and if you have a full disk that leads to more issues, etc.
But supporting data type changes without rebuilding is not ideal. It will lead to data quality issues and complexity on the application. Integer -> String example is simple. But what about String -> integer, how are the consumers of data supposed to handle the situation where the field in some records has a string value and in some has an integer value? They will have to add type checking which complicates each of these consumers.
Then some of the downstream consumers such as data warehouses depend on strict data validations. We went through this problem at Uber- I blogged about it here https://www.uber.com/blog/dbevents-ingestion-framework/
Solving the exact problem you stipulate.
I did sth pretty similar last month: https://rxdb.info/rx-storage-foundationdb.html
It supports indexes, mongoDB queries etc. to store and query JSON documents via RxDB on top of FoundationDB.
You write "giant hash map that will guarantee that items are in sorted order". Where can I get more info on that?
You can see some more details about it here: https://apple.github.io/foundationdb/data-modeling.html
This is worth correcting at source as well as down the comment tree. FoundationDB was 100% closed source until 2018, when it was open-sourced by Apple.
For anyone else who's confused, FoundationDB went closed-source in 2015 but went open-source (Apache 2.0) again in 2018.
Apart from the size of the document, there is no limit on the size of the array or the depth of nested data. We plan on substantially increasing the document size limit.
I suppose this is due to the FDB limitations, so this obviously isn't a blobstore nor will it ever be (?). For example, we need to store video and image files which are easily 100KB - 1GB in size. Tigris or FDB are great to store metadata in (and the metadata is just as important to us as anything) but the blob storage is a bit of a problem. Would be interesting to integrate something with, for example, MinIO or S3.
The native Javascript SDK is on our roadmap. In the meantime if you would like you can generate Javascript client using openapi generator such as https://blog.logrocket.com/generating-integrating-openapi-se...
We would be happy to accept your contribution.
This is one of the most confusing aspects of the modern data infrastructure industry, why does every new system have to completely rebuild (not even reinvent!) the wheel? Vendors are spending so much time rebuilding existing solutions, they end up not solving the actual end users’ problems, although ostensibly that's why they decided to create a new data platform in the first place!
In this post we talk about our approach to building Tigris - the open source developer data platform. We talk about why we chose to build on top of FoundationDB, one of the most reliable distributed KV store with an amazing correctness story. We also go into detail about our experience using it.
Good luck with your project, FDB fucked it's users back in 2015 when it abruptly closed shop and went closed source. Hopefully some good can come of it yet.
This is simply untrue - it was not open source prior to acquisition either. The first point at which FDB was open-source was 2018.
Gotta get that Apple money.