Json-Base – Database built as JSON files
github.com
github.com
The first problem he encountered was that multiple connections couldn't both be using the database at a time without clobbering each other. "No problem," he thought, this is a good use case for micro services. A service sitting on top would ensure that there was only one operation being performed at a time.
Next, his problem was that the database would get corrupt sometimes when something bad happened in the middle of writing the file. His solution was to put the entire JSON format inside of a JSON string. If it could be parsed successfully, then he knew the whole file was written. Then all he needed were "backup" files for each table, in case the current one was corrupt.
Next, his problem was that querying and iterating through a large table performed badly, since it required parsing the entire thing first. Querying several times required the whole file to be parsed every time. The solution was to move SOME of the tables over to JSON-inside-SQLite.
EDIT: Oh yeah, the next problem was how to structure the data inside of sqlite. He decided to make a single table called "kitchen_sink" that held every JSON value. There was a column that said which "collection" it belonged to. There was another column that represented the row's primary key. So you could quickly query for a collection name, and a primary key, and get the full JSON row.
So the next problem was that you couldn't query quickly for things that weren't the primary key. So new columns had to be added called "opt_key1" and "opt_key2" where certain rows could put key values, and indexes could be added on those columns, so you could quickly query by it's first optional key, or it's second optional key.
Just use SQLite, people. Even JSON-in-SQLite is still likely to be an improvement.
You can have structured relational databases (RethinkDB for one, but there are others)
But the real answer is that our team was very siloed. No one knew what anyone else was doing. The other problem was that he was actually solving real world problems, and he was a very high performer. He got stuff done. Arguing to start over a project that's already working is a difficult position to hold when talking to management.
Really? At the expense of everyone else who has to deal with this monstrosity for the foreseeable future, or worse yet replace it with an actual tool that can be reliably used.
This JSON-inside-sqlite-inside-JSON-inside-a-JSON-string beast should never have seen the light of day.
You're not paid to be entertained, sorry. You're paid to be productive. As productive as you can, and to put the needs of the client and the long-term success of the company hopefully first but certainly before any resemblance of entertainment if you're getting paid
Did I mention you're getting paid to work?
This sort of stuff is what deters me from being a developer sometimes. Fuck the salary, get me out of here.
If it's any consolation, it fell on to me to maintain this code after he moved on to something else, which is why I know so much about how it works.
That's not any consolation... If anything, it's all the more reason for you to be pissed off. He should have dropped the project, rolled out a future-proof tool and taught to do differently next time.
Anything short of that is just enabling the dude's delusion of grandeur and therefore a mistake on everyone else's part...
This has happened to be before. I disagreed with a technical direction, it was implemented anyways, and then I'm left to maintain it. Very frustrating.
Just like the never ending "turn this Excel workbook into an app" stream of work, refactoring older apps will be a constant. Focusing today's conversations on yesterday's mistakes only detracts from the work left to do (which is to say if your architecture change arguments are valid, there should be ways to justify implementing them today outside of "it should've been done this way in the first place because then we wouldn't have had those problems that are now solved anyway")
I've dealt with this before, it is a form of gas-lighting. Some people are good at making everyone else into a bully when THEY are the actual bully. Like the kid who keeps splashing you in a pool, but runs off and cries and tells when you splash them back.
Standing up for yourself sometimes makes you look/feel like an asshole. That doesn't mean you are wrong or that you shouldn't do it.
edit: Check out the book "Radical Candor" if you regularly struggle with expressing negative feedback
< "Well all the work was already developed, and it would take too much time to rewrite it. You should have said something earlier" > "When?" I asked, considering she had just put up the (big) PR's and PR's ARE the time to review... < "Check her commits as she pushes them to the repo" - as in her bugfix/feature branches, not master...
My jaw dropped. Especially since I was hired on as "Lead" and had all the accountability but no actual power.
It's incredibly frustrating because during code reviews I will request changes so it's not such a broken hack job, and the response will basically be "No, it's not worth changing". At which point I'm the one "holding up development". We wasted hundreds of development hours during the last project because of this persons "inventive" code, and nobody seems to understand what's going on.
Shame the job market is a bit crap right now.
Two pieces of actual useful advice I can offer are:
1. A review style I picked up based off of RFC 2119[1] basically the reviewing software we use allows us to mark particular comments as blocking of non-blocking and I pair that with the usage of MAY/SHOULD/MUST within the comment language i.e. "We're using the old `array()` syntax here instead of `[]` we MAY wish to use the more modern syntax" this allows me some room to elevate necessary change while keeping in the nitpicks I really want to throw in (and I do try and minimize them) without lowering the power of the strong comments. I've used MUST maybe three times always for something incredibly terrible like pages not loading or migrations to the DB that are unsafe and cause data loss.
2. Agree on syntax and style rules and enforce them. It's easier to get people to agree to rules once than try and argue for them on each PR - anything like brace placement or line limit shouldn't come up repeatedly since it wastes everyone time and makes folks feel belittled.
Post-mortems are great for many reasons. For the case of GP, one particular advantage is that they align senior peoples' understanding: we shouldn't do X again. If you have a strong narrative for why a project failed, post-mortems are a formal setting in which you can present this narrative with concrete evidence to higher-ups.
In the future, when you see warning signs that a mistake is approaching repetition, you can raise the concern up the chain, invoking the memory of the post-mortem to motivate their intervention.
I also totally agree that a sincere and high-quality code review process is required for high quality code. Your 2119 recommendation is excellent. I'd also recommend doing some reading on commit message templates that smart people follow, they've improved my commit game, big-time.
We've had that process on for quite a while, and while there are some big weaknesses and holes in it we've also adopted a principle to keep PRs as small as possible[1] with those two tools we've had some pretty reasonable success with a lot of our biggest incidents being related to times when we've made large changes or a review was skimped on.
1. Even if that isn't measure in LoC - moving a dependency and updating references to it is something I'd count as a single action - but one I'd want isolated from any logic changes.
The commits I was told to review if I wanted to stop thecraziness were the personal ones going to the bugfix/feature branch.
Many good companies enforce a no-origin-branches policy, with rare and well-justified exceptions. Because, used as you describe, a "feature branch" is just a future massive diff in disguise (when it's eventually merged), and massive diffs are a big no-no because they're a huge pain to iterate on via code review.
The workflow looks something like this.
git pull master; git checkout -b my-feature; ... ; git add -A; git commit
At this point, you submit the code for review, and upon approval the branch is merged into master and pushed. It’s not possible to push a commit hash to master that has not been reviewed.
If you have a feature that’s composed of many steps, you can “stack” multiple commits, and review/merge them in order.
If you want to develop the entire stack at once, you’re most likely doing something wrong (according to this culture). You can incrementally merge pieces of code to master in such a way that’s impossible for it to be deployed, and your final diff can be what makes it deployable.
Encouraging smaller changes isn't nearly as useful if those changes aren't isolated - if it's just half the picture then you can't accurately review it.
I hit a similar sort of issue recently - I've been incrementally developing a complex data migration, each change to the migration has worked on its own and been reviewed separately but I'm still going to go in and request a full review once the piece of logic is fully assembled. This is also happening on an integration branch on origin - we do try and keep these to a minimum but we're making a backwards incompatible change that would be quite expensive to do in a fully backwards compatible manner.
There are things that are infeasible to reasonably do without an integration branch (nothing is impossible technically, but it might be a huge waste of time) but even those things are pretty few and far between. If integration branches are common place at your company it might be good to examine coding practices and see if you can slice up tickets to be smaller.
But, I do believe it pays off in the form of a higher quality end-product (fewer bugs, more testable/legible components, more extensible), which saves you time in the long run.
I'm pretty sure the cascading series of "his next problem" sentences implies that there were plenty of problems with the architecture that weren't identified ahead of time, and they had to encounter and then fix as a series of bugs.
> And still, there's a lot to be said for keeping your developer's entertained so they stick around.
There's a difference between keeping your developers entertained and letting them infect production with ill-conceived projects that cause problems for all those that interact with them.
This project is reimplmenting something already solved multiple times. There are many document stores, and JSON interfaces and addon to traditional RDBMS, so what was being solved here, other than letting someone scratch an itch at the expense of the division he's working in. You're better off giving him 20% time for his own projects and calling it a day if you really think entertaining your developers is important enough to warrant it.
There are times when rolling your own is useful. Generally when there's some extreme requirements for space or performance, but even that becomes rare when the area is mature and explored thoroughly. A database, even a JSON document store of some sort, is so mature that to make it worth while for one person to roll their own when it seems to need all the common features (locking, remote access, different clients), that to actually recoup the cost of building our own (much less the future cost of troubleshooting and bug fixing) is almost impossible unless you're somehow hired a genius workaholic for peanuts.
Once you've abstacted it to a service, your API is what you and your client (should) care about, and many of the arguments for more specialized implementations no longer apply. Personally, I think the only reason I would go with something like sqlite instead of Postgres/Mysql behind a microservice is if I was baking the date into it with each release, so the sqlite data files are shipped with the version released. Even then, I'm not sure there's any reason I would do anything other than sqlite though. Even if I had need of lots of JSON files, I would probably have my build procedure process them into an sqlite file I tested and shipped with, if only because I would then avoid having to deal with all the problems this guy encountered by trying to make his own database.
The GateKeeper process isn't something you want to index on too heavily - but you also need a mechanism to counter-balance the possibility of a dev saying "I built a prototype last week that does 95% of the things we want" and 3 months of iteration later identifying that it only did 5%, and that getting the remaining use cases will require a re-write.
The problems he encountered with his dumbass solution were EASILY foreseen by an even noob coder. What did he "get done"? How did writing his own shitty version of a database add value to the company? He is good at finishing his own pointless tasks quickly, maybe, but if I was in charge of the team he would be looking for a new job after this stunt.
Sick-to-death of these cowboys. Nothing is ever "done", the majority of expense in software development comes in during maintenance, not during initial implementation.
As OP said they were siloed from each other and he definitely needed some mentorship. I've seen devs like that turn to incredible coders just after a couple of months of pair programming.
That is 100% a pet peeve of mine: Places that hire perfectly capable jr. engineers and then fail to give them the support they need.
Given the team dynamics and lack of involvement from this person's manager, I wouldn't move to fire them. I'd move to rethink the entire team, admonish the manager, and possibly remove them. The team itself wasn't working, and this was a symptom: someone had a bad idea, pursued it for too long, nobody did enough to stop it, and then they couldn't go back.
This is a classic consequence of a manager who has stopped paying attention to their own team. The team was most likely also overburdened with too many tasks, which is why everyone was working on something separate and independent and nobody knew what anyone else was doing. In reality this developer shouldn't have been given a project like this without being paired with a more senior engineer to supervise it, but that would cut down on the number of story points the team could get through and would thus be discouraged in a dysfunctional environment.
Depends, I’ve known people who have gone through similar experiences and still poo-poo all those “unnecessarily bloated” solutions like a proper database.
we ended up with a lot of shitty solutions to problems that were hard to maintain, hard to extend, and hard to use because the more "complex" solution was really just a fancy version of a folder and some text files.
So I built my own database in PHP.
Enough said
Like, initially allocating a large file block, and then you subdivided that yourself to get individual sector access?
It's all fun and games until you realize DynamoDB works more or less the same way: https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
My question is, why didn't he go for any of the existing solutions when setting them up would've still been faster than rolling his own DB-in-a-JSON-file solution?
I've been building an open-source alternative on mobile that based on similar concept (SQLite + FlatBuffers): https://dflat.io/ SQLite own schema is already awesome, but in this way, you can have sum-types, better schema upgrade guarantees, index building can be asynchronously etc.
I'm not sure i understand how this can happen... unless you try to update JSON in-place (which is a very bad idea for any text-based format), what you do is encode/write the entire JSON from scratch. So either the file is written properly or it isn't written.
Honestly from the entire message it doesn't sound like JSON was a bad idea but that your coworker didn't know what he was doing and if he was doing something else then he'd still be doing big mistakes.
2. Lightning struck the 12 V feed and upped the voltage to 10 MV, turning all 0 and 1’s into 6’s
3. Someone spilled a New England Pale Ale on the server
4. The process was assinated by the mysterious killer only known from his modus operandi of leaving OOM written in blood across the syslog
5. Birds nested within the server and fed all the SATA cables to their babies
Seriously though, disk writes aren’t atomic.
1. Write your updates to a copy of the file.
2. Do an atomic rename of that copy to the original.
The article linked above never explains that part, it only assumes that it will happen. From the code it sounds as if the crash can happen in the OS itself (but then the entire kernel will crash). At that point things are completely outside your control and you might as well running on broken hardware.
His sources seem to disagree, certainly with that kind of blanket statement. For example:
Our study takes a pessimistic view of file-system behavior; for example, we even consider the case where renames are not atomic on a system crash.
So this is clearly considered an outlier/unusual.
(e.g., a single 512-byte write or file rename operation are guaranteed to be atomic by many current file systems when running on a hard-disk drive)
[https://www.usenix.org/system/files/conference/osdi14/osdi14...]
I remember reading quite a bit about the (performance reducing) lengths filesystems go to in order to ensure consistency of directory entries even in case of a crash, and for example how "soft updates" were introduced to accomplish the same consistency with less of a performance degradation.
Looking at it from another angle, if you are running on top of a filesystem that cannot keep itself consistent, then you are SOL, there really isn't anything you can do to mitigate.
Just like we can't guarantee that we will be able to persist data that's in memory to disk if the OS is free to kill us at any time. "Best effort" it is, which means getting the data to disk as quickly as possible and not corrupting what is there.
1a. fsync() the file to ensure the contents are durable.
2. Do an atomic rename of that copy to the original.
2a. fsync() the directory to ensure the rename is durable.
> if there's a crash during the write
It never brings up how can there be a crash in the first place (also the entire article is too filesystem specific).
It's impossible to make crash-free systems. If your goal is to maximize the chance that your data remains valid, you have to plan on that.
I imagine if this database system is contained well enough, it shouldn't be so difficult to swap its internals with something else. Especially if it's all just JSON-like.
https://github.com/skorokithakis/goatfish/
It's actually quite good as a quick-and-dirty datastore. I do need to move everything to SQLite's JSON field, though.
Had he pushed further.....
Had he pushed further, he'd have raised funding for the newly invented NoSQL DB, and built a startup company on top of it.
The is the mentality that plagues the industry, that anything more than a few years old is obsolete, and therefore experience is worthless, and therefore the wheel must be reinvented every time because those old programmers must have been dumb, why would they use SQL otherwise. Why real engineers don't take "software engineers" very seriously (and in turn why software engineers don't take webdevs seriously).
Only downside is you need to format ext4 with type small otherwise you run out of inodes before you run out of disk!
https://www.cs.ait.ac.th/~on/O/oreilly/perl/cookbook/ch07_09...
Worked like a charm, never had a problem with it.
It was actually put in as a placeholder until we had time to think about a real storage solution, but it turned out we never needed anything more sophisticated, and were actually the fastest and most reliable clients we had. In fact, every time I encountered a performance problem I was hopeful that I would finally have a good reason to do that real implementation, but it invariably turned out to be a simple bug.
- Cocoa has -writeToFile:atomically:, which writes a new file and then renames, so no write-corruption
- We were lucky that lists had just the right granularity for a single file to be read/written atomically
- We likely wrote (quite) a bit more data than absolutely necessary, but I/O tends to have large fixed overheads so medium files tend to take around the same time as small files
- We did not do anything with the data on disk except read it, so not a DB
- We really did use files, not JSON strings inside SQLite
- We flushed to disk asynchronously, but as quickly as possible
Looking at the code there's:
- race conditions everywhere.
- bad and inconsistent formatting, which doesn't help with the
- huge if-else monstrosities.
- Also uses synchronous IO and asynchronous IO randomly.
- Uses try-catch liberally, doesn't check the caught errors, and just re-tries blindly forever in some cases.
If you do any parallel updates/inserts/removals with this "database" you're pretty much guaranteed to lose data. Updates are essentially: 1. read table, 2. make changes, 3. save table. Which at least would work if it was all synchronous.
I know this is going to sound harsh, but building databases is hard for even the most experienced coders, and whoever wrote this is clearly at the other end of that spectrum.
https://github.com/Devs-Garden/jsonbase/blob/master/tables.j...
Or maybe there a different locking mechanism in place that my cursory look missed?
I volunteered to write a medical visit recording app for an NGO in a developing country (a friend works with the NGO and asked me if I would help), and they have almost no budget, no guarantees of internet connectivity when their folks are in the field, and the likelihood that they may be using this software for years.
So I wrote a C# app that uses Winforms, and stores all data as JSON files, the 'table' structure is basically directories in the file system.
It lets them share visit file by import/exporting a zip file of the JSON via sneakernet USB drives [super naive last record written wins], does not rely on an internet connection anywhere at all ever, and all files are stored in plain JSON so that they can conceivably in the future do some data analysis on it. Their alternate plan was to continue using paper, or some terrible regular reconciliation of excel spreadsheets.
Having said all that and defending my decision on this single use-basically-I-wanted-to-have-independent-JSON-instead-of-SQLite-so-in-the-future-maybe-have-a-web-function-to-sync solution,
This feels like a different use case.
sounds like you should use CouchDB, it a database/webserver, so you could make a simple html form on localhost[0], CouchDB is built with replication/sync (over HTTP) as one of it's main feature[1], and on the field, an offline-first webapp with PouchDB[2] and Service Workers[3] could have the exact same form
[0] https://docs.couchdb.org/en/stable/best-practices/forms.html
Because then someone has to run and manage a webserver, and there is no guarantee that ServiceWorkers will work like they do in 5 years or on an ancient Windows7 laptop running IE7.
I want this to be able to run for years without my intervention. :)
couchdb
there's nothing to manage.
On Windows, you install couchdb.msi or whatever, installed as a windows service, it automatically boots at startup time.
Start IE7 go to localhost:5984/_utils, you get the DB's UI. At that point, all you did was installation.
One click later, you created the first db called 'somedb', a click later, you created the first json doc called 'somedoc'. Now you can access it from localhost:5984/somedb/somdoc.
For the HTML form, just after you created 'somedoc', you can click on "add attachment" and upload someform.html, then go to localhost:5984/somedb/somdoc/someform.html from IE7, no need for anything fancy. After you're gone, someone with the most basic HTML knowledge can make some changes if need be. No Internet required. Will work as long as the laptop works.
Basically, you don't have a bad idea, and if I were a couch expert or were not on the other side of the world, I might have chosen that. But since I know C#/WinForms well enough, and if we went with the browser I would have to support mobile phones and I don't want to support mobile phones for this use case for a lot of other reasons.
I also suggested Couch because outside of western countries, you seldom find laptops or desktops (outside of cities), but smartphones with a recent browser are ubiquitous, so if it worked on IE7 it would run anywhere, even in the most remote area with no/crappy network. And the first time I used couch, my programming "knowledge" was very basic HTML (no JS).
> if we went with the browser I would have to support mobile phones and I don't want to support mobile phones for this use case for a lot of other reasons.
Yeah... all in all I completely misinterpreted the requirement of your use-case.
That probably explains a lot: you have a hammer and everything looks like a nail.
CouchDB makes no sense whatsoever for the requirements described.
Actually, that comment made me see things differently on a project that had me scratch my head for the last couple of weeks, so thanks a lot!
Don't know if it's related but I tend to feature creep.
If you want the long version, or you are interested in ways you could also be involved in that kind of project (or know C# and WinForms and want to help??? :) ) my email is my hn username at gmail.
For a community that loves it some Jepsen analysis, I can't for the life of me figure out why this has been up-voted so many times. This is just saving JSON file to disk. I'd argue this is harder than using Redis (flushing to disk) or (vomits in mouth) Mongo. Or shit, just use your filesystem and `jq`, you'll have something likely faster, safer, and more maintainable.
Just uncomment the tests you want to run! Easier than using a testing library IMO.
And just because you're not using a library doesn't mean you shouldn't have assertions. All these "tests" do is log. What do I check the output for?
All you need to do to have a somewhat respectable build is uncomment those tests, make them clean up after themselves, change the console logging to be assertions instead, and make them run on GitHub.
Actually, another view is that there's nothing wrong with tinkering and DIY. Perl, JS, Redis all came from people hacking their own solutions (as far as I know).
Also, many big software orgs build extensive internal tools themselves.
Plus, making your own stuff is a lot of fun. You should try it sometime (if you haven't already) :)
What is bad is putting them in production when you don't have a clue about the domain.
Also, they have to have some clue about the domain, because the domain is their own problem and they're writing a solution for it. So I don't think we can really just someone as not having any clue about their own engineering challenges.... especially if they're working solutions to them....
Antirez said literally he didn't know about existing solutions when he went to write redis, and he and redis are awesome. nothing bad about that
but I get your point about bad solutions are bad but that's sort of a tautology, doesn't add much value, and who are we to judge someone else's solutions are bad we don't know everything about their use case.
Again... even if we can say that you choosing someone else's technology for your problem is not a good solution we just can't criticize the author because it's your responsibility what you choose. so I just don't think it's valid to criticize the author
They can be lifelong experts on their problem, yet have no clue about writing a database engine and low-level programming in general.
> Antirez said literally he didn't know about existing solutions when he went to write redis
Nobody is born with knowledge. The difference is that Antirez studied previous solutions, studied how to do it, and then applied that knowledge right.
Instead, that person did the equivalent of building a bridge disregarding everything humans learnt about it since the Roman empire. It will not be a surprise if the bridge ends up collapsing.
Also, the writing in the README feels sloppy, which doesn’t inspire confidence. For example, you might want to decide if it’s called jsonbase, JSON-base, JSONBASe, Json-Base, JSON-Base, json-base or jsonDB.
They’ve called it a database. They have said explicitly “ You can use this as a backend for your ReST APIs.” But it doesn’t meet the table stakes for a database and encouraging folks to use it in a production environment is actively harmful.
I wish more folks were up front with the trade offs they make. I respect an OSS author a lot more when they are honest and upfront with what a thing is good at and where trade offs have been made.
When I don’t see that, I assume that either the author doesn’t know/care (red flag) or they can’t be bothered (annoying).
It is also impractical to expect the Creator to anticipate all the use cases and potential benefits and pitfalls that people might find in those different use cases and express them.
Second it's fundamentally a violation of a boundary about choices. The people who make the choice to adopt software or not are the ones who are responsible for the technical debt or credit they allocate by making that choice.
Instead of criticizing creators for not adequately disclaiming their new products because of a hypothetical or real harm that is incurred because people choosing that, you should criticize the people selecting things for being irresponsible with the projects they are responsible for.
If your evaluation of a project is simply based on reading the readme at a superficial level then it's nobody else's fault but yours if you end up with problems with the tech that you choose.
I'm not saying you're being mean here I think this is just a misguided attempt to try to avoid technical debt but it doesn't focus on an effective way to do that. What I feel is disappointing is how this sort of criticism is often leveled at new projects as a way to dismiss or I think unfairly criticize these creations and their authors, maybe as form of "concern trolling." if I understand that term correctly.
Like, "don't use this new project in production" is sort of a tautology of "be careful about any tech that you choose that it's suitable for your use case", which is pretty obvious and I think low value thing to say, but it's often said about new projects in a way that suggests "this project is terrible and the author is bad for suggesting that people even think about using this". which I think is very toxic to a culture of creation, invention and tinkering and it's disrespectful of people who put in the effort to make something. it also encourages something which I think is harmful which is the need to think "I need to make this project perfect and bulletproof before I even think of releasing it" which I think means there's a lot of projects that could have benefited if they were appreciated at the small flame level, but maybe people are discouraged from putting them out there because of this sort of misused criticism.
even though I'm not really a fan of his I think Paul Graham said something about this point regarding startups that's like a startup is like an idea that's just being born and it's very fragile so you have to kind of protect it but it can grow into something really amazing.
Now I’ll accept that XML databases has their use (especially if it involved storing and transforming third-party XML) but I can’t think of any good use for this when there’s SO many better options.
This has all happened before, and it will all happen again.
The main issue is that contrarily to a DB, any modification will shift everything after it, so any indexing will have to be corrected. I suppose that if the document is not stored as is, but instead broken up in pages (filesystems are likely doing that already, so piggy backing on that could help), then indexing could be improved, but then storage starts to look like a regular DB, rather than JSON.
Interesting nonetheless, time will tell.
Consider:
{
"name": "Sam",
"age": 12,
"friends" [
{
"name": "Tom",
"age": 15
}
]
}
Attributes like name and age are properties of a person entity, when placed in a JSON hierarchy something else is happening, the one dimensional relationship the things have between each other is also being saved into the structureThat's dangerous because relationships should be formed on read, not on write, otherwise you concrete all future reads towards whatever it was on write, and if you're particularly sloppy the data gets duplicated which is even worse
The solution is to normalise your data store and use relational algebra to reify relationships at runtime
The problem with mainstream databases is they don't force normalisation, automatic indexing and pulling off attribute level normalisation is unworkable performance wise, so in most teams this doesn't work but this idea does work if you want to try this out learn Datomic
Specifically for storing the results/state of a set of manually-executed management scripts. The scripts needed to query the data from previous executions, do some stuff, and store the output. Think poor mans version of terraform.
Everything was dumped in a git repo that was shared across a few people. It was a quick and dirty solution to manage some alpha customers before the "real" system came online.
LokiJS has multiple persistence options, with JSON files in the filesystem being just one of them.
Alternatively, you could also just use it in-memory or with IndexedDB.
I've been playing around with Rust recently maybe I will do a simple implementation in Rust-Lang which will keep it memory safe and efficient.
To reduce the IO overhead, I batch multiple json values into a larger file.
I don't need random access because I'll replay all the changes when the server start.
Going to open source the library soon.
However we are using Postgres as a backend.
Our code is already technically Open Source, but it's not relatively stable and not doc'd yet, so I won't link.
That's probably beside the point, though. Hopefully no one ever uses this so the non-standard async pattern doesn't matter.
You have to run a “freeze” command before editing the database directly (so it can flush the current version of the database, and redirect writes to memory + log), and then “thaw” so it can read your changes and apply the log of updates to it.