Turns out that a run of the mill RDBMS will fit 99% of the problems you're trying to solve.
Turns out that a run of the mill RDBMS will fit 99% of the problems you're trying to solve.
Point is that sqlite could take people way further than they think. Postgres incredibly far. By the time you need a caching layer where key value stores shine best you have many options you can do and resources to do them.
Up to the millions of TPS actually.
https://akorotkov.github.io/blog/2016/05/09/scalability-towa...
Cheap, "faster," less-reliable stores became very big around the Mongo explosion - and don't get me wrong they can have a place, particularly Memcached and Redis which are great for ephemeral data - but you can do a lot of that with a RDBMS.
https://adamdrake.com/command-line-tools-can-be-235x-faster-...
So a lot of caching that many of us had to do to scale systems a decade or two ago is now not necessary until/unless you get several magnitudes up in scale, and while the number of internet users has gone up, so has competition, and the proportion of services that gets successful to need scale up to several magnitudes beyond the old chokepoints at the time is tiny.
Conceptually, the relational model has relations and atoms. The algebra is built on the primitives conjoin, disjoin, project and rename. You can't get much simpler while being as expressive.
The SQL model adds a good deal of complexity, but at least it's standardized and you can generally ignore the complexity you don't need.
A KV store is conceptually simple if you stick to using it as a KV store. Once you try to implement any kind of schema, you have to build that from scratch. As you add business logic to it, you wind up inventing an ad hoc data model, which is likely to be conceptually more complex than either the SQL or relational model.
You may not be aware of the complexity of your ad hoc model, but math will inevitably remind you.
I think this idea needs to be expressed more clearly in CompSci and programming. What is the exact nature of the complexity explosion you are talking about here? Is there something analogous to the increase in complexity going from regular expressions to stack machines? What math are we talking about here? It seems like everyone just explains the relational algebra, then leaves it right there.
Fair point.
> What is the exact nature of the complexity explosion you are talking about here?
I'm thinking in terms of the entities of Occam's razor: "entities should not be multiplied unnecessarily." "Entities" is pretty abstract, we can't identify what complexity is directly, but what we can do is imagine, "what if we tried to build a mathematical model that captures a real life application?"
If we did that, and had a mathematical description of a thing we wrote, we can formalize it by trying to reduce it to some minimal set of axioms.
And then, your more complex rules are derived from those axioms, and if you get your math right those complex rules will be consistent. If you're very clever, you can make it reasonably intuitive.
If you have something that's very complex, what you'd observe after modelling it is you have mostly axioms and very few rules are able to be derived from those axioms. That is, the rules are just the rules and there's no broader reason for them to be so, or deeper consistent patterns. And, maybe some of those rules wind up being contradictory, and they may lack orthogonality.
The relational algebra, being an algebra, is a set of operatations that are closed over the universe of relations, so it's very nicely orthogonal and reduces to a small set of primitives. As relations can be visualized as "tables" they're relatively intuitive, and using techniques such as normalization you can also structure around potential anomalies that add unwanted complexity.
You can also control where your complexity goes. An integrity constraint can enforce some rules about what may be in a relational variable, and if that's enforced, all code that wants to use data from that relation can simply assume that the data maintains that structure. Thus the complexity can be centralized in the system.
Now, that's in the ideal world with a True Relational DBMS, we have to deal with SQL and vendor-specific SQL at that, so we have a rather complex underlying model. But it's typically Pretty Good.
When we build such a model ad hoc, we go by our intuition and are constrained by the market. Our intutition leads us to use more complex structures, that is, structures that if they were expressed mathematically would use a larger set of axioms to describe them. Then, as we want those structures to interoperate, we are unwittingly merging two sets of axioms.
Worse, because we're typicaly expressing it in application libraries, we wind up repeating code, so the complexity is spread all over the application and it especially multiplies if "conventions" are copy-pasta'd.
Hopefully that explains the nature of the explosion of entities / complexity; let me know if I can expand on anything.
It does not satisfy me. One can show that a regular expression or finite state machine is limited in specific ways, as compared to a stack machine or a Turing machine. One can write proofs concerning the number of states a specific machine can be in, given an input of a certain length. The explosion in complexity can be quantified, as can the impact on the effectiveness of testing. By comparison, "using techniques such as normalization you can also structure around potential anomalies that add unwanted complexity," is just an aphorism.
For more context and depth, see previous discussions...
https://hn.algolia.com/?query=GraphBLAS&sort=byPopularity&pr...
[1] GraphBLAS http://graphblas.org
In most cases though, Mongo fits use cases, we never try to use it as a wholesale RDBMS replacement.
As an aside, if you're doing it anyway, it's also a good point to dump to .json.gz in S3 (or similar) as a secondary backup system.