Personally, I find a non-relational database useful when my data model is non-relational.
Personally, I find a non-relational database useful when my data model is non-relational.
A tree is relational. Each child has a relation to its parent.
I can't imagine data that has no relation (no connection) to anything else. Maybe what you meant was heterogeneous (e.g. data elements that do not all have the same attributes) - but even then I can't readily come up with an example.
Yes, I agree that most types of data can be stored in tabular form (or an "n-ary relation" per Wikipedia). I'm just wondering what concrete types of data one would rather store in a document.
I don't think there is a good example. The decision to store some data outside of an RDBMS must have more to do with the processing model or something else.
What else other than the processing model and business requirements would determine how you model and store your data?
I'd recommend the paper What Goes Around Comes Around[1], the first paper in Readings in Database Systems[2]
[1] https://scholar.google.com/scholar?cluster=73661829057771494... [2]redbook.io
https://en.wikipedia.org/wiki/Entity%E2%80%93attribute%E2%80...
EAVT is great as an intermediate format but it is absolutely useless to query for since most of the time you are trying to find a set of attributes for a given entity i.e. full table scan.
What you want is a "wide table". One entity column and all the attribute columns to the right. Often with most of the values set to null.
This is the dream use case for MongoDB since it you can ignore sparse values yet when you query it via their drivers it will appear as a wide table. You can't do this at all in PostgreSQL since you will hit a column limit.
JSONB is designed for exactly this, isn’t it?
The lack of a Spark driver alone renders PostgreSQL useless for most companies.
This is what indexes are for. An index on the entity id should avoid any full table scans.
And you want to build indexes on half the table ?
Good luck with that.
Your math is at odds with your own requirements, null values don't need a row.
Sparcity is an issue for the wide table not the EAVT form.
I still can't imagine what sparse heterogeneous data exists in the world that makes sense to store. Any type of querying or processing requires some kind of structure (even if implicit in the code) which you can just put in different table structures.
You have to make sense of data to process it and that kind of implies a structure, doesn't it? Am I missing some obvious example of heterogeneous data?
One customer column, tens of thousands of attribute columns.
If you need everything about a customer it is a single, O(1) fetch operation which makes it perfect for driving chat bots, call centres, websites, operational decisioning engines, dashboards etc. Almost every large company will have one of these.
You can't really do it in relational systems properly because (a) you hit the column limit, (b) often it is sparse i.e. lots of NULLs everywhere, (c) you need this system to be distributed since it often gets a lot of load.
But where you get into tens/hundreds of thousands is when you have machine learning models automatically selecting and storing important features from the data.
In years past, people called these "data warehouses" and essentially took snapshots of their production DBs and denormalized the hell out of them so that aggregations wouldn't crash the server.
Try modelling a cyclic graph in a relational way and you'll quickly tie yourself in knots trying to update and query it.
The point is, relational databases a great for storing data that you've decided to model relationally. If you decide not to, then you probably want some other sort of database.
Sure. But that database is not MongoDB.
The "you only need relational databases" mantra bugs me though, because it's so obviously not true.
I bet two years ago there was someone out there saying "if you're going to do NoSQL you better use RethinkDB over MongoDB"
How great the technology is, is absolutely not the only factor. Good thing people can take in many different factors when making their decisions.
This talk is a pretty nifty perf overview:
https://www.percona.com/live/e17/sessions/high-performance-j...
That said, if you know beforehand that horizontal scaling will be a crucial factor, probably postgres isn't the first choice. But with how fast CPUs are these days it's usually not important for a long time.
That’s a very naive statement to make.
Absolutely majority of the companies will do just fine, because the hardware improves faster than their demands. Starting out with a distributed system "because one day we might need it" is just silly, because chances are you'll never hit it, and you'll have to pay for the overhead of having a distributed system (which is non trivial).
Actually my company started using PG and had presentation and someone asked if we considered a distributed database so we can scale. The presenter nicely said it was evaluated and this solution worked best, but that was too nice.
1. It's only about 100GB of data
2. The hardware is barely utilized, we didn't tune it (except some standard memory settings), because there's no need yet.
3. Our data is relational (in fact most data from most companies is relational)
Which means you can't use it for any big data/analytics use cases. MongoDB has fantastic client libraries e.g. Spark, Java.