Shouldn't this be "otherwise you wouldn't be capitalizing on the advantages of a schemaless, document-oriented database"?
The fact that it's JSON (vs. some other format) seems immaterial
https://blog.jooq.org/2014/10/20/stop-claiming-that-youre-us...
But your main point is good—- JSON isn’t the essence of it.
And yet it’s part of it, in that you could store XML and still be a document DB; MongoDB definitely chose JSON as the medium deliberately.
Literally every project I saw using MongoDB ended up going back to SQL within the first 2 years after realizing the data is indeed very much relational and theres no clean way to model it using documents.
You always end up with either tons of duplication across documents, which is hell to maintain, or tons of multi-document queries with hacks to look ACID, which is also hell to maintain.
Sure Mongo makes it easy to prototype applications, but it makes it very complex to build robust and maintainable software. Its especially bad if you think your data isn't relational, because it almost certainly is.
Disclaimer: I believe Datomic to be the game changing database; because it values simplicity and composition and these attributes drive the entire design.
A simple example would be a marketplace.
Player buys item X with Y gold from another player.
1. Server checks that item X exists.
2. Server checks that the player has at least Y gold.
3. Server removes the gold from the player
4. Server gives gold to the seller.
5. Server removes item from marketplace.
6. Server adds item to inventory.
What if someone maliciously crafts two requests in a way that step 2 of the second request happens before step 3 of the first request? The money is deducted properly but the account can now have a negative balance and there are now two instances of the item.
* Trading in game items between two users (needs multi document atomic locks if you don't want duplicate or lost items) assuming your "schema" is a document per user
* You want to rename or restructure an attribute in the future, with no schema it's not possible change migrate data easily without writing ad hoc code (maybe you can use third party tools) or changing queries to expect data in multiple "schemas" which quickly gets painful
Good Luck!
So yes, it's annoying for a very small % of what I'm doing, but 99% of my updates/writes are within a single document, so I find it very nice for development.
You can have schemas with MongoDB. There are various libraries to facilitate database design by schema specification.
Also renaming or restructuring your data is not necessarily an easy task with SQL. The nature of a database dictates that how good it works for your application depends on up to how well thought-out your schema is. Having to change your schema around is tasking. One of the reported advantages of document stores when they were becoming trendy was that it was easy to change your schema since your schema is essentially determined and regulated at the application layer.
Also MongoDB has ACIDic transactions now (freaking finally) so if it’s as-advertised then I feel like half of your argument is not really a strong one any more.
This data is meaningful because it allows to analyze what's going on over massive systems, detect when problems will happen, find bottle necks in applications and infrastructure, among many other use cases.
Application domain data tends to be relational I'd agree. But in general, this makes up a very small percentage of meaningful data in the world.
I find that hard to believe. Maybe not that the raw data isn't already relational, but that there are no relations real or implied.
If logs contain info about 'things' and any of those things can be considered to be the 'same thing' for multiple entries, then there's a relation right there – entry to thing.
And even metrics and network event data I'd expect to be full of cryptic IDs that reference some 'thing', i.e. a typical 'code' for which it's really nice to have a table with at least a friendly description.
Admittedly some of this data – or maybe even most of this data – isn't very 'deeply relational', but it definitely seems that claiming that "there are [no] relations in it" isn't strictly true.
But it's all semantics really at that point.
Anyways, a relational database is a poor solution for this type of data. The stored data gains little to nothing, and may even negatively affect it's integrity (at time t, the event DID have this ID; it DID have this label), when stored relationally. Each event is discrete and there will be many of them which optimizes better for scale than relational organization.
I guess my point was there is vastly more useful data appropriate for a non-relational database than there is for relational databases. You might say it still has a "relation" in an abstract sense but this data does not need relational semantics within the database it resides in.
Have you used it? If so, what was your use case?
In my experience, Mongo is most often used with ORMs that emulate joins, like Mongoose. And the possibility of data inconsistency due to lack of transactions is ignored, or patched over with cleanup scripts after the fact.
For any "real" system that is going to be in production for a long time this becomes a real problem
There are tools to "migrate" data but they come with all the limitations of the Mongo isolation model
Typically you either
* Write ad hoc (possibly using some tooling) code to iterate over your old data adding or mutating the field(s) in question
* Write queries such that they can handle the data being present, absent or in different forms for all of time. As you could expect this is a large burden
It hurts when the next requirement comes along something like...
"As a user I want to have a home, work, and mobile phone number"
Now you have 3 "versions" of your implicit "schema" to contend with
1) No phoneNumber 2) phoneNumber and mapping it into / out of one of the three phone numbers in the UI 3) objects with three properties homePhoneNumber, workPhoneNumber, mobilePhoneNumber etc
Then the business comes up with "As a user I want to have arbitrary phone numbers that I can label" now the developers start to squeal
RDBMS + SQL is no panacea but having DDL operations like the following (all probably syntactically invalid but you get the idea) out of the box is incredibly powerful. ALTER TABLE user RENAME COLUMN phone_number TO home_phone_number; ALTER TABLE user ADD COLUMN work_phone VARCHAR(32) NOT NULL; CREATE TABLE phone_number (id BIGINT NOT NULL, user_id BIGINT NOT NULL, name VARCHAR(64) NOT NULL, phone_number VARCHAR(32) NOT NULL);
I have had reasonable success using MongoDB as a store of "things that happened" and will never change
Depending how much graph relation stuff you need you might be better off just using the graph API. I have no experience with that though. Or they support a MongoDB API if that covers your needs too.
Like I said I've only done basic stuff, but I really liked it - it's performant and really easy to set up and use. I used the Python API and it was really easy, then I switched to the Node one to try using it in Azure functions (Python library imports aren't really supported there) and that's nice too - it uses promises and works great. It also doesn't feel like a giant lockin (IMO) - their APIs work anywhere and there's no magic in Azure AFAIK to make you put your compute there if you're using the Database.
[0] https://docs.microsoft.com/en-us/azure/cosmos-db/sql-api-sql...
Mongo could have implemented SQL on top of their storage engine a long time ago minus the joins. Instead they built their own query mechanisms. Mind you, mapreduce can't be reproduced explicitly in SQL, but SQL expressions can compile to Mapreduce (see Apache Hive), so even that was not an excuse.
Edit: NoSQL served a purpose to remind people that there were other options other than relational databases (including those that predate the relational model and those that came after it), but man, what a terrible and misleading misnomer.
Speaking as someone who ported an application from MySQL to MSSQL there is still a lot of work required to remove those custom extensions, but the core of what you're doing can remain the same.