Can you share the use cases? Why Mongo works better than Postgres?
Can you share the use cases? Why Mongo works better than Postgres?
Circa 2013 or so I was working for a regional newspaper group. They decided they wanted a weather section. So we subscribed to, iirc, Accuweather.
I wrote a script that would pull it down at the update interval (Every 15 minutes I think it was)
This was, for the time, for a shoestring outfit with only a few servers, a lot of data.
Something like 50k locations, with around 30 “rows” for each location - current conditions plus hour by hour going out a bit more than a day.
Loading this into our poor little MySQL server was… slow.
So I just setup a small mongo instance, and loaded/read from that. Worked great… we didn’t care about durability so we dodged that whole minefield. Mongo had a way to atomically swap two tables… so the import script would load into a scratch table, then swap that into the one the web server actually read from.
Loaded super fast since we didn’t care about durability, and read performance was more than adequate for the very simple queries we ram against it.
(To give you a feel for the era, this is when MyIASM was the default table store, and didn’t support fks or transactions)
With Mongo, its distributed by default. And if your data keeps growing, you simply add more nodes without any additional changes in the application code.
Of course it haa a perf impact - nothing is free.
Live resharding indeed helps for a sharded cluster that already exists.
JSON indexing is quite advanced in mongo. Postgres is catching up (gin indexes on jsonb arrays). So are aggregations on json
MongoDB Atlas - fully managed and by MongoDB core committers
We don't have a lot of multi document ACID requirements.
We are on GCP and cloud sql is a very basic postgres deployment and we don't want to mange infra
The most obvious one being where you are requesting data and related data based upon a unique business identifier.
NoSql DBs mostly make this both trivial and fast.
Postgres and other SQL options become relevant where aggregation and reporting become primary considerations.
NoSQL Databases arose from a realization that sometimes you don't need all the features of a relational database, for example:
- Redis which at the most basic level is a key value cache
- MongoDB which stores (denormalized) documents
NoSQL Databases like MongoDB are best when:
- All data related to a certain entity can be reasonably self contained (i.e. does not require relating to other possibly dynamic pieces of data)
- referential integrity (knowing that User.storeId can be treated as just a string) does not need to be maintained
- Constraint checking does not need to be performed or can be performed at the application level
- Schemas are either wildly variable or do not change at all, or checking them just isn't important.
The best example I hear often is a news site or blog -- if your main model is something like an Article that contains the author, content, tags and all data necessary to display one entry (and 99% of the time you get one Article and are not required to request related data that exist in separate collections), then NoSQL makes sense.
It's a tired conversation I think people have had repeatedly for a long time, so I won't go into it again (I'm also VERY biased towards RDBMS and in particular Postgres) -- would love to hear from experienced MongoDB proponents though!
One case I know Mongo was trusted for very early on was large scale out use cases, and after WiredTiger their execution engine improved immensely as well. I assume it's still great for that, where scale out is often quite lacking in RDBMS or not in the core experience.
[0]: https://en.wikipedia.org/wiki/Relational_database#Relational...
RDBMS puts a lot of effort to support a wide range of ACID properties. However which also make it very difficult to be used as cluster.
So if your bussiness data is too large to be hold by a single instance of database, you will need a database cluster. And maybe someday, you want to get better performance, then you change your schema to remove some data constraints. Like removing FK, use application logic instead of storage procedure, etc.
That's the reason why we need NoSql database.
If OLTP is not the key part of your business, and it is expected to store a large amount of data in future. Then it would be better to carefully design your data schema so you can use NoSql database instead of a traditional RDBMS.
E.g. use nested document instead of FK to store 1 to 1 relationship. So you can take advantages of single document atomicity. Use array to store 1 to many relationship instead of an additional mapping table.
> RDBMS puts a lot of effort to support a wide range of ACID properties. However which also make it very difficult to be used as cluster.
Note that while common, ACID is actually not a requirement for RDBMSes -- there are NoSQL datastores that provide acid guarantees, like FoundationDB[0].
> So if your bussiness data is too large to be hold by a single instance of database, you will need a database cluster. And maybe someday, you want to get better performance, then you change your schema to remove some data constraints. Like removing FK, use application logic instead of storage procedure, etc.
I agree with the point, but I want to note that the vast majority of apps do not have data too large to be held by a single instance, especially with the eye-watering density of hardware these days.
Some proof for my essentially wild conjecture:
- Reddit started with just postgres[1]
- LetsEncrypt (the reason most of the internet will have TLS in the future if not already) supports over 235MM sites with just one MySQL box[2]
Now it's not that you can't mis-use RDBMS (Postgres), or that it's always the right tool for the job, but I just want to note that it's very possible to get very far with application size without scaling out horizontally. Scaling out horizontally & compromising your data model should be the last options you pursue, IMO.
Also, that said, Citus for postgres is now fully open source[3], TimescaleDB has been open source and only got open-er[4] (they have a clustering mechanism) so this diminishes the use case for NoSQL going forward. This means that the specific use case for NoSQL is eroding somewhat.
> If OLTP is not the key part of your business, and it is expected to store a large amount of data in future. Then it would be better to carefully design your data schema so you can use NoSql database instead of a traditional RDBMS.
I'd argue that you should default to OLTP and only deviate when you need the extra power, but we'd probably agree to disagree there :)
[EDIT] Highscalability.com has a great writeup on this from 2010 (!) that as I skim through still looks mostly relevant:
http://highscalability.com/blog/2010/12/6/what-the-heck-are-...
[0]: https://www.foundationdb.org/
[1]: http://highscalability.com/blog/2013/8/26/reddit-lessons-lea...
[2]: https://letsencrypt.org/2021/01/21/next-gen-database-servers...
[3]: https://www.citusdata.com/blog/2022/06/17/citus-11-goes-full...
Yes, most of NoSql provides some kind of ACID guarantees, however not as good as RDBMS. MongoDB also provides single document atomicity. But in RDBMS you can configure the ACID properties.
> but I want to note that the vast majority of apps do not have data too large to be held by a single instance
I agree with that. For me, the order would be: 1. Do not use database 2. Use SQLite 3. Use single instance RDBMS 4. Use single instance NoSql 5. Use single instance RDBMS + NoSql cluster 6. Use RDBMS cluster
> I'd argue that you should default to OLTP and only deviate when you need the extra power, but we'd probably agree to disagree there :)
Personally I would prefer NoSql when the business model fits some specific patterns. But I agree with you that.
The article of Highscalability.com is very useful, and also some other articles. I did not know the site before, thanks for sharing.
Yeah it’s so funny, HN seems to rediscover way of thinking every so often! I think it’s why SQLite projects are so popular (and of course how well built and popular SQLite is)
And multi-document, in sharded clusters. 2018: https://www.mongodb.com/blog/post/mongodb-multi-document-aci...
>>>But in RDBMS you can configure the ACID properties. https://www.mongodb.com/docs/manual/reference/write-concern/
Am I missing something here?
I don’t think that’s right though? Requesting data and related data is the purview of RDBMS, unless your example was a completely denormalized single object (I know Mongo supports links between objects as well but that doesn’t seem to be what you’re saying?)
Also the line point about SQL based databases being relevant when reporting and aggregation misses the importance of relations and operations in relations that is the point.
Maybe the comment was focused in on “SQL” rather than using it as a shorthand for “relational database” like most people do. Yes SQL makes aggregation and reporting easier but it’s primary job is data retrieval, and in most cases via an identifying primary key.