Bye Bye Mongo, Hello Postgres (2018)
theguardian.com
theguardian.com
That gives you not only a long window to test and validate, but also the option to rollback if things suddenly don't look so good with the new system. It seems like people sometimes get impatient and want to do a 'big switchover' with insufficient production testing and no real rollback plan, but in my experience that's almost never a good idea.
The cost of doing a massive, irreversible migration that doesn’t work seamlessly the first time is significant. And often more significant in engineering time, lost revenue from services going offline, trust both from external users and internal teams watching the shitshow, and possibly data loss or backwards-incompatible changes than those server dollars.
It’s funny that one consistent theme of the best engineers I’ve worked with is less their “amazing 10x ideas” than the number of “sounds like a good idea if you’ve never done it” ideas that they shut down. Huge migrations are one.
I don’t know how they got there but that’s expensive, disrupts their supply chain, and I suspect they’d cheerfully pay four times as much not to have to slog through this mess.
For those who, like me, never heard of TLV prior to this post, it's type-length-value encoding.
In Oracle 21c and the Autonomous JSON Database, we introduced a new Native JSON Datatype with it's own Binary Storage format called OSON[1]. With OSON, you can perform partial updates on a json document. The added benefit is that this results in significantly lesser redo log size.
https://news.ycombinator.com/item?id=18717168
When is Mongo ever a good fit? so I've heard stuctured logs, but they can be shoved in a dedicated RDBMS themselves, or just a file system. Unstructured data? But PG now has supports for XML or even JSON. I've heard it's also easier to administrate and scale, which I'm sure it is but then why some businesses do go back to MySQL or PG if Mongo is easier?
You don't have to think that much about your relational model, so you can postpone all kinds of nasty questions about data integrity and multiplicity early on. This enables a team to move fast and break things.
The switch to something like Postgres comes later in the project where you know your requirements better and you need a system where you have more control over your data. You might even have a person responsible for maintaining your data for you on the payroll.
Personally, I usually just grab Postgres these days and use `json` columns for the documents and loosely fitting data, then drag out things into a more relational model as we go.
how about these assumptions or myths?
a. treat it as in-memory document database. expect to lose data if it does not successfully write to disk periodically.
b. concurrent reading and writing. it claims that it is resolved for specific scenario in multiple document-level concurrency,while put a lock on single document issue.
c. Total data storage should not be bigger than memory, although it was fixed in v3.2 with new database engine.
d. multi document queries introduced in 4.2 and 4.4, however it is conflicting with their own recommendations that says Each document should be independent, denormalized data model.
I like their syntax, feel so natural. postgres jsonb syntax is unreadable, a bit hacky for me. should i be concerned with mongodb myths?
This is giving you the choice where in the CAP triangle you want to be.
As you can see the default of w:1 and j:null when run in non replicated mode does not write journal to disk before ack. If you want better guarantees read some best-practice articles, most of them recommend w: majority, especially when run in replication.
c) no issue
d) $lookup works fine and is as old as 3.2, but as with any join, if you try to combine 2 tables with millions of rows each, performance will be bad. That said it's not as good as RDBMS here. No foreign keys or voodoo under the hood optimizations for example. Mongodb sales people will always say that you should model your data differently, without really giving good examples of such. But really - even if multi document queries exist, you should not continue modeling data as BCNF and think of mongo as a drop in replacement for RDBS, the domain should be suited for denormalized documents to begin with.
That was the only time when I experienced a Mongo-based system which didn't suffer from any Mongo-induced problems.
MongoDB offers a fully-managed database-as-a-service as well (Atlas), for what it's worth.
If you are starting with things that are impossible to define (like the infamous phrase "internet scale") or are plain wrong, then you still don't have to shut up but maybe make sure you like your db of choice for the right reasons.
There are people who love mysql because they use PHP and others who love Mongo because they use nodejs. It's understandable: Those stacks have a lot of users and it's easier to get support. Nevertheless, it's not strictly a strong starting point when people are just discussing databases.
So, don't feel oppressed, just be clear with your reasoning.
I like to think of that as "personal blog with a couple visits a day".
;)
With Mongo and null-friendly languages (Kotlin, PHP) I can just start writing documents with new fields right away.
I definitely appreciate the flexibility of a document store but I despise not having schemas for clients. Interface contracts reduce errors and ultimately make life easier for everyone. The database also has tons of information in RAM it can use to search and sort data, pushing those operations to the client balloons bandwidth requirements unnecessarily.
I'm a Pgsql fan yet the months of work that followed don't seem justified. I wonder if they're still happy with it given how document oriented journalism appears to be from the outside.
The other way is that you could store documents that fit a relational structure, into a document-store. It depends on use-cases, and those use-cases exist and are valid.
We can store KV structures in a SQL database, but we're seeing projects using Cassandra and co.
This may not be that big of an issue if you have a few different types of entities and you're not keeping that many relationships between them. But you're still going to face problems when you need to change the relationships, an issue which is trivial to resolve in a relational DB. If there are more than a few relationships and they're even just moderately unstable, you've probably made a critical mess of the application.
I think a key difference here is illustrative: you typically only see that when the problem is well understood to need the characteristics of those specialized KV stores and were willing to pay the costs of working with those trade-offs.
In contrast, Mongo was overwhelmingly favored people who were speculating about problems they might have without much experience supporting that call. Usually they never actually came close to that level, likely because they were spending their time on the much harder problem of trying to build SQL database semantics into their application instead of working on their business.
The key lesson I drew is the powerful value of sticking with proven tools unless you know you can’t. Reading about things FAANG work is cool but people cost themselves a lot by not asking how many orders of magnitude separate their workload and team size from yours.
:(
Postgres JSON operators are cryptic and made my head hurt every time I wanted to accomplish something more complex.
To express it differently, if mongo and postgres where identical operationally, I would choose mongo every day. More friendly and fun. But there are so many horror stories that I am still not sure that it is the right choice for a safe production environment.
JSON support in Postgres is really not intended to replace - even as a stopgap - what nosql was used for.
I first started working with JSONB columns in Postgres around 2 years ago, I think, and had the same initial impression - the operators are certainly quite alien compared to more typical SQL, and it took me a while to figure out how to do more complex things.
I kept at it though, and now it all feels more natural. I haven't touched Mongo for about 10 years, so it wouldn't be fair of me to compare, but I love the performance and power of JSONB in Postgres. And having all this right alongside regular relational tables in the sr database is fantastic.
I often work with TimescaleDB too, which again just works alongside everything else, in one database - Postgres really is an incredible database!
To be fair, I use mongoDB for 8 years now compared to less that 1 year of psql JSONB, so I can't really answer objectively.
There were some issues, some our fault, some Mongo's, but I think anything will have teething problems at that scale.
Just make sure you don't treat it like a bucket. Use well defined schemas and indexes, like with any DB.
long overdue.
NoSQL databases are great for storing arbitrary JSON-like structures, retrieving them by ID and occasional queries/aggregations on the inner fields of those structures.
NoSQL is absolutely not suitable for most business logic and the lack of a schema is actually a major drawback. It feels like a solution (because inserting invalid data will succeed on NoSQL compared to a conventional DB which will reject it based on constraint violations, missing fields or mismatched column types) but in reality you're just kicking the problem down the road and it will come back to bite you because your application now has to be able to deal with this inconsistent data (and in most cases that isn't accounted for with dealing with NoSQL, and the result is predictable).
I've been on a project where the main database was MongoDB and while it worked fine for the most part, we'd get exceptions when the application tries to read some records (most likely from earlier on in the business' lifetime) that had missing keys and would predictably explode. This isn't the fault of MongoDB by itself - the whole point of it is to be able to deal with unstructured data - but the fact that someone chose to use it while they actually needed structured data and constraints to ensure only valid data is inserted in the DB, which traditional relational databases do provide. Of course, it's easier to blame MongoDB rather than admit "we've been stupid and shouldn't have chosen the wrong tool for the job".
It's a poster boy for making a shoddy product and then trying to paper over the deficiencies with sales and marketing.
Part of this marketing includes creating a super easy, slick set up for beginners and just burying the grenades for beginner users to discover later when they're already locked in.
NoSQL is indeed appropriate sometimes, but Mongo never is.
In your case you didn't pick the wrong tool - your type system and interfaces to the DB (ORM) were probably lacking.
yes
I wonder what this ultimately cost. Migrating from a document model to a relational DB requires code change in pretty much every single location it’s used. That’s a lot of engineer time to develop, test, and do the data migration.
According to the post they paid Mongo a lot of money for support and they didn’t have the correct understanding/skills either.
It sounds to me like the core problem is their requirement to run Mongo on their own AWS stack rather than Mongo’s, and while Mongo claimed to support it seems like they didn’t really.
No doubt this was a huge migration but from my personal perspective going from Mongo to Postgres feels like a move in the right direction in the long term. For one you have a lot more options for support!
It is reasonably straightforward to deal with relational databases in comparison, this is a well-known space.
Clearly you don't.
They did have the correct understanding/skills to manage a Mongo DB, and on top of that they bough the MongoDB-Inc-recommended management tools and the MongoDB-Inc-recommended support plan.
Basically they were doing everything by the book, as advertised and advised by MongoDB itself.
Yet they were having problems, so much so that even MongoDB Inc's own support engineers weren't able to help.
After a great lenght of discomfort they evaluated that MongoDB was not a good fit for them, and moved to PostgreSQL.
> That’s a lot of engineer time to develop, test, and do the data migration.
Downtime is more expensive.
EDIT: Also, the article says that they spent at least two months a year planning and executiting database upgrades, since the OpsManager tool from MongoDB Inc wasn't able to help.