Never use MongoDB (2013)
sarahmei.com
sarahmei.com
But then in real life people built great software with all the above, so I’ll just say a great classic: pick something you know, use it well, build something good, end of story!
No tool will fix wrong assumptions or bad design, we can dive into philosophy here but I’m more of a practical person so... :)
All tools, including programming languages can be bad no matter the skill of the user.
All tools can be bad, but that means that probably you cannot use it well or build something good with it. Every good enough tool can be used to build something good enough, or else it wouldn’t be good enough. And it’s also realistic to assume that every tools popular enough is good enough for something.
But you see, it’s starting to get philosophical over here
All of these have at least a little bit of truth to them, and you should know about the downsides of technologies, even if you decide to use them anyways.
A good database system does a lot to catch errors (including consistency problems), isolate them, and roll them back. Moreover, it will allow you to move performance problems into the database, where it's likely to be handled more efficiently with less application code (e.g. joining in the database is likely to be more efficient than a naive join algorithm implemented in the application).
Some will argue that these are misfeatures and should be handled in the application. In some cases, that is true; but you are probably going to need some other aspects of the stack to be very robust and performant to get reasonable results.
In other words (please excuse my examples as they are intended for illustration and not flamebait), PHP over Posgres might be fine; Haskell over MongoDB might be fine; but PHP over MongoDB is playing with fire.
I'd still say that, in most cases, the database layer is the first place to start to work toward a robust system. Even a proven-correct Haskell program can fail miserably if there was a minor bug three versions ago that wrote some bogus data that wasn't caught by a good database layer.
Saving this to MongoDB directly, rather than SQL seems to simplify things.
With anything else, at the end of the day, you are enforcing FK constraints anyway, so might as well use SQL.
I never had issues with MongoDB performance.
One caveat to this is that I am yet to see a project that needs database sharding in real life, and I have worked on projects with millions of entries in a table and hundreds of writes a minute.
Specifically for analytics/reporting.
The problem is that most analytics/reporting/BI tools SAY they support MongoDB and then they tell you to write a "connector" for each entity, at which point it's easier to just move it to SQL.
It's impossible for EVERY tool to be good. This isn't reality. There has to be tools that are patently bad to use and people have used these bad tools to build great things. But it doesn't change the fact that a tool can be horrible to use.
I would argue that at the time the article was written, Mongo was definitively a bad tool. Things have changed, but not all things.
This kind of attitude and softness towards criticism - "All tools are great" is not how we need to operate. Professional, well articulated and constructive criticism needs to be on the table.
I downvoted the GP for this reason.
I advise everyone here to listen to criticisms and write them as well. Don't be afraid of some kind of a backlash, express freely.
I'm not about making things that simple by default, but I don't like those absolutist titles like "Never use MongoDB", also because the years since 2013 actually proved the article to be kind of wrong.
"Never do X" is what I tell to children about matters that they wouldn't be able to understand, and making a point like "we used an immature tool that wasn't the best choice for what we were building, and on top of that we used it wrong, so you random guy should never use it for anything, ever" sounds like fearmongering to me and it's not something suitable to my taste.
But I get your downvote :+1:
In the case of MongoDB, many people thought that with it you'll have all of the benefits of, say, PostgreSQL, without having to think much about your data schema. History has shown that this is not the case -- as usual, it's about tradeoffs. There are no absolute wins: if you want an RDBMS, you have to put more work in X, if you want a document store then you have to try really hard with Y.
"MongoDB vs. RDBMS" is a very old and tired argument by now but it basically boils down to: people started using MongoDB wide-eyed, optimistically and with more enthusiasm than engineering skill and of course, there were harsh reality checks.
It should be noted that VCs also burnt a lot of marketing money to deliberately create that blind enthusiasm and optimism.
I really appreciate that people took the time to express their opinions starting from a simple statement like the one I wrote, and I don't think that each and every comment needs to be a complete analysis of every possible nuance of the subject we are commenting about.
I'm not a monolithic person, I prefer a conversational approach when discussing opinions, but that's just a matter of personal taste.
Anyway, to make my point I'll use as an example two projects I was personally involved and still bring food on the table: NodeJS + MongoDB (started 7 years ago, still on it); MariaDB + PHP + FORTRAN (started 4 years ago, after the first two years only consulting from time to time). Fun fact: right now I could swap one DB for the other and I wouldn't mind the difference, maybe some minor tricks here and there.
You can get burned with MongoDB (totally inconsistent data model? Well, that sucks bad), you can get burned with PostgreSQL or other SQL databases (thousand line long business critical stored procedures starts getting inefficient? Self join time explodes after some time in production? Ouch...)
The main point I wanted to make is what jeff-davis (your sibling) summarized for me with words way better than my own: "If you are building something and excited, then keep going, don't stop because a blog told you 'never'"
Too often I read blog posts like the one we are commenting when I was starting with software development as a job, they scared the shit out of me! I didn't have any experience on the tools so reading that some much more experienced guy working on some Silicon Valley billion dollar corporation was saying "This is shit: you touch it, you die" didn't bring anything useful to me, just stress...
Over the years I've heard everything and its contrary about... anything! There is normal people here on HN, there are young devs starting right now that read those comments, not everyone is building mission critical rocket firmware, billion users social networks or AI/ML/Blockchain fanta-finance trading bots (not that you are implying anything like that, far from it) but my advice after some years in the business is what I wrote: know your tools, get to like them, get proficient, use them well and try to build something good... your personal experience, be it success or failure, is worth a thousand times more than the success or failure narrated on a random guy/girl blog from N years ago
This is a weird perspective. Most of us work for money and other people are calling the shots on what's in priority to develop right now. Your statement reads like you're commenting about hobby projects?
> Too often I read blog posts like the one we are commenting when I was starting with software development as a job, they scared the shit out of me!
Not to be reductionist, just trying to understand: your point of view boils down to "I skipped a good tool because I bought a fear-mongering tech blog article", is that correct?
If so, obviously one has to form their own opinions as they gain experience. But I wouldn't feel that strongly about those articles. My takeaways are: (a) critical thinking is important and (b) you can't ban people from posting. ¯\_(ツ)_/¯
> I didn't have any experience on the tools so reading that some much more experienced guy working on some Silicon Valley billion dollar corporation was saying "This is shit: you touch it, you die" didn't bring anything useful to me, just stress...
Not sure what point you're making here. Our job is riddled with stress in general (even if you only take this as an entry point: the brain hates learning and wants us to stay still by default). Confronting that reality means exiting software development for many.
> your personal experience, be it success or failure, is worth a thousand times more than the success or failure narrated on a random guy/girl blog from N years ago
That is absolutely true. But we should recognize filter bubbles regardless. I for example only heard one successful MongoDB story and the team switched to another DB 3 years later.
---
If your general takeaway is: "use whatever you want" then sure, nobody can dispute that since we live in a free world. However:
(a) When comparing tools for certain jobs, there are objectively worse tools compared to others. This reality should not be avoided, or denied, or brushed aside with generic statements.
(b) "Just use whatever you like" is not a sound career advice. Some tools will give you better career opportunities, others will make people frown at you for being the bleeding edge adopter but can give you a competitive advantage, and some will just burn you.
It's not about if you can achieve X with any technology. For the most part, you can do pretty much anything with everything. Still doesn't make it viable or the solution with the least friction and most reward though.
I felt your comments were too generic, too optimistic, and rather uninformative. Hence my responses.
- here almost every single company is incredibly small (average is 4 employees)
- we have lots of companies (66 every 1000 adult citizen)
- IT panorama here is like 15 years behind the US
- after university or technical school you know almost nothing, just some basic abstract notions
So you have to learn every single thing by yourself, and if you want to work with today technology you will not learn it at $WORKPLACE because your boss probably is a 60 years old guy who wrote COBOL programs but want you to build "something like Facebook in the next two weeks".
The only window we have on today technology is on the internet, many tools that come cheap somewhere else are quite expensive for us and you must be lucky to know that one guy that can give you some hints.
In my case I've studied in technical school and then did IT engineering at the local university, which I dropped with a couple of exams to go after I had to argue with a professor that solved it's own exercise in a broken way and couldn't understand why my solution worked. In the last 15 years working, until now, I didn't know anyone here in my hometown who can work with anything outside some legacy ASP stuff, basic PHP on CMSs like Wordpress or some basic Java applications (mostly Android), DBs are always MySQL. Just a friend of mine who uses QT, which he learnt by himself and after changing three jobs now he is finally in a dev team, but on the previous three he was the only dev in the company. (I know other people that work with "today's" tech but they are not from my province)
I don't know if I can communicate my point, but reading that $PERSON who works at $COMPANY says that $PRODUCT is shit because he had problems scaling to 100mln users is a totally useless point for all the people outside the few tech bubbles around the world.
It's not about "everything is good", I never said that everything is good, but in the real world™ people built great things using almost any tech, maybe even "uncool" tech: it worked, they made money, they were successful, their company served customers for years and years, and maybe the chosen tech had a positive role in the process. My point is about people that say "Never", "Always", "This is broken", "You'll get burned", "It doesn't work"... that's the part that is not constructive: it didn't work for you, for your case, maybe you did some mistakes, but $PERSON must stick to his/her own reality, not try to extend their experience to everyone else.
Can relate to this a lot and I agree. US / Silicon Valley tech blog articles aren't always informative and one has to apply a lot of critical thinking if they want to extract any value from them. I am with you here.
> It's not about "everything is good", I never said that everything is good, but in the real world™ people built great things using almost any tech, maybe even "uncool" tech: it worked, they made money, they were successful, their company served customers for years and years
I would immediately agree with you if the result-oriented people actually made good money [almost] always regardless of technology. But it isn't universally the case, so very often you have programmers who know modern (and very good) tools but are required to use ancient ones because "they just work".
Truth is that no, very often they don't just work; they "just work" due to the suffering of countless underpaid programmers who can't say "no" because they are afraid for their family's livelihoods. Let's not conflate concepts. ;)
I've seen people replace legacy PHP + MySQL + ssh/rsync scripts for deployment with a managed service like Heroku (or Docker/k8s, or bare metal on two instances) and a more modern programming language -- hosting bills dropped by anywhere from 50% to 800% and like 3/4 of the IT personnel became superfluous (and were fired in a few months). There is a lot of efficiency that's routinely left on the table and is never reached for.
It's a nuanced discussion. Generalizations like yours and like mine can't nail it. I am 100% with you that not everything has to be "new" and "cool": absolutely! But this same argument is also used to drag down or even stop measurable progress, productivity gains and expense savings.
There's a balance to be struck and both programmers and businessmen have too extreme points of view -- on the both sides of the spectrum -- and don't achieve it.
> "Never", "Always", "This is broken", "You'll get burned", "It doesn't work"... that's the part that is not constructive: it didn't work for you, for your case, maybe you did some mistakes, but $PERSON must stick to his/her own reality, not try to extend their experience to everyone else.
Agreed in principle, but not in your particular example. MongoDB is supposed to actually store data. For several years in a row it was failing even in that. You can't claim you are a database and lose people's data. So I definitely would side with that author's viewpoint that "MongoDB sucks" because yes, it does, or at least it did for years in the past and nowadays I am too burned by it to try it out again.
> I don't know where you're from, I'm from Italy so let me put down some points:
RE: the tech landscape of the local market, I am from (and in) Bulgaria and it's almost the same here. I was getting Ruby on Rails offers in 2020 and it was touted as a "bleeding edge modern technology". Those recruiters gave me good chuckles.
In the meantime I am working with Elixir for 4.5 years and Rust for 2.5 years now...
I would reword your comment as: "If you are building something and excited, then keep going, don't stop because a blog told you 'never'". I think that's what your main point was, and it's a good one.
Equally, it’s important to also tone down “X is nothing but the best” and praises should also require equal and opposite constructivism.
Certainly I could endeavor to build amazing modern software in C, but unless for some reason C is an absolute must, I would rather try any other language first. I doubt this is a controversial statement, and yet someone will always be there to defend the opposite stance.
What does 'good' mean? Gcc is a good C compiler and a bad Java compiler. It's even worse at being a document database.
I don't think 01acheru was saying that all tools are equally able to do all tasks. I read them as saying that people have used tools with recognized flaws to make good stuff and that being snobby about whether tools are "good" or "bad" in a general way isn't super useful for anyone. Instead, we should say specific things about specific flaws and let others decide if those flaws matter to them.
In particular, this post isn't really saying that MongoDB doesn't work, it's saying that the MongoDB data model isn't useful for what the author was using it for. Even if you are sure that your app was "the perfect use case for MongoDB" all you can really speak about is your use case. The real headline for this article is "we couldn't make MongoDB work and we're skeptical anyone can," which is totally fair, but shys away from the grand claims that 01acheru (and I) are critiquing.
> No database makes you more productive
> The most popular database for modern apps
What we have here is not really something intended for a specialist usecase, as implied by your comment. Mongo is clearly pushed as being the best database, used by most "modern" apps.
For me, they appear to claim to be popular and focused on developer productivity. Those quotes don't seem to show them claiming to be "the best". Like...it's pretty standard "put your best foot forward" sort of high level description and I think it's a bit much to say that they're making some universal claim?
> Gcc is a good C compiler and a bad Java compiler
Meaning, some tools are meant for a specific purpose and it's not fair to judge them if they do poorly at things they were not designed to do. GCC does not claim to be a java compiler (any longer [1]).
I don't think this apology applies to MongoDB. It is described as a "general purpose" database, good for most modern apps. This is asserting exactly the opposite of what was argued in the parent comment. It is then fair to expect that it should do a good job of representing basic data like graph relationships.
Once we start talking about C compilers (or general databases), I wanted to express the idea that terms like 'good' or 'bad' are too board to be useful. The only way MongoDB would be a 'bad' database is if it didn't do what it said it did (which seemed to be true for a long time[2]). 'Bad' only makes sense, to me, as a synonym for 'broken.'
Instead, tools that perform the game general function tend to focus on different aspects of that function. MongoDB focuses on 'productivity' (whatever that means). My impression was, for a time, GCC focused on overall performance while Clang focused on IR introspection through LLVM[3].
I think the article OP links is valid criticism of MongoDB. They're very clear that Mongo didn't work for their use case and I'm sure they're correct. I just think they go too far in saying "Never use MongoDB" and that what they're really saying is something along the lines of, "we couldn't make MongoDB work and we're skeptical anyone can."
P.s. I don't think GCC ever claimed to be a Java compiler - your link is to GCJ, a different program also made by GNU.
[1] Such as SDCC: https://en.wikipedia.org/wiki/Small_Device_C_Compiler
[2] https://stackoverflow.com/questions/10560834/to-what-extent-...
[3] I don't think this is true anymore? I don't write a lot of C.
I got more of a "stop complaining about tools. Pick one you like, use it and build something with IT instead. I feel the same way. Developers are a fickle bunch. One tool works, but then in order to be "cool" you have to bag on it and then propose some other obscure tool you think is better that nobody has ever heard of.
It seriously reminds me of people arguing over music. Its totally uncool to like a mainstream band because everybody else likes that band and its not cool to like them. So then you have all the "cool" people who listen to all the obscure "awesome" bands who Rolling Stone magazine tells you to listen to, so then you go around telling people you listen to the Shithouse Rats. "Oh you've never heard of the Shithouse Rats? Well, they're kind of obscure." and now you're one of the of the cool kids.
Hard to get using it approved at work, less usage means fewer bugs are caught and features developed, etc.
So people evangelize their favorite tools because it benefits them directly to have them adopted more widely.
But every once in a while you have a case like WhatsApp, which sold to Facebook at a price of $500 million per engineer, which never could have happened without Erlang.
And by the way I really like you music analogy, and it’s emblematic of something even larger: they read about Shithouse Rats on Rolling Stone, named after a song of Bob Dylan or Muddy Waters or the group itself (don’t know which one) and all of them are quite famous and mainstream.
You’ve got to love mankind, we are awesome!
Think of it like horse drawn wagon vs. a car. There might exist a guy who in his humble opinion thinks the wagon is better so he use it to get places instead of a car but is that guy a reasonable guy? No.
The analogy I mentioned above is more apt because Mongo was indeed at one point in time more of a wagon rather than an unpopular piece of music.
And it's a bit insensitive to suggest that you support incest and child rape. Or am I reading to much into your post?
Let's be honest here. You aren't suggesting that you support raping children anymore then I am suggesting the Amish are unreasonable.
This cancel culture attitude of constantly calling out and classifying everything as some sort of infraction against a culture or a race has got to stop.
I agree. Unfortunately, sarcasm doesn't convey well in text.
I concur. Version 3.0 is dated March 3, 2015 that uses the WiredTiger engine which fixes much of the brokenness.
I did some workaround work on a MongoDB v2.x app. It did suck and was inconvenient operationally, but it also did scale so had its uses.
However the discussion now should be about how it is to be used/not today and not back in 2013. So fair to say it did suck or you shouldn't have used it, but that doesn't have much relevance.
Don't those tools die out and get forgotten?
There is no correlation between something being widely used and it being good at its job.
And no design can fix a bad tool, we can dive into the practicalities here but I'm more of a philosophical person so...:)
Also JSON became the standard way to ship data around, and RDBMs systems of the time couldn't really handle JSON. So you either write a bunch of code to map complex nested JSON to relational tables, or just dump it into an un-indexible text column.
There was vendor hype, just like there was around Object databases in the pre-internet days.
If you were starting a new project you needed to decide if you were going to use a document store and an RDBMS or just on or the other. If it was just one you would choose a document store if you anticipated you would need to handle a lot of unstructured data.
Today the situation is revered. A document store only does documents well. A good hybrid database like postgres gives you the best of both worlds. Throw in hosted database services and resource constraints are much less of an issue. So people aren't running back to an old school RDBMS. They are moving to a much superior and evolved data store.
Another benefit of this kind of approach is starting to learn about a challenging subject. Say you want to deepen your knowledge in a branch of mathematics that you find interesting and useful. The history of that branch will tell you so much more than a typical lecture-style conglomerate of concepts. It provides a great overview of important actors, their relationships, cause and effect of discoveries, the culture, the problems and so on. On top of that it is easier to remember and internalize concepts if you know the story behind them.
JSON is the standard way to ship data around the internet, yes. Though grpc is catching up and more and more often I see people relying on grpc in their architecture. And grpc conceptually is a lot closer to RDBM, given that you have a code generation step and everything in your data needs to be defined(aka statically typed).
Recently I started several personal projects and though I struggle to find time and motivation to work on them on my own, document related databases are completely out of the question. postgre and potentially redis as a proxy for heavy loads and that's that. I wouldn't call postgres a hybrid database. It does support json datatypes natively but in it's core it is the definition of what RDBMs are. The best example for a hybrid database(from a developer's perspective since it isn't open source and I do not work for google in any shape or form) is spanner.
I've worked on plenty of both SQL and Mongo projects, and honestly the process around schema migration is pretty much the same. Just for Mongo you write it in the code instead of SQL.
That time is still here if you're running enough read nodes and QPS.
My theory is that it's easy to add a field by adding logic into the app instead of munging tables relationships. Moves the logic to where developers are more comfortable. Scalability/etc is irrelevant for most use cases anyway.
I literally can't parse what you mean by this
Search, with mongodb can do $all query, which is hard to replicate at sql level without aggregation. However I'm still waiting for aggregate-level $elemAt.
Logging, you can attach anything to a property, then it'll be queryable.
Draft records, it's easy to just insert and insert the records because it's schema-less. Validate during creation and validate again during publishing or approval. It's queryable and you can use a generic collection for that.
For logging and draft records, sql JSON field may be able to handle them, though I don't know how good it is at querying.
Traditional migrations for relational databases are really painful. Document databases make this much easier, and if you've faced the operational pain of needing to migrate a large database (for example, it's so easy to accidentally lock an entire table in Postgres), you might be pretty compelled.
(That said, I think the pendulum is swinging back away from document databases. So you're in luck ;))
If you enforce a schema on the Json structure how do you handle the changes on the live system?
The most common use case is, "I need to store data where the schema is unknown or can change without notice, and have my shit not break." This is what we used Mongo for.
The other use case I could see (and this is pretty much only with Dynamo) is, "I want to build an application that's cross-region native. Most of my data is relatively static, so I accept eventual consistency on changes. I will have a separate data store for transactional data and data that cannot be eventually consistent." I want to build this project, but it will never happen because it's too easy to RDBMS in a single region to start.
Well, filesystems are pretty good. It's the only document store I use (and mostly enjoy).
But then you look at the trade-off with some think like just Maildir, and you really start to wonder if this schemaless document store thing is so great?
I suppose the real shame is that proper object dbs like zodb or gemstone gets much less attention - they to have big trade-offs - but I feel they at least give back in terms of consistency and simplicity.
This might sound jaded but my feeling is that a lot of developers just looked at JSON objects that they were already working with and thought to themselves "actually, it would be cool to just store this directly".
Which, in itself, isn't a bad idea but writing a completely new solution from scratch to a problem that's been solved for decades seems a bit like hubris.
AFAIK many relational databases support JSON today, so I'm not sure what the argument would be to choose something like MongoDB today from scratch if you had the choice of anything.
Querying nested values is is nothing like Rethink or Mongo.
Keys could be rows or keys.
You’re still having to make one to many relationships for something that should just be an array.
What? Her database can't possibly be indexed properly.
Comparing fetching from a normalized design to a denormalized one isn’t really a fair comparison.
But then the way tables don't map well to 1-to-many mappings and joins still returning data in tables this can also be a problem. Especially if a large field get duplicated a lot. RDBMS really should go from 2d-Tables to proper nested types for results IMHO.
Active record is awesome in many ways, but it can shoehorn you into n+1 solutions.
I would suspect something like bad statistics or some other reason that caused a pathological query plan. In any case this is not a good representation of any potential performance difference between Postgres and MongoDB.
> joins will kill you when you have very large tables
nitpick, but table size doesn't directly matter (much). if your queries are very specific and only return a couple rows, then you can have huge tables and join across them without issue. Joins only get particularly painful if you're doing aggregation/reporting queries across large parts of it
I agree, if you index your tables. Relational databases are very capable; when there's a performance problem, it's often due to simple things like failing to index what should have been indexed.
No tool is perfect for all use cases. There are cases where relational databases won't work. But when I try to store data, I first consider storing it in files, and if that is unpleasant, I consider relational databases. These are both relatively simple time-tested solutions, and it's usually good to start with simple & time-tested unless there's a reason it won't work well.
Doing the same in SQL requires a lot of intermediate tables.
Schema evolution is supported by pretty much every ORM you'd care to use; it's not the job of the SQL database to handle the migration. I'm using Prisma and you literally change the software spec for the schema and say "migrate", and it creates the migration SQL and applies it to the Postgres DB programmatically. That gets you deterministic schema evolution and not the "my schema isn't actually reliable" that NoSQL/no-schema databases rely on.
And then you have CockroachDB/YugabyteDB that give you extreme horizontal scalability, with full PostgreSQL compatibility...
And bang, the last reason to use MongoDB vanishes.
That said, I will admit the change streams feature is amazing. That completely changed the way I thought about building reactive applications
I've never been anti-Mongo, but this one little piece has made CouchDB an affordable choice for people like me who are not equipped to otherwise defend the choice.
Is there a missing piece that could deal Mongo back in the next time I try to convince someone that there's as straight a path to the sysadmin solutions I generally compose?
https://github.com/mongo-express/mongo-express
But if couchDB works, it works. Personally I'd love to shoehorn LISP into everything I do, but most of the time I just use python and bash because things tend to get done faster when I do.
Maybe if you're trying to build a massive ($B) company, starting with PostgreSQL makes more sense for you. For everything else, MongoDB works just fine.
Just use the technologies that you know and can move fastest with. Startups rarely succeed/fail because of which technologies you choose to use.
In many companies the need to regularly access the database for analytics and BI is a thing; this isn't limited to $B companies. Most of the tooling available works best with SQL databases. (Though the BI connector at https://docs.mongodb.com/bi-connector/current/ looks interesting for this purpose)
If you don't know what your needs are, you should always start with an RDBMS -- it's not that difficult to go "up" to a NoSQL db from there (you're only losing information, and if you can't safely move because of loss of ACID... you'd probably have been really fucked if you started off without it), but you can't easily migrate back "down" to the RDBMS -- a NoSQL database stores almost no information about your data or its constraints.
And your application will almost always want transactional guarantees and to model relationships properly -- generally only small chunks of the design (design-wise; data-wise it might be 90% of the app) can be treated with eventual consistency and have real scaling needs, which you can shift over to your nosql system.
Apps are generally just metadata tracking with a dash of real work.
Unless you mean for solo projects? In that case, it may work for you, but if you're not a database expert, it makes sense you would work around your limitations. That doesn't necessarily mean it's a best practice.
All my proof of concept (and some production) stuff just uses that until I need features or concurrency that it can’t provide.
but even still, you can always use json columns in postgres if you don't want the db to enforce a schema.
Complaining mongo doesn't scale is like saying your Mita can't off-road
Oh you always have a schema. Its just about whether you want to enforce it or not.
My prototyping always starts with
DB = {}
at the top of my file. Sometimes it grows to serializing to disk / loading the dumped object from disk. Often times it's all I need to know that my idea was crap and needs revised. And it always keeps me from faffing about with infrastructure.The idea is if you have just a couple of days to crack out of MVP you don't want to waste time with postgres or whatever. The problem ends up being then your boss is like, all right this works keep going with it
It's less work than setting up mongo, postgres, or whatever long-running data store. Just deploy the web application... and you're done.
I generally dislike when people try to write off a technology just because they don't know how to use it. Don't hammer in screws, but it doesn't mean you should throw out all of your hammers
It's very easy to get going quickly and find out that you've tripped over antipatterns, and now your database is "a database, a Ruby app, and 30ms of latency on every call". It's easy to think "I'll model this later" and end up hitting the database five times in a row to answer a question.
With these systems, up front modeling of your data model and access patterns is essential if the trade off you are trying to make (far less functionality for smooth performance at ridiculous scale) will ever make sense.
There's never a good moment to switch database, but there is always a new problem caused by storing relational data in a document store.
You probably shouldn't use that too much, not for anything remotely production, but it's there if you need that flexibility or an easy place to dump javascript objects.
When you are just getting started, it's super easy to manipulate SQL data structures regardless.
Caring and feeding for a database (of any type) with any type of HA has a learning curve. Hence the growing number of PaaS services that handle the setup and maintenance for you (AWS's RDS, Mongo's Atlas, etc)
It was really, really bad 3-4 years ago. Regularly entering irrecoverable error states while performing basic management operations via MongoDB's management GUI. I've noticed a significant improvement in the past ~1 year.
Given that this happened to me across companies, teams, and a half decade of time, I've decided that this is a case where the problem is Mongo and not me.
Setting up and caring for a DB cluster is a complicated thing, to the point of there being non-BS certification courses for nearly every major HA database including Mongo. It very well could be that there was a well-documented flag that you never learned.
This doesn't mean you're being unreasonable. I'd be cranky with any DB getting itself in a split-brain scenario, but my conclusion wouldn't be a bug in the software but rather it's a bug in my understanding. It's worth noting that getting in this state should either be impossible, or it should be obvious about how it arrived at such a state with links to relevant documentation.
(There's also little incentive to make it super easy to run in production on your own. They sell that as a service after all. It wouldn't surprise me if the product had invariants that assumed production-level configurations and that nobody's tested with whatever configs ended up making it go nuts.)
My startup ages ago wasn't big enough to have perf issues with mongo, but I certainly experienced data loss (ore wiredtiger).
A company I worked for (with the most capable team of DevOps I've seen in my career, incidentally) eventually moved from self hosting mongo to atlas, hoping to ease the pain. This was already a few years after wiredtiger came out and fix a lot of issues.
Performance problems kept happening and we just started switching to postgres.
At my new job, DevOps people keep complaining about our mongo clusters on atlas.
It's again, terrible perf and occasional weird issues.
We did a disk upgrade and the cluster took 3 days to move all the replica from starting up (and replicating data) to running. During those three days we were running with one less replica and everything was to slow. A complete nightmare.
Adding another replica wouldn't have solved anything because it would have taken a day to finish booting.
They've been trying to migrate off mongo for years but there is never time, as usual.
1) Document database - rather than a strict rigid schema, you can store nested json documents in tables/collections. Or the idea of soft schema where the whole database doesn't need to be blocked for a schema change and you have some leeway in integrity.
2) Relational database - Ability to make complex sql queries that join data from multiple tables.
Mongodb has some support for joining but it doesn't have a sql variant. If your data is mostly key:val store then it's great. You can shard it, and have replicas. It's easy to make a fast reliable backend with mongodb. Many popular sites run on mongodb backend.
However with new json types in MySQL and Postgres, it too has support for inserting documents and querying subkeys. It can be sharded and replicated (albeit with a bit more configuration).
Couchbase which is like mongo (in its document store capabilities) N1QL which offers agility of SQL and flexibility of JSON.
So like any tool, it has it's tradeoffs.
Then again kudus to the author for evoking our reptillian brains: "Never use MongoDB" incites emotions and gets you on top of HN. If it was called "When to use MongoDB", it wouldn't get the same reaction.
That is exactly what I'd expect, and that is how small websites like IMDB work. I am on the page for a General Hospital episode, and via the actors in the episode or whatever other part I can click through to Babylon 5, or the other way around, or anywhere else.
Urmmm...How?
So yes I would agree with that. Never use MongoDB.
https://jepsen.io/analyses/mongodb-4.2.6
... not good.
2016 https://news.ycombinator.com/item?id=12290739
Discussed at the time: https://news.ycombinator.com/item?id=6712703
If you use metadata documents to model your relations you might get away with the most dangerous foot guns, but then why not jump straight into graph databases?
Mongo now supports multi-document (and multi-node) transactions, joins, and has a decent storage engine.
So you might even have a chance of keeping your data actually consistent.