Everything You Know About MongoDB Is Wrong
developer.mongodb.com
developer.mongodb.com
This is what eventually consistent means! If i wrote (key=X, value=Y) to a primary, then read X and see a value thats not Y (because the secondary node hasn't caught up yet), that is inconsistency. Strong consistency (e.g in a single node SQL database) would mean its impossible to read stale values.
That is literally the definition of eventual consistency. How you achieve it doesn't matter to being eventually consistent, just that it is.
See my other comment above for distinguishing between EC and Sequential.
OP is arguing the blog post is trying to change the definition of "eventually consistent"... the author is saying the delay between the master and slave is not "eventually consistent"... when, exactly as you say, it is.
In “eventual consistency” you could read the committed value at some time, no value, or another committed value depending on the exact implementation. It’s a much, much weaker consistency model than Sequential Consistency and Session/Causal Consistency is the closest you can even get to sniffing a stricter model.
I had a longer comment about this new page and MongoDB as a whole but I’ll just swallow my words. They know what they’ve done and don’t really care.
Writes can get lost if the primary dies and the secondary is promoted in the meanwhile.
This is literally their "we are the fastest database just turn off commits" claim, but now in a distributed manner.
I don't recall them ever saying this.
They shipped with syncs turned off by default.
EDIT: actually it was journaling of writeconcerns. Later they backtracked and enabled it by default, but by then their reputation was cemented (or bifurcated into fast, or terrible, I suppose).
https://shekhargulati.com/2011/12/08/how-mongodb-different-w...
Those don't sync to disk on every write.
It will let you delay commits — if you ask it to — so it may batch multiple commits together on a single fsync, but it will NOT report to you that it has saved data durably when it hasn't. Because this is a safe operation, PG will let you do this easily: even on a per-session or per-transaction basis.
Similarly, you can ASK it to commit eagerly without waiting for fsync, again on a per-session or per-transaction basis. But it will never turn on by default, and the documentation of this option clearly calls out the risk: "an operating system or database crash might result in some recent allegedly-committed transactions being lost"
Or more too: you can ask it to turn off fsync (and all other kinds of sync) entirely and risk total database inconsistency — say, because you are running it on a filesystem that guarantees consistency, like ZFS — but it strongly recommends against it just on the off-chance that some over-eager user carelessly turns it off. Because this is generally not a safe thing to do, PG will NOT let you do this easily: you MUST change the server config on disk and restart the cluster.
I don't know about MySQL, but I'd be greatly surprised if they did anything like what MongoDB used to do. There's a reason MongoDB got that reputation, in a world where MySQL — and its surprising data-eating behaviours — already were popular.
As you say, eventual consistency only comes in with multiple nodes. And unless you are willing to sacrifice availability, eventual consistency is the only game in town in a multi-node scenario.
Yes, but the product has been advertised from the beginning on the basis of how easy it is to grow horizontally; the whole "It's webscale!" argument.
----
> Eventual consistency only comes in when you introduce multiple nodes / secondaries. As long as you connect to the primary for reads, you'll never get stale data.
1. Writes can be lost if the primary dies and a secondary is promoted before data from primary has been replicated to it.
2. The whole point of adding secondaries — again, the "webscale!" argument — is to scale read capacities
Is this any different than how things work in eg Postgres?
> 2. The whole point of adding secondaries — again, the "webscale!" argument — is to scale read capacities
The main driver behind secondaries is redundancy/durability. If the primary goes down, there's no/very little overall service interruption.
If you need the latest data, read from the primary. If you care more about faster reads than perfect data accuracy, read from a secondary.
> Is this any different than how things work in eg Postgres?
1. The comment of mine you are replying to applies to any DB that uses Eventual Consistency, as can be seen above it.
2. Yes, this is different in PG simply because PG doesn't do Eventual Consistency by default.
>> 2. The whole point of adding secondaries — again, the "webscale!" argument — is to scale read capacities
> The main driver behind secondaries is redundancy/durability.
In general, maybe — or even maybe in your specific use-cases. But Mongo's prescribed use-cases — the ones they marketed, heavily, even going so far as to make up terms like "webscale" — are about 'performance' and 'scale'. I specifically call this out in my comment.
1) If that's actually true, then both Mongo's marketing department and its technical writing staff have been doing a really, REALLY poor job for years.
2) "You don't know nuthin', dummy!" is not a great lead-in to correcting people's misconceptions.
It also appears to be the case that what I knew about MongoDB was actually true all along, and most of these "myths" are just Mongo trying to redefine the terms or problems to be something they're not. :-/
On the other hand if you look for reasons why mongo db is bad by searching "mongodb wrong choice" or "mongodb bad" and you find this post, by the company itself, it'll likely be kinder to the product than the posts from other blogs.
I actually disagree. Lots of companies have to rebrand, especially when their leadership and/or product have changed.
The core problem here is that the article seems to be an attempt to fix Mongo's reputation through misleading PR (e.g. using different definitions of words than everyone else in the industry), rather than technical solutions.
It's typical of Mongo as a company, though. They don't have technical merit and try to make up for it by trying to convince us all that the sky is green.
It's like those supplements that say "scientifically tested!", which is a true statement even if the scientific test found that the supplement is ineffective and does nothing.
They're trying to establish a narrative about their product which is an already accepted, defined terminology, and which has existed for decades in the Computer Science community.
Eventual consistency doesn't need to be redefined.
If the general public assumes a database to have relational features, maybe it's time to rebrand mongodb into mongostore or something?
If mongoDB tries to be the "eierlegende wollmilchsau" of databases, they have a too large pivot that they're trying to deliver to as a market. Mongodb isn't made for ERM based scenarios, maybe stop trying to push it into that. If people would've wanted that in the first place, they'll likely have chosen ArangoDB or an SQL based database anyways.
To be blunt, I don't understand why they seemingly are so irrational to push the narrative.
Maybe they've seen that most of their customers stop using their product after a while?
If so, then it might be time to stop promising things you conceptually cannot and should not want to deliver.
Also, dear mongodb folks: Web scale is definitely not mongodb as a product. For decades devs have used memcached with SQL based databases and it scaled beyond imaginations, and beyond what a non DDoS testing environment can reflect.
If your product isn't webscale because you do not want to be held responsible for adapters that you refer yourself in your own docs to, maybe it's time to take responsibility and introduce a q&a step regarding a minimum performance standard that every adapter has to fulfill to be recommended?
Regarding debian package version: I personally am gonna stop here, because I hate people complaining about shit, not taking responsibility, and expecting others to do work for free. If debian's policies annoy you, host your own damn ppa. Or sponsor them to help them work on that. But this... this is not okay.
My personal opinion, is that a database should not lose data _by default_
If defaults are touchy, it sounds like they would still benefit from a closer cooperation with Linux distro packagers, perhaps even taking upon the task.
MariaDB has shipped as “UNIX sockets only” and “Only listen to localhost, no default auth”. This increases the difficulty of trying out the database, but for a mature product for mature people, maybe there is no excuse for bad defaults anywhere.
The nice thing is, there is an API compatible version of mongodb distributed with most Unix systems, installed by default in /dev/null. I applaud the authors of mongodb for achieving this level of market penetration in such a short time, defying years of experience from the real world.
Timeless.
- Why did you move to Postgres?
- What was perf impact/improvement after moving to Postgres?
Thank you in advance!The performance boost was at least 10x and even more with complex LINQ queries all thanks to entity framework. At that job, we did a lot of risk analysis, bucketing bonds and creating a composites. Entity framework core and it's LINQ to SQL was a key in performance boost. Not to mention entity framework provides multiple layers cache which we used as our risk analysis algorithms were basically custom version of bin-packing ...
Our experience was about the same - both cost- and perf-wise. It just that our perf boost number was order of magnitude bigger on some tasks :).
And hot garbage for performance. Unlike relational indices on foreign keys, mongo simply... does a lookup for each doc in the pipeline step. Indexing the looked up collection does nothing extra in aggregation, you're just doing a repetitive manual join. A simple aggregation query that I wrote that added a value from a second collection based on its timestamp compared to the original document's timestamp, took at least an order of magnitude more time than without the lookup.
Aggregations on a single collection are very performant. But never try to lookup or join anything.
Scaling data is mostly about RAM, so if you can, buy more RAM. If CPU is your bottleneck, upgrade your CPU, or buy a bigger disk, if that's your issue.
Never listen to this person about anything.
The mongo I'm dealing with now scales by database. Each entity has between 20-100GB of data in its own database; we're adding entities continually. If I try to replicate for performance, I'll be replicating everything--there's no selective replication. If I shard, I'll be sharding within a collection, which is the equivalent of striped RAID--great if that's what you need. I don't. I need to shard at the database layer. I need my queries routed according to the database at which they're aimed, not by the sharding key. Can I? Not a chance in hell with any of the existing scaling mechanisms from Mongo. My current mongo VM is already the largest Azure offers. How do I add more RAM to that?
The way sharding works in general is that your data gets an additional key (unless it has one that works for sharding already), and a router in the stack does traffic control on the query to send it to the correct shard; on the way back, the data is reassembled into a single result. This enables parallelism in your query, boosting performance by using something like map-reduce. Secondarily, your shard management layer can do a lot with shard-level redundancy and dynamic sharding to spread the data evenly across shards.
The scaling axis here is the size of your data--in mongo's case, the number of documents in a collection. As the number of documents grows, you shard the collection to keep queries on it fast. Mongos, the sharding router, only manages sharding with a sharding key on a document, so the only sharding possible is spreading documents from a single collection across multiple shards. If you have 10 databases that each have 100 collections, you get that on every shard, but each collection only has a subset of the collections docs.
I haven't looked deeply into the mechanism, but I imagine this floats on top of mongo replication, where the replication layer cooperates with the sharding manager to replicate only the shard's docs (as Mongo replicates by tailing the oplog, all a shard has to do is ignore oplog entries for docs without a shard key in the correct range).
It's the fact that it works at the collection level that makes it useless for us. Each database, for single entity, is a set timespan of timestream data, with one collection per timestream. Entities vary by the number of timestreams/collections they have, but as they're all the same type of entity, their maximum size is pretty consistent and we don't have problems querying a single collection. We don't need collection-level query performance.
Our axis of scaling is the number of databases, not the number of documents within a collection. What would be literally perfect for us is a sharding manager that routes by database. We could put each entity's database on its own mongo instance/cluster, or an instance holding X databases, and scale horizontally almost indefinitely. We're lucky in that per-entity data falls into a clear range of sizes; we're unlucky in that no such router exists for such a common scenario, which is bizarre to me.
"We're really bad at explaining Mongo and nobody knows how to use it."
It’s just a difference of like, opinions, man. Oh and in 6 months they can claim “that version of MongoDb was version 4.2.6, we’re on version 6.11.42, that’s an oooold version they did that report on”. It’s the always the same tactics.
Yeah. If you dismiss the part that it took them 4 versions and 9 years to be able to make that claim.
ACID by compliance. Definitely not ACID by design and hence not trustable for transactions. Even this article says that MongoDB shouldn’t be used for transactions.
MongoDB seems great for when you need to capture large streams of data in a create-once/read-forever type of way. Probably why is popular among large enterprises.
So if a company doesn't add a feature at version 1.0 then it simply doesn't count ?
That said, mongo (inc?) is has a p good marketing team.