Jsonb: Stories about performance
erthalion.info
erthalion.info
The hate on MongoDB is strong on Hacker News / Reddit but I suspect most of you never had to design a solution when master/slave is not enough.
Also: https://www.mongodb.com/mongodb-3.4-passes-jepsen-test
The simple scalability of mongo (sharding/replication) is something that, while the recipient of much mocking initially, now at 3.6, gives enterprise level availability and scale with minimal oversight or expertise.
I'm really liking the features in 3.6 (over the wire compression, replication hardening, streaming query subscriptions to name a few) and every release packs a ton of features and improvements without breaking anything.
Postgre is a great relational database (one of my favorites) but its json field (which even mssql sports nowadays) is not a stand-in replacement for mongo in anyway way.
As for scaling, it certainly scales out of the box with partitioning, but admittedly a partitioned table doesn't offer all the features that a non-partitioned sql table offers. A great assessment of the current state of postgres scaling can be found here: https://blog.timescale.com/scaling-partitioning-data-postgre...
If you have concrete criticisms say them so we can have a discussion, don't write snarky lines like this is some kind of "us vs them" game.
I've increasingly found that the number of situations where this is actually a good idea are quite, quite rare.
In any real world system I've worked there are always entities that are purely self contained, but are represented by many different database tables. When they're saved, the entity is decomposed into its constituent parts and written to those tables. When they're retrieved, those tables are joined together to re-create the entity we care about. And that's all that ever happens, it's just serialised and de-serialised.
Isn't it premature optimisation to normalise that entity right out of the gate? Why not stuff it in a JSONB column. When - and only when - you find yourself needing to dig inside of it for queries in other parts of the system, then you normalize it.
CouchDB actually solves that problem really well IMO. You upload a validation function in javascript to the database that will validate every document that comes in. I wonder if the same thing can't be done with trigger functions in postgres.
But programmatically, i do ensure only one microservice (out of process) or module (in process) is writing to a collection. When I'm using an RDMS and developing a system, I also ensure that only one module is writing to a subset of tables that make up an aggregate root. I would never have a system where a bunch of programs are writing to the same tables willy nilly without going through a common interface.
You basically can. The difference is that it's quite obviously notable to a developer of any level that changing the schema in the database is something that should be done with some forethought and care. And then when that change is made, all applications accessing the database get the same new view of the schema so you're not going to have two different clients with a different idea about the schema trying to operate on the database at once because schema is global.
Making a change that suuuuubtly alters the way data is stored in an unstructured object is the kind of thing that's really easy to overlook in a pull request.
Your "only one module" rule is laudable, but you've also got to make sure that module can do everything you'd ever want to do to the database, including all the "one off" data mangling admin tasks anyone occasionally has to do. Otherwise it will get bypassed. Case in point I used to use an ORM heavily, but it wasn't possible to express everything in that ORM, so where that broke down we had to resort to manual sql. Add a few developers to the equation and you don't know exactly the format of your data.
Good schema management practise also means you end up defining all of your schema mutating operations as migrations, leaving you with a rather vital log of how things have changed over time and a good pinch-point to catch inadvisable changes to schema.
And if they change the schema in Mongo, with the official C# driver, by default, any program that tries to read the collection<T> and the document in the databases doesn't match the C# class, you'll get a loud and obvious error at runtime.
Of course if you want to have the flexibility of having objects change in the data store not affect downstream changes, you can decorate your class with [BsonIgnoreExtraElements] attribute.
You'll get all types of And then when that change is made, all applications accessing the database get the same new view of the schema so you're not going to have two different clients with a different idea about the schema trying to operate on the database at once because schema is global.
My contention is that you should not have two "clients" operate on the same schema. All clients should be using one module. I'm not advocating that microservices are the one true way. I'm saying that all clients should be using the same module, microservice, or stored procedures (well I hate stored procs for that type of stuff, but that's another rant)
Making a change that suuuuubtly alters the way data is stored in an unstructured object is the kind of thing that's really easy to overlook in a pull request.
There is nothing "unstructured" about a data class (POCO/POJO) in a strongly typed language. Yes you can operate on a Mongo Collection in .Net by getting a Collection<BsonDocument> but most of the time you're going to get something like a Collection<Customer> and the compiler is not going to allow you to insert, search, etc. anything but a C# defined Customer with C# defined data types.
Your "only one module" rule is laudable, but you've also got to make sure that module can do everything you'd ever want to do to the database,
That's the beauty of .Net, Linq, and expression trees. You can define your repository with a method of something like:
Find(Expression<Func<Employees,bool>> query) {...}
and the client can pass in any arbitrarily complex, strongly typed, Linq expression and your repository can pass that expression to the Mongo driver and it will translate that to Mongo query. It gives you even more flexibility because if later on you decide to use an RDMS, those same client queries can be translated to Sql by the Entity Framework driver.
But even if you don't want to go as far as "one true module" for your data, there is no reason not to have one module that defines your data types. Anyone querying against or inserting into a collection<T> will always have the right document definition of course you can put all of the definitions into a Nuget package and it will still be obvious when changes were made because your nuget package will be versions.
including all the "one off" data mangling admin tasks anyone occasionally has to do.
A one off type of admin task wouldn't be part of your production flow and you would probably bypass any abstraction that you had - including stored procedures, views, etc.
* Otherwise it will get bypassed. Case in point I used to use an ORM heavily, but it wasn't possible to express everything in that ORM, so where that broke down we had to resort to manual sql. Add a few developers to the equation and you don't know exactly the format of your data.*
There should always be one team responsible for your business objects (in DDD your aggregate roots). Why would you let any one developer change your underlying data store without a code review and/or pull request?
Good schema management practise also means you end up defining all of your schema mutating operations as migrations, leaving you with a rather vital log of how things have changed over time and a good pinch-point to catch inadvisable changes to schema.
If I'm using a strongly typed language like C# and I'm defining my documents based on a class. My Employee.cs file (and all of the child classes) are in sync with my Mongo document store. The language ensures that at compile time. "knowing my schema changes" is simply a matter of doing a "git blame" on the Employee.ca file.
Even if I'm not abstracting all of the data retrieval in a repository, I don't see any reason not to at least have a Nugst package of all of my POCOs that define the documents in my collection.
But when you have more than a couple of developers on a project and have ... mixed ... levels of experience, things easily get out of hand while your back is turned. Before you know it you've got a dataset with fields that are... well, you don't know what they are - they could be strings, they could be true, false, entire sub-objects or arrays - they could even be json-null. And then you have to potentially cope with all those possibilities in every place you use that field.
Yes, of course you can attempt to put some sort of constraint on the data format (jsonschema perhaps?) but that's not exactly easy to perform as an in-database constraint which is the most useful and powerful form of guarantee you can have.
Of course, it also leads less experienced developers away from using postgres' rich types (even datetimes are too rich for json/jsonb!) and leaving it impractical (or impossible) to perform logic on data in-query (which is often the most efficient and robust way of doing it). This road also leads to people trying to make references to other objects within json blobs. Try ensuring referential integrity on that!
It all comes down to: you have a schema, whether you think you do or not, and if you let the database know about that schema it can help you manage it.
Is there an image missing or something?
Replication and failover don't look so great in postgres either.
My point would be that there are no "right" use-cases, I pointed it out because the article was talking about a specific thing: as the article points out in the outlined scenario: "It’s clear now, that throughput for PostgreSQL and MongoDB now is almost the same before the spinlock performance degradation hits MongoDB."
So the article points out there is a significant performance degradation after a certain number of connections in MongoDB, such things won't surprise me anymore I'll add it to the heap of things wrong with MongoDB, thats why I was making the comment.
> "We’ve got so far that throughput of PostgreSQL and MongoDB is the same for read workloads. Is it surprising? Not at all - I’m going to show you, that under the hood all our databases use more or less the same data structures and the same approach to store documents."
I'd argue that there is no reason to use MongoDB because there are no use-cases other more mature databases won't be equally or better suited for. MongoDB is horrid, you won't know all its deficiencies when you get started with it, but over time you will come to despise it, it is really far away from being a mature database and I would advice anyone against using it for any task, and I'm not the only one who was burned using it and who adopted this attitude towards it.
What a remarkably convenient API it had. NoSQL/"Document store" stuff was relatively new (to me, at least). Did the convenient API contribute to whatever design problems they had? No idea.
But if they had problems with durability then it really is all for naught. I didn't get beyond prototyping so I never saw whatever issues caused them to get so much ire.
You're being hyperbolic. Yes, there are so many better choices for databases out there. There are also many worse choices. Mongo is very hard to recommend nowadays, but unless you go in completely ignoring its problems its not like your company is going to go up in flames.
Do you expect to hear them say, "Yes, choosing MongoDB was a bad decision and our production infrastructure is currently on an unwieldy, slow, dangerous piece of crap" and then go on to sell the company for tens of millions of dollars?
Do you expect people to admit the real reason they're using MongoDB, which, in 99% of cases, is "we don't know how to use SQL and are really perturbed that anyone thinks we should have to learn"?
It is true that MongoDB only spontaneously combusts occasionally, but that doesn't mean the choice is of negligible importance.
All of the complaints I've seen against Mongo are from people who read Aphyr's blog from their SQL High Horse and say "ha, look at those dirty Mongo peasants, can't you see that one in every ten million reads are inconsistent under high load? Why would anyone use this crap?" Use it in production, even under significant load, and you realize that sure you might hit a snag every once in a while, but its tolerable. It works fine. Its not trash. Its adequate. Understand the problem you're trying to solve. Understand Mongo's limitations.
First, word gets around, "professional to professional" or not. Second, yes, people will subconsciously shift the blame to whatever they feel the least responsibility and/or enthusiasm for (aka "confirmation bias"). It's hard to find people who are sufficiently secure to even admit such errors to themselves, much less their colleagues, whom they want to think highly of them. Large doses of confirmation biases and other significant misconceptions that lead people to genuinely believe that their "shit doesn't stink", as it were, would definitely help whilst closing a multi-million dollar.
I've found that most of the technical choices that people make are simple fashion statements unless the technical limitations strictly disallow it (i.e., there is only one solution capable of providing a workable outcome). Software "engineers" and "architects" lie to themselves to justify whatever makes them feel more special and important. You don't have to look very far to see this in action (see: anyone claiming MongoDB was a good decision outside of a very unique or limited use case).
>All of the complaints I've seen against Mongo are from people who read Aphyr's blog from their SQL High Horse and say "ha, look at those dirty Mongo peasants, can't you see that one in every ten million reads are inconsistent under high load? Why would anyone use this crap?"
Please be assured that I read "Aphyr's blog" only when it hits an aggregator or when it comes up in a Google search. I haven't read his post on Mongo and edge cases where it may lose a write are of minimal relevance to my dislike for it (though this isn't something to trivialize in a database).
> Use it in production, even under significant load, and you realize that sure you might hit a snag every once in a while, but its tolerable. It works fine. Its not trash. Its adequate. Understand the problem you're trying to solve. Understand Mongo's limitations.
Yeah, I mean, I'm leaving a caveat for the 1% of cases where people do understand their workload and Mongo was the stronger fit for whatever weird edge case is in play.
But in all the cases I've seen Mongo deployed, it absolutely has been a very straightforward matter of "SQL is old and crusty, JSON is new and neato and it doesn't make us think about the structure of our data before we start spraying random values all over the place so let's use that! Woohoo!"
If we saw more MongoDB deployments that were strictly replacing EAV tables, the argument that people generally are using MongoDB for some legitimate technical function would be a lot stronger. But per usual, people are first, cargo-culting, and second, deploying something that trades the upfront effort of developing a schema with the not-so-distant development disaster that surely awaits those who refuse to plan their schemas.
You're right that Mongo doesn't just spontaneously combust, and I personally haven't contested that (I'm not the guy you originally replied to). I administer multiple Mongo clusters at $DAY_JOB and if left alone, they don't cause much fuss. But that is not the same thing as Mongo being a great engineering decision.
Nope, my dislike of MongoDB the technology and distrust of 10gen/Mongo the company is 100% based on real-world experience.
By the way, one-in-ten-million means several times a day. And it was a LOT more frequent than that...
We're looking for thoughtful comments. That doesn't require changing your views at all, just presenting them with more information, and an orientation to good conversation rather than venting.
I've head of a lot of logging (i.e. application logging) companies have great success with Mongo because they can store their entire client's set of data without need to know the schema ahead of time.
Large numbers of clients — really, any more than there are cores — are notably not a Postgres strength.
> I assume this one is most important questions for you now. Why PostgreSQL is underperforms so significantly? The answer is simple, we’re doing an unfair benchmark.
If PG doesn't perform well, it's unfair.
That's hardly biased after and very much before.
/rant