It's like a weird version of the Turing test where you have to decide whether someone's speaking seriously or in jest when they talk about NoSQL.
It's like a weird version of the Turing test where you have to decide whether someone's speaking seriously or in jest when they talk about NoSQL.
https://en.wikipedia.org/wiki/Poe%27s_law is the term you're looking for...
So that said, when comparing NoSQL solutions, ending up with an apples versus oranges comparison is almost inevitable.
The number one reason for why people want NoSQL is horizontal scaling and for that PostgreSQL is terrible, with all available solutions being hacks that don't work.
I would refer you to an earlier comment:
> In general NoSQL solutions are optimized for certain use-cases at the detriment of others
Your definition of 'right' is absolutely not the only definition of 'right'.
More to the point: there are a lot of problem spaces where the data integrity provided by MongoDB are more than sufficient.
Define tons of data. I'll bet a beer that what you say can fit on a single postgres instance and even be small enough to run fine on a cheapish developer laptop.
Sure MongoDB may not be the best fit for it but my point is more that these are scenarios where horizontal scaling is an important consideration makes more sense for some nosql solutions than for sql solutions. It's not just about single box performance.
Remember, sharding isn't scaling the database. Sharding is admitting your database can't scale so you're offloading the problem to another layer.
But more common than not, inexperienced devs are using Mongo and similar to store relational data, simply because they were sold the 'MEAN' stack and didn't realize that, while it's easy to get a quick prototype running, a year or two later you eventually need things like transactions and joins, and NoSQL is absolutely the wrong technology most of the time.
But I do think http://www.scylladb.com/ is great.
There was a post on LinkedIn last year, "MongoDB: The Frankenstein Monster of NoSQL Databases"[0], by the CTO of SlamData, which I think produces an analytics package for MongoDB or the like.
The piece is an interesting dive into why MongoDB is what it is (at least as of March 2016) -- and given his connection to MongoDB and its employees -- being a partner of theirs -- it's quite eye-opening (and I'm surprised with almost 200 comments in this thread that no one posted it previously):
Much like Mary Shelly’s Frankenstein monster, MongoDB’s data access layer is sewn together from ragged pieces that don’t fit together. Pieces that were never designed to fit together.
The result, depending on your point of view, is either an Enterprise-grade NoSQL database destined to supplant Oracle, or an unholy abomination of nature, deserving of an angry mob bearing torches and pitchforks.
Let me dissect this creature so you can decide for yourself.
[0] https://www.linkedin.com/pulse/mongodb-frankenstein-monster-...
You might expect a database company to announce improved scalability, reliability, or performance. Or perhaps announce some of the countless features that users have been requesting for years, which are collecting dust in MongoDB's ever-growing issue tracker.
These are sane, logical expectations for a database company, so you'd be forgiven for being shocked at MongoDB's announcements that the company is investing massively into every product category except database technology.
Yet, to those who know MongoDB's troubled history, these announcements come as no surprise. In fact, I even predicted the launch of Stitch just over a year ago, at the last MongoDB World.
MongoDB didn't start as a database company, it's never acted like a database company, and in my opinion, it doesn't really have the DNA of a database company.
If his assessment in this next snippet is correct, one wonders about those who invested in this IPO:
MongoDB has no ecosystem. There are no analytics tools for MongoDB (except SlamData [N.B. this is the author's company] ), no backup software, no recovery software, no data integration software, no query optimization software, no data management software, nothing. Zilch.
There's just MongoDB.
In hindsight, this is an inevitable consequence of a company without database DNA trying to build and monetize a database. MongoDB couldn't figure out how to build an Oracle-sized empire on a database—partially, I'd argue, because they couldn't figure out how to build a database—yet they have to hit their sales quotas.
If you can't sell the database at scale, you end up trying to build and sell an ecosystem around the database. Slowly, bit by bit, MongoDB went after their early partners, trying to put them all out of business to drive a few million here and there.
The result is a "database company" that sells everything under the sun, including database management tools, data exploration tools, cloud hosting tools, cloud hosting services, and soon, BaaS (Bubble makes a return!) and BI software. Everything except, you know, a database.
[0] https://www.linkedin.com/pulse/mongodb-world-2017-lonely-sto...
The development community doesn't seem to care about fixing bugs, and when they do fix things, they reliably introduce new, often devastating defects. "Move fast and break things" is unforgivable at the persistence layer. We're stuck using an ancient, unsupported version, because it is, per our empirical testing, the least bad. But we made that choice, and for now we're stuck with it.
Please, let our pain be a lesson to you. I am totally willing to be someone else's "That's how you get ants..." on this point. Our emoji for it in Slack is a burning poo. It's that bad.
EDIT: phrasing.
Which is this least bad version, if you don't mind my asking?
You say you're stuck with it. Can't you just change a few things, shut down for a little while, export your data, and insert it into a new DB? I know it is hard to do it with writes happening at the same time, but it seems like you could freeze it as read only, extract the data, and put it into a DB that has been prepped ahead of time and tested.
I know we did this more than once BUT it wasn't with public facing data and we were able to migrate while still working on existing data. We just couldn't add more data while doing so.
The issue I've found with the sync gateway is that the queries you can perform are more limited than what you can do directly on a Couchbase store e.g. joining data is difficult.
BTW, Couchbase should not be confused with CouchDB. Similar name and might have some history in common, but Couchbase is more fully featured.
In this context: please, for the love of God, improve your documentation and dev tools. I love the underlying technology but my experience getting a basic service up and running has been pretty mixed. One of the reasons mongodb has been so successful is that there's 50 bajillion articles showing you how to hack up some crappy code that roughly does what you need it to do. Some of those articles are less bad than others, but in a pinch you can find something to at least get you on the right track. Couchbase doesn't have that deep well of experience to draw on. If you run into a problem, or experience strange behavior, it's up to you to figure out what's going on. That would be ok, but even the official documentation and standard dev tools are not good enough. To get people to adopt couchbase you need to do more to get them started.
Javascript-specific whining: it's particularly frustrating to find an official ODM like ottoman and discover that half the features don't work and that there are tons of bugs that haven't been fixed for about a year. These aren't minor bugs either; some of them stop you using headline features. Forget full text search, N1QL is mostly unusable! Check out ottoman bug #153 for details there.
(Is that the same idea?)
(I use sublime..but I admire those who've jumped into vim for the productivity boost that brings).
Not entirely true, but nuance often precludes persuasion.
Many of the ideas behind NoSQL databases can be very valuable, given the right context, but there are a lot of good reasons relational databases have been the de-facto default over the last four decades.
Let's scrap everything that was invented in the 70's or before.
But recently, I've been exposed to a fairly big and complex SQL one with several references between entities and lots, lots of X_has_Y tables. This makes me think that with growing complexity (which seems to be a general trend), NoSQL databases seem more practical at some point, or at least something less rigid than classic relational ones. I'm not saying SQL is obsolete, but it seems like its domain of usefulness is shrinking.
Meanwhile, Postgres was quietly plugging away adding new features and continuing to deliver solid performance for a wider range of workloads, including better performance on JSON document storage.
Even at its best, there is essentially no reason to choose MongoDB over Postgres with JSONB-type columns. They are essentially the same data model but Postgres gives you better guarantees of data consistency, plus a forward migration path to relational data when the day inevitably arrives when you need to model relationships between entities.
At this point Postgres is where most open-source RDBMS development work is concentrated. It's not only a solid codebase, it's piling up features pretty quickly and there are relatively few niches it doesn't fill at least adequately. All of these niches are covered by some commercial products built on top of Postgres (eg EnterpriseDB or CitusDB). It's pretty much a one-stop shop for application development. You can use it for everything from GIS to machine learning [0] pretty efficiently, and it pretty much will just do the right thing without you watching.
NoSQL really fits best around the margins, like as an auxiliary system for analytics. There is really almost no use-case where "user inputs data and we lose it" is an acceptable application behavior, so consistency is a business requirement for your master database whether you realize it or not. And consistency across a distributed system is hard so it almost always makes sense to sidestep clustering until the last possible moment. Buying more machine is cheap, replication/failover is a lot easier than consistency between distributed masters, and if you are really up against the wall there are those commercial products that can do this with Postgres.
If you want to make an analogy... Oracle is the suit, Postgres is the hardworking small business that is slowly but surely eating up Oracle's lunch, and MongoDB is a trustafarian with a hot-dog detector app. And that's why there's a lot of resentment towards MongoDB.
[0]: The 9.x series and 10.0 release have been absolutely jam-packed with new features, it's absurd how fast development is moving at the moment. One of my favorites... indexed cube queries. A cube is a data-cube type, an N-dimensional cube of data. One feature of this is distance queries, which have obvious applications in pattern recognition tasks (eg k-nearest-neighbor). One of the features in 9.6 is index functionality for these, so you can now do indexed KNN searches on your data...
https://www.depesz.com/2016/01/10/waiting-for-9-6-cube-exten...
I'd say it also fits well in two niches: document datastores (so long as there's some JOIN support, via referencing nested documents vs direct nesting) and graph stores.
I remember 10+ years ago working on storing nested sets in the RDBMS and it wasn't pretty. And the RDBMS schema for Magento 1, with key-value tables all over the place which NoSQL would have removed the need for.
Which has its own problems. PG does this just fine, with a full battle-tested relational system to back it (and you) up.
> with key-value tables all over the place which NoSQL would have removed the need for.
Product X having a stupid schema is not a good basis for an argument for or against a particular product.
What other way could Magento have implemented user-defined columns at the time, using a RDBMS? In 2009 when MongoDB was released, JSON columnstores were something to dream of and the alternative was storing serialised data in a BLOB field. That "stupid schema" did not have an alternative I can think of, except NoSQL.
Postgres supports hierarchical/nested structures using the "ltree" column type. There is nothing stopping you from defining a primary key of (eg) "set1.set10.set100". There is also support for recursive views/etc to operate on these kinds of sets.
Again, if you have some kind of "sparse" column, it can make sense to put that into a JSONB column. This is effectively the same thing as attaching an unstructured document to a record for this use-case.
Joe Celko popularised them. I've been unable to find when they were first introduced; but a search of my source code archive points to having written one ~14 years ago.
I'm pretty sure that mmap was the only storage engine available for MongoDB for most of the hype period.
I bet they are still waiting for that join....
If I’m doing analytics on a time series this gets even better with partition pruning and hash joins or bitmap indexes. And if I have a columnar database, that blows up this whole complexity argument.
My point is that, layout and indexes should never be assumed to be one size fits all. Keep in mind if I’m doing analytics I want to bring my time series data into memory at least as a stream, as I need to calculate / filter / transform the records. Not everything is about rendering a page on a website.
Interestingly enough, many of those techniques are common in the NoSQL world as well — a billion records is enough to require thinking about data flow anywhere — but the difference is that you have to deploy them more frequently.
Less rigid than what? The schema is going to exist somewhere...
Most other models are for high performance of specific access paths and punt integrity management to code, which is a throwback to the 70s. They’re glorified file systems and data structure caches. Mongo has a reputation for losing your data.
There are cases where scale and availability kills you and you need something like Cassandra or whatnot. But there are few as flexible and general as Oracle or Postgres.