MongoDB Is Raising Another $100M
techcrunch.com
techcrunch.com
It really depends on what you need... there are situations where I would recommend MongoDB, RethinkDB, ElasticSearch or Cassandra... The farther you get from traditional RDBMS, the more you have to consider your needs and any trade offs.
The whole point of master-master is not dealing with failure situations or trying to scale out performance (even though master-master handles those situations naturally), the point is that it's a different philosophy. The idea is that there's no single official true state of the database. Reality happens to map that idea very well. The data doesn't exist in a centralized official place, it exists in multiple places which may not be perfectly connected in real time all the time.
The data exists at multiple servers at multiple databases. It exists on your mobile device, sometimes disconnected from the server. It exists on thousands of browsers on the same time. There's no single master.
The best mapping of this reality is the master-master ideology. Yes, it requires application-level merges of data, but the benefits you get from implementing this can be tremendous. And that's why I'm excited about using CouchDB. Or rather I'm excited about master-master databases. CouchDB happens to be one of those, and I hope future databases will be built around that too.
ElasticSearch is fairly close, but I'm unsure if it will scale as well in practice. ElasticSearch handles the middle-to-high ground very well. I'm not sure I would use it for authorative data... It works incredibly well for logging (with logstash) and the front end utilities (kibana, etc) are nice too. It's primary use is as a search, if you are very read heavy for searches, or write heavy for analytical data it works well. You can tune your use to separate the storage and reads in interesting ways.
MongoDB works incredibly well when your core data is composed of mostly self-contained documents and in need of certain flexibility. A typical classifieds website is a great use case. It's also a very natural fit in a lot of programming languages, js (node) in particular is a natural fit. There's much less disconnect between the data and your application models.
RethinkDB is very similar to MongoDB, but has a more traditional mindset when it comes to it's use. I think the programming interfaces are a little better thought out, consistency and data security is at the forefront here.
In general it comes down to... very large loads, use cassandra... easiest use is mongo... rethinkdb I've been waiting on the replication story to get better, and the geosearch support is fairly recent... ElasticSearch search or for logging (write heavy). In many scenarios where I would use ElasticSearch, it would be along with an RDBMS as an authoratative data source.
That being said, CouchDB does have a nice built-in in gui.
Cloudant is doing some very interesting things relating to querying on a CoudhDB system. They created a query syntax so that you have an option besides map/reduce [1]. I watched their webinar on it and it seems pretty slick, but I have not yet played with it. Also, it is not part of CouchDB yet so there is no option to run tests against locally to verify syntax, response errors, etc, etc.
Cloudant is working to merge many of their changes from BigCouch back into the CouchDB project, so one day I expect CouchDB to have multiple query options [2].
[1] https://cloudant.com/blog/introducing-cloudant-query/#.VLCSc...
[2] https://cloudant.com/blog/update-from-nebraska-the-cloudant-...
1. Indexes are built and appended to at read time.
2. If you change the view code an index rebuilding will occur. (A hash of all the views in the design document is taken and compared with the previous hash. So even a small thing such as adding a space will in effect change the view code and trigger rebuilding of views.)
3. No change to any other part of design document will have any effect on views.
I wrote in detail about it here (http://staticshin.com/programming/does-updating-a-design-doc...)
Which means that you can write an erlang module and call it with couchdb show functions and have the benefit of a ready made http api for you.
A practical use of this would be to use mnesia as an in memory key-value store(for similar things that you would use redis for) and then call the functions in the module from a show function. All the benefits that you get with erlang you get with couchdb.
One of the killer features (for me at least) is user accounts management. In 3 http calls I have login-logout facility in my app ready. I can't tell you how impressed some of my clients are when I show them a V1 of their product in a couple of days.
To make matters worse, the used JS engine is fairly old and the exceedingly wordy documentation is kinda hard to follow.
I'm a lot happier with RethinkDB.
CouchDB 2.0 adds dynamo-style clustering support (similar to Riak and Cassandra) in addition to the replication(sync) protocol used by the CouchDB ecosystem [4]. It also includes all the fixes / performance improvements based on Cloudant's operational experience over the last 5 years.
That said, document databases are certainly a niche. CouchDB is a good choice if you need strong durability (writes are always fsync'd - to multiple copies when clustering) and consistent performance as your database scales (querying options may seem restrictive but are designed to scale very well). As others have pointed out, the RESTful interface to CouchDB makes it a good fit for web and mobile applications which can query the database directly without an app tier. PouchDB [5] and various mobile datastores which implement the sync protocol [6][7] allow you to also take your database to a browser or mobile client and work with it when disconnected which is pretty compelling.
[1] https://speakerdeck.com/wohali/putting-the-c-back-in-couchdb... [2] https://cloudant.com/blog/introducing-cloudant-query [3] http://blog.couchdb.org/2014/12/19/couchdb-weekly-news-decem... [4] http://www.replication.io/ [5] http://pouchdb.com/ [6] https://cloudant.com/cloudant-sync-resources [7] http://www.couchbase.com/nosql-databases/couchbase-mobile
Do people actually consider using MongoDB for new projects? Do they want to add it to existing infrastructure? Why, with only a cursory search on the limitations?!
We use other hosted SAAS services for analytics (MMS, Google Analytics, a few others) because well they are cheap/free and that isn't our core strength.
When I work with people who claim that their models are not relational, I usually have to contend that they are. The argument goes like this: you may be able to model your problem as documents or hierarchies, but can you model all of the questions you want to ask about that data in the same way?
The major vendors of relational systems have first-class support for hierarchical data structures, recursive data structures, graphs, KVs, and documents, and they can be used in conjunction with the basic relational features. Modern SQL is more that just SELECT...FROM...WHERE...GROUP BY; it has powerful, fast analytical functions, domain modeling, and reporting features. The top engines can partition your data and parallelize your access patterns to get the most value out of your commodity multi-core/SSD hardware.
The support for such systems is ubiquitous in todays software libraries. These systems even have tailored hardware platforms to support them if your problems really lie far out on the curve.
The downside is that none of the free/OSS systems are quite as capable. The commercial systems often require the top-tier editions to support all of the above.
The good news is that it's really a good financial deal if you actually need it. A $40,000 license for Oracle or MS-SQL is 1/4 of the annual cost of an engineer that can coerce similar functionality out of a lesser product. Their are plenty of consultants that can help you get there on a one-and-done basis.
PostgreSQL is getting there, too. Query parallelism is, for me, the biggest gap. There are some neat aftermarket solutions, but it's not quite there yet.
I'm not passing judgement on if that's a good thing or not, but many teams I see today are looking at the storage engine as nothing but that: a temporary place to put things that can be swapped out if there's something that does the same job faster, where 'job' is glorified K:V and possibly sorting.
Ceding advanced functionality to the database is what is being avoided: my own app code is usually easier to troubleshoot than an obscure Oracle error.
And EVERY database has limitations. You just need to be pragmatic and determine if you will ever really hit them.
Given this, why don't they just use Postgres?
If I am building an application with a document model (which actually isn't that niche for SPA sites) then MongoDB is a great choice. It is much easier to use and manage than PostgreSQL which is important if you are trying to get something off the ground. This is why I would use it. But that doesn't mean it's the right choice for everyone nor is PostgreSQL, Oracle, MySQL or any other database.
I've been watching hierarchical data models for some time now - I honestly can think of only a few, limited applications for something like MongoDB.
I'm not the only one who thinks so - see http://www.sarahmei.com/blog/2013/11/11/why-you-should-never...
Edit: Look, I know I'm being negative on MongoDB. But I really wanted it to be awesome, but it's inherent limitations are just so start that I feel that most people who use it are doing so because they are misinformed or misled. If you have found success with it, that's great. I just have very strong reservations about the entire data model for most people.
As someone has posted above: if you have data that is relational, then it's relational - you should use a RDBMS.
At no point have I thought so condescendingly as you do that MY choices are the best for everyone. They aren't. And I promise you that picking the one technology/approach for everything doesn't work anymore. It's a heterogeneous world out there.
And that link you posted is pathetic. MongoDB is not suited for social network style data neither is many other databases. Doesn't mean they are useless for every use case.
Ease up there buddy! I wasn't being condescending. I just think that for most people a relational model is probably what they are looking for. I don't think the article I posted is "pathetic", as that's a bit condescending... Just showing one datapoint that shows where people think they need a hierarchical data model, in fact they need an relational model.
That's quite true. My experience with MySQL comes from working with company intranet installations that didn't see that much traffic and some web apps, all with a high read-to-write request ratio. They all ran off a single DB server (some hardware, some VPS), so my administrative work was limited to automating fairly straightforward tasks with things like Ansible. I'm curious to hear an example of the kinds of problems you've run into with MySQL, since I assume you deal with more complex and larger-scale deployments. (And I'm looking for anecdotes to help persuade customers and developers alike to give Postgres a try.)
Mostly we were bitten by long-running MySQL sillinesses:
* its strange idea of UTF-8 (we call the MySQL version WTF-8 - see http://geoff.greer.fm/2012/08/12/character-encoding-bugs-are... )
* InnoDB's galloping disk consumption
* the Debian/Ubuntu package's default stupid behaviour of putting all the InnoDB databases into a single file ibdata1 (I hope this is a Debianism and not something that's default in upstream)
* ibdata1 never shrinking ever (bug #1341, open since 2003)
* binary replication issues (e.g. bug #68892, which is fixed but that doesn't help older or distro versions)
* several others I've mercifully obliterated the braincells that were holding them. But all of these were long-known issues that will never be fixed for one reason or another (backward compatibility with past mistakes, or they just can't be bothered).
Here's a good crib: http://grimoire.ca/mysql/choose-something-else
tl;dr MySQL: the Comic Sans of databases. Except Comic Sans has use cases.
One challenge that companies find as they try to scale is they have to add a lot of Sales and Marketing costs well before they receive any revenue, so there's a cost bump despite low growth in engineering. This is compounded if there's a professional services component.
You can see their test matrix here: https://mci.10gen.com/
and the driver tests here: https://jenkins.mongodb.com/
Also they run https://mms.mongodb.com/ which does backup and monitoring
Although the next generation is moving to ElasticSearch, we haven't had issues with our use of MongoDB (which is a very good use case for it).
I think you are thinking there will be an objective reckoning of features/performance/reliability and somehow since OSS wins all of those they will choose it, and as far as I am concerned, thats not actually why they choose their database in the first place.
A million b2b apps are developed on sql server because visual studio/microsoft makes that easy for their .net stack, and oracle sells to executives or other manager types and gets shoehorned into projects or set as a requirement before smart people get involved ALL THE TIME.
A lot of it still comes down to enterprise pricing, support, integration, and name recognition.
However, I am thankful my own bosses can count, and went "WHAT" at the last Oracle bill.
So we're actively seeking to move our own stuff from 'Orrible to PG, and to get rid of the vendorware depending on Oracle.
We just got AppDynamics in (ridiculously versatile and useful monitoring). Speaking to the AD sales engineer, he said a lot of their Oracle-using customers are eyeing up PG similarly.
I like PostgreSQL, and plv8 looks incredibly cool... when replication and promotion are in the box (not needing enterprise or other complex addons), It'd be my first choice for most situations.
(edited comment to make it less snarky)
Replication has certainly gotten much better in the past few years, though.
Mongodb literally has a 1-line command to add a replica set member and it works. This is still a huge advantage.
I don't care either way I just use RDS and don't worry about it. (You can add replicas in one click with RDS.)
And if you think that is as simple as MongoDB's replica sets then frankly you are crazy.
Wal-E was written for use in cloud which has its own challenges, but if you want to run in your data center the closest thing is Swift in OpenStack.
As for Mongo, my company is currently using it, and I wouldn't say it's any easier especially when you try to use it in public cloud, when you no longer have guarantees that the instance you set up won't be terminated and recreated in different AZ. There are plenty of challenges.
I'm currently working on convincing other teams to drop using Mongo and instead make queries directly to Postgres which is our authoritative source of data. The idea was to simplify our infrastructure, and I was not expecting much difference in performance but oh boy. In all of my POCs that I did so far PG is beating Mongo that makes you feel sorry for it. Both in performance (you need to understand what you're doing and use right types, indices and queries) and data size (after moving, the data is much smaller so it no longer requires being distributed, and also the instances can be much smaller).
Perhaps you can point to a relatively simple walkthrough to setup PostgreSQL with plv8 for replication (with an easy promotion of a slave to master), that doesn't take a commercial support license...
Everything I've seen seems incredibly convoluted and more difficult than say MS-SQL, MongoDB, RethinkDB, ElasticSearch or several other databases at data replication and spinning up new nodes, or handling a primary failure.
Also, most of my work with Mongo has performed very well, if your data is a good fit, which I will admit it isn't all a good fit. Honestly, I'd rather use pgsql with plv8 over mongo, or elasticsearch, or ms/azure-sql... The support costs and my time are important to me, and better spent working on architecture or development. If I can make an operations level decision that works well enough, or is easier to scale then the development time is almost a wash.
If you need scaling in MongoDB, you're looking at replica sets combined with sharding. Not all workloads, as you rightly point out, need to scale from the get-go, but there are an awful lot that need HA.
Who are you to lecture others on what is/is not important for their needs ?
Don't forget that a lot of people have realized that MongoDB is NOT a good fit for what they want to do. Then they have a resourcing problem - a big one.
And there are plenty of people who are switching away from PostgreSQL and other SQL databases to MongoDB. Again it is quite a popular database (hence the huge amounts of cash they seem to be easily raising).
i think that's the nub of our difference of opinion. i saw that as a ruse, i don't genuinely believe that the presenter never heard of them.
[0] - http://vimeo.com/2723800
Disclaimer: I used to work there.
http://www.postgresql.org/docs/9.4/static/datatype-json.html
You have the flexibility to store arbitrary JSON blobs when you need to, the stability, maturity, and performance of Postgres, and the ability to migrate data to a more rigid schema once your project matures to the point where data validation is more important than raw prototyping speed.
The main drawback so far is that the query interface is a little clumsy.
You'd be surprised how much MySQL can do and how great the documentation is.
Disclosure, I have recently joined up with them as a Solutions Architect.
Things like ElasticSearch (search database), Neo4j (graph database), and Redis (key/value store) seemed to be used along-side a traditional RDBMS, and have specific use cases that make them superior than trying to shoe-horn the functionality into a traditional RDBMS.
We retrieve documents using the unique id generated in mysql table.
Add arbitrarily complex (json) structures to a table, without a database migration.
With are relational model, your working set needs to be joined together from pieces, so you want ACID to ensure that everyone sees a consistent set of pieces.
But with a document model, you can 'pre-join' your working set into a single object that has everything you need. And that object doesn't have a fixed schema, so it can grow and evolve over time.
While Mongo doesn't provide ACID across documents, it does guarantee atomicity, consistency, isolation and eventual durability for SINGLE documents.
IF your application can live with a universe that consists of a single, arbitrarily complex object - then Mongo is as within epsilon of being as safe as a regular ACID transaction system.
I have 2 MongoDB mugs. I could use 4 more for a nice set.