CouchDB vs. MongoDB
blog.panoply.io
blog.panoply.io
It fails to mention that CouchDB now has Mango, which is a MongoDB-compatible query language.
Since 2.0, CouchDB also has Dynamo-like clustering thanks to Cloudant's open sourcing of the BigCouch code.
I wonder if the MongoDB side of the comparison is more up-to-date, or equally stale.
The last part I was worried about is the self-contained bit, so I'm going to give CouchDB a try, and use Mango to prevent redoing everything
Apparently their new replication protocol isn't as utterly broken as the old one, but my understanding is that data recovery after server crashes is still a problem.
At the end of the day (see my other comment in this thread), I'm bullish about the developer experience of Mongo/Meteor, and I think that 363 days a year of significantly increased developer productivity is worth the 1 day of devops hell that might ensue from that stack choice, and 1 day of realizing that your performance problems all stem from Mongo query performance being much more reliant on manually creating indices on fields vs. an unindexed Postgres table of the same size. Sigh. On the reliability end, it helps that our (largely append-only) data model is such that inconsistencies are possible to clear up manually if needed. But on those couple days a year, I'm definitely not a Mongo fan.
I've dealt with many apps and teams that use Mongo, and the amount of time spent tracking annoying bugs because of a lack of schema, writing migration scripts and the abysmal performance of the database leads me to believe that the only reason Mongo devs think they are more productive is because they can get 3 lines of code to make a new collection + record in the database in 5 seconds, and ignore the 5 months of time they spent down the line.
MongoDB 3.4 seems like it has a lot of great stuff and I hope to move to it soon.
Meteor actually provides exactly this for MongoDB; it has a "minimongo" package in the browser that supports Mongo's query language, running it synchronously against an in-memory copy of the collection [0]. And with Meteor, you can specify "subscriptions" declaratively that enable bidirectional synchronization while their owner components are in scope.
Mongo certainly has some reliability issues (see other comments here) but I've yet to find a full-stack system so painless to develop in, especially if you need realtime support. With things like ToroDB Stampede [1] and a general approach of "write all your code in React with Meteor dependencies factored out into containers," there's a clear migration path towards the relational-based separate-backend-frontend world when you need to go there.
This is the main feature I sell when pushing CouchDB.
Use it to project events and you'll see what I mean.
I recommend anyone shopping for databases with easy master/master replication with eventual consistency and no single points of failure to consider Couchbase as well. It's not the same will have it's own set of pros and cons.
The author clearly has a different definition of "strong consistency" than most. I don't see how any claims of consistency (in a data usage, not CAP sense) can be made of a database that can't properly store a number as a number or even guarantee that it's a number at all.
Also, does anyone actually like the Mongo query language? It was cute when I first saw it but I pity anyone trying to do anything complicated by manually writing those JSON strings.
And yes in a schemaless database you are required to manage the schema in your application layer as opposed to within the database. If we wanted a database with a rigid schema we would just use a SQL database.
And my point is that if you query a single document and the field is a number then it will be returned as a number i.e. it physically stores and understands numbers.
MongoDB allows a free form schema for every document. You're again comparing apples and oranges.
What do you mean?
Couchbase-Lite: https://github.com/couchbase/couchbase-lite-ios https://github.com/couchbase/couchbase-lite-android/
Cloudant Sync: https://github.com/cloudant/sync-android https://github.com/cloudant/CDTDatastore
We've had a good experience with CouchDB with an Android native app, iOS native app, and a web app (using PouchDB with it there).
Have you ever had an issue with conflicts, where multiple instances of the app read, modify and write different things to the same document at the same time?
This problem is what pessimistic (select... for update) or optimistic (using a version column) locking is for. If you don't want any race conditions to sneak into your code, as a rule you should probably be using one or the other regardless of whether or not you use postgresql JSON.
CouchDB and Couchbase both support only whole document updates. So you get document conflicts in those document stores as well, meaning your app needs to understand and handle 409's. But those conflicts are relatively easy to handle in most cases, at the cost of a new round trip. Mostly it's a matter of downloading the new document state and merging your change to it and re-post. If you're using Redux/Vuex/Event Sourcing this becomes trivial to support. Another way to handle it is to split a single large document into smaller pieces and write a map/reduce view that returns a composite document. That should be possible in Postgres as well with a prepared statement.
From my experience MongoDB is fast, but CouchDB really shines when you have a read heavy application.
Also the article didn't mention Mango queries, which is a blessing (fast indexing as erlang views), but in my opinion this feature can be a lot better with stale results, for instance.
I do see the point of storing documents rather than rows for some use cases, or dbs extra strong on searches, etc.
But what is the problem with an "add column" or "change datatype" operation is sql..?
Many other engines, PostgreSQL for example, can add a new column without constraints as a nearly instant metadata change only. Data type changes that do not require validation (expanding a VARCHAR vs CHAR to INT) are also rapid.
In my experience, "schemaless" just means that I'll have to manage the schema manually through some "updater" scripts.
Any data without schema is just noise, so if you have any data, it means that you have schema for it somewhere.
EDIT: I made a statement about our preference, for our environment. I did not make broad claims about these DBs for other people. How am I upsetting HN?