CouchDB 2.0
blog.couchdb.org
blog.couchdb.org
If you want to make my life a tiny bit easier:
- blog.<project>.org needs a link to <project>.org because half of you visitors want to read the "it" before they read the "news".
- github.com/<you>/<Transmogrifier>For<ThatProject> should have a link to <ThatProject> in the description or the readme's first paragraph. I often come across interesting plugins without knowing anything about projects' they are for.
(and yes, I know there are actual problems in the world. But these are easy to solve)
- Clustering https://blog.couchdb.org/2016/08/01/couchdb-2-0-architecture...
- New Query Language https://blog.couchdb.org/2016/08/03/feature-mango-query/
- New Admin Interface (written in React) https://blog.couchdb.org/2016/07/27/fauxton-the-new-couchdb-...
If you want filtered replication, we designed Couchbase Sync Gateway because we thought db-per-user was to heavy. What's fun is to think about the options for mix-and-match across the stack.
It would be cool to read more about where "replication apps" are being used and how.
I'm sad that the Couchapp thing never took off or is getting de-emphasized. That was one of the really brain-bendy ideas about Couch that I loved, that these DB-side applications could also replicate to other people.
I'm only just catching up with Couch and finding the db-per-user stuff. Do you know of anyone shipping a Couch instance inside thick clients, like a traditional desktop application?
I don't understand how projects like remotestorage, for example, aren't using Pouch and Couch and instead try to develop their own replication protocol without success for many years.
BTW, you are spot on about continuous replication. Try using continuous replication on Cloudant and you will end up with a fat bill.
Cloudant recently changed its pricing structure [1] (via IBM Bluemix) which should significantly reduce the cost of continuous replication - that may be worth a look if the current model is problematic.
[1] https://developer.ibm.com/bluemix/2016/07/26/cloudant-public...
Of course there is loads more, but this is the stuff that keeps us devs going.
Yes, the server probably runs CouchDB, but it could also be Cloudant or CouchBase etc. None of what is mentioned there is a USP of CouchDB itself.
That doesn't mean it's the best technology but it got me using it. It has been an enjoyable enough experience, not a lot of complaints. A better querying interface and the addition of an update operation would go a long way but those both seem to be included with this release.
I might even prefer it.
> The replication protocol, which supports multi-master, has changed little
Basically true, though with many interoperating, independent datastores it's a tricky thing to evolve. 2.0 adds an additional endpoint, _bulk_get which can significantly reduce the number of requests when CouchDB is paired with an on-device database such as PouchDB, Cloudant Sync or Couchbase Lite (the endpoint was inspired by the same feature in Couchbase). The CouchDB replicator itself has had many performance and stability improvements [1] and continues to be a significant focus for active development.
Also, CouchDB 2.0 introduces internal cluster replication using distributed Erlang. If you currently use CouchDB replication to bi-directionally replicate between machines on the same network for HA, replacing those with a CouchDB 2.0 cluster should be a big win.
> In other words: everybody seem to be looking at CouchDB as just a very poor and limited MongoDB.
Query doesn't pretend to be MongoDB-compatible - it provides a syntax that should be familiar to MongoDB users and more query flexibility than views allow. I think Query still has a fair way to go - this is the first release - but it's a move in the right direction.
As to whether CouchDB is viewed as "a very poor and limited MongoDB", they are very different databases. CouchDB is a good choice if you want a rock-solid JSON datastore which comfortably scales up to multiple TBs / many machines, with multi-master replication over unreliable networks. Query support, as you say, is not as rich as some other databases, so if that's more important to you, there are probably better options.
> Filtered replication was implemented, but it is slow to the point that no one recommends that you use them.
The new _selector filter [[2] in CouchDB 2.0 offers a significant performance improvement for filtered _changes. It should be a small change for replicators such as PouchDB can take advantage of this.
> About Couchapps, the special database features that powered them in the first place were left aside
I don't speak for the project, but it seems there has been much debate about this in the CouchDB community and the conclusion was that there are better solutions to most Couchapp-shaped problems than running application logic in the database. The features that combine to enable Couchapps haven't gone away and will benefit from the general improvements in 2.0, but they haven't been explicitly developed.
[1] https://blog.couchdb.org/2016/08/15/feature-replication/ [2] http://docs.couchdb.org/en/2.0.0/api/database/changes.html#s...
The folks working on cockroachdb recently did this[1] and it was a good read.
[1] https://www.cockroachlabs.com/blog/diy-jepsen-testing-cockro...
So, yeah, I think even if eventually consistent, and not immediately consistent, they want you to feel they have taken care of your data when designing the thing, so seems prudent to do.
I am particularly really excited to see the announcement of the Mango query language, since querying Couch was one of lesser-easy things to do back then. I'm also very excited to hear about performance improvements, as this has been particularly interesting to me as I've been tracking various system's performance as I have worked on our own (Mongo with Wired Tiger, Cassandra, Redis, and even Chrome V8 engine as we have built towards 30M+ ops/sec, see https://github.com/amark/gun/wiki/100000-ops-sec-in-IE6-on-2... ). However clicking through on the performance links didn't lead to any numbers or benchmarks. I would love to see that!
Really happy that Couch is on the homepage of Hacker News. I really feel like they made lots of correct decisions that got passed over by the NoSQL craze, and have lately not been receiving the type of attention as it should compared to (what I biasly think) unfavorable but hyped up Master-Slave systems. People should really check into Couch's Master-Master replication!
Fetched and unpacked the source tarball, then did mostly what the INSTALL.Unix.md said to do;
sudo pkg install erlang icu spidermonkey185 gmake gcc curl help2man py27-sphinx
./configure
gmake release
sudo pw useradd couchdb -u 5984 -c "CouchDB Administrator" -L daemon -s /usr/local/bin/bash
sudo cp -R rel/couchdb /home/couchdb
sudo chown -R couchdb:couchdb /home/couchdb
sudo find /home/couchdb -type d -exec chmod 0770 {} \;
sudo find /home/couchdb/etc/ -type f -exec chmod 0644 {} \;
Note how I set the uid to 5984 ;)Unfortunately, when I try to actually run it...
sudo -i -u couchdb ~couchdb/bin/couchdb
...I'm just getting messages like: [error] 2016-09-20T23:57:55.321596Z couchdb@localhost emulator --------
Error in process <0.547.0> on node couchdb@localhost with exit value:
{database_does_not_exist,[{mem3_shards,load_shards_from_db,"_users",
[{file,"src/mem3_shards.erl"},{line,327}]},{mem3_shards,load_shards_from_disk,1,
[{file,"src/mem3_shards.erl"},{line,315}]},{mem3_shards,load_shards_from_disk,2,
[{file,"src/mem3_shards.erl"},{line,331}]},{mem3_shards,for_docid,3,
[{file,"src/mem3_shards.erl"},{line,87}]},{fabric_doc_open,go,3,
[{file,"src/fabric_doc_open.erl"},{line,38}]},
{chttpd_auth_cache,ensure_auth_ddoc_exists,2,
[{file,"src/chttpd_auth_cache.erl"},{line,187}]},
{chttpd_auth_cache,listen_for_changes,1,
[{file,"src/chttpd_auth_cache.erl"},{line,134}]}]}
[notice] 2016-09-20T23:57:55.321650Z couchdb@localhost <0.323.0> --------
chttpd_auth_cache changes listener died database_does_not_exist at
mem3_shards:load_shards_from_db/6(line:327) <= mem3_shards:load_shards_from_disk/1(line:315) <=
mem3_shards:load_shards_from_disk/2(line:331) <= mem3_shards:for_docid/3(line:87) <=
fabric_doc_open:go/3(line:38) <= chttpd_auth_cache:ensure_auth_ddoc_exists/2(line:187) <=
chttpd_auth_cache:listen_for_changes/1(line:134)You'll have to create these tables:
* _global_changes
* _metadata
* _replicator
* _users
It's a bit of a pain!
Well, "moving over" might be a slight exaggeration seeing as how my little project is in such early inception that all I had prior to switching it over to CouchDB 2.0 was a few SQL-files defining the schema of various tables, as well as a couple of other SQL-files with queries to run and a bit of sample data. But anyway, the point still stands that said little project is now using CouchDB.
When I started with CouchDB it wrong choice for so many reasons, client had <30gb of data, couchdb was cooler than node.js, and I was frustrated with SQL Server. In hindsight sticking with SQL Server or Postgresql would of been better - older/wiser today.
With clustering you now get that. By way of “oversharding” even on a single node. Speed up is linear with number of shards / CPUs
> - space reduction,
2.0 has the better compaction format. There are still ways to improve, but we are getting there.
> - ES6 or even ES5
Our custom wrapper around Spidermonkey 185 is getting long in the tooth. It didn’t make it into 2.0, but we are well aware that this needs updating.
Anyway, sounds like we got there in the end ;)
A spiffy log package was the first thing I thought of and indeed that's one idea listed on this page: http://docs.couchdb.org/en/2.0.0/intro/why.html
But I think the "logging" case is hard to argue vs. "manually" grepping (or ag-ing using the silver searcher) over in-place log files and aggregating/rendering dashboards via static files.
Another good use case I had was for JSONP request for an auto-complete input field on a webpage. Again, the database was publicly readable.
I have also used it to aggregate data for graphs. The data changed daily, so the ability to cache the results of a view until the data changed was nice. But the first request of data each day still took a while. I don't think I got a performance boost in this case, but I did get free caching.
All my other uses of CouchDB were mostly for fun and could've been implemented in traditional SQL.
Looking at the online draft version of the book, the tour chapter [0] still has version 0.10.1 in the example.
I wonder if there will be a second edition of the book covering mango, clustering and the new admin interface.
Just a heads up, Damien didn’t work on that.
The admin interface of 2.0 will probably be familiar since I've already learned my way around older releases. Nonetheless, I think that there is no question about the fact that if there comes a second edition of the book that is going to cover 2.0, and seeing as how the first edition has some stuff about the admin interface, that's obviously going to need to be updated, no question about it.
So when I buy the next edition of the book, I will probably just skim through the admin interface stuff, but not all readers will have used a previous version of CouchDB, so to some of them, having that covered and up to date will be valuable like it was to me the first time I wanted to know about CouchDB.
So if you think I'm saying that I need the book to learn the admin interface, I think you did not understand what I meant, because that is not what I meant. I meant that I want to learn about mango and clustering, and that I also expect the next edition of the book to have the parts about the admin interface updated.
I was, however, under the impression that you actually needed a book for a lot of things and I didn't like that feeling.
However, now that you explained, I understand how good can a good book be. I also read that CouchDB book almost entirely before going on to use the database, before thinking about it, or having a need for it, and that made wonders to me. I had forgotten about this fact of my life. Thank you for reminding me.
Nice name :-)
I haven't used Couchbase, but I understand besides the "Couch" prefix and that CouchDB's original author working there for a few years, it doesn't have much in common with CouchDB project.
CouchDB 2 is 99% API compatible with version 1.
Whilst I really like couch I'm just not convinced on using it in production due to its glacial pace.
MongoDB makes a lot of sense to folks already familiar with relational databases. Collections are like tables, documents are like rows, and it's easy to use document IDs to establish relations and perform queries like you would in a relational database. There are a lot of problems with using it like a RDBMS, but still a developer with zero knowledge of MongoDB can become productive with it very quickly.
When learning CouchDB, the first WTF moment is when you realize there are no tables and that all the documents regardless of the type of data they contain go in the same place. Eventually you learn that you can add a type field to the documents to distinguish between them, which does feel hackish. The next WTF moment is when you want to query the data and realize you have to use map-reduce to do what would seem trivial in any other database. The early version of the admin tool definitely didn't make this easy since it required writing a JavaScript function and escaping it so that it could be stored as a single-line JSON string. It also didn't help that the map-reduce code was stored in special documents that used magical document IDs to distinguish them from other documents, which again feels hackish but makes sense eventually.
That said, I am a big fan of CouchDB, and I hope that with the query language and new UI in 2.0 that CouchDB will earn the respect it deserves.
I am looking forward to see new features for sure.
I love the CouchApp idea because it's a natural extension of "save whatever you want." As long as we're saving json, why not save static files too? It lets you prototype quickly and, in our case, let us develop and test out specific features for clients by quarantining risky crap code on the fly into one database. I like to think we're beyond that stage, but at the time it was invaluable to keep multiple versions of our app running concurrently.
I have 3 solid years of experience working with CouchDB and in the end I find myself longing for PostgreSQL with one bson column for the unknown or "volatile" attributes. Our data was semi-relational, as I would argue is most data. Meaning there were honest to goodness has-many or belongs-to type queries that could have been simplified and easier to maintain outside the application code.
I know that's just a style, some people would use key-value stores for everything with no validation on anything. If you're not that far gone and you still like freedom then CouchDB might be for you. As for me, I wanted an error from my database if my application code asked for something or tried to store something invalid. But for rapid prototyping for us I don't think we could have gotten a better database.
It is not enough for actually relational data, but surely you know that CouchDB has validate_doc_updates functions specifically designed to check user rights and document structure ?
http://docs.couchdb.org/en/1.6.1/couchapp/ddocs.html#validat...
It promised master-to-master replication, so has that. Doesn't lose your data (actually fsyncs, yay!). Has a helpful integrated web interface, is HTTP + JSON interface.
I don't know I've shipped lots of products on it. I say it lived up pretty well to the hype...
I used it for a variety of projects and seem to love it.