PouchDB, the JavaScript Database That Syncs
pouchdb.com
pouchdb.com
Getting it to work inside React Native was initially very challenging. Keeping our shim up to date with the recent changes to PouchDB has also been challenging. We are currently using PouchDB 5.4.5 because there was a breaking change in 6.x and I haven't had a chance to dive into it to figure out what is going wrong.
The PouchDB community (especially Nolan Lawson) does a great job of showing examples, answering questions, and responding to feedback.
Cb lite was complicated since it needs a native plugin rather than being just js. That wasn't too bad.
The thing that put the nail in it for me was all the setup on the server side with sync gateways, multiple databases that then sync with one another, etc etc.
couchbase-lite apparently also works with couchdb (protocol compliant) and cloudant.
I am currently working at a gig where the use of couchbase-lite with sync gateway and couchbase are required for a project.
I've upvoted davidbanham's comment because that seems like a more insightful answer to your question.
1) Realm works only in React Native and not in React web. There is a regular web component to my project as well and the thought was we could reuse our knowledge of PouchDB across mobile and web. This has paid off in the prototype mode because it allowed us to build some simple reports in our web product. However, this is probably not the right architecture for reporting and we will likely change it soon.
2) Realm was very new and I felt we were already using too many new frameworks. So it was added risk on a tight timeline to get a prototype out the door.
3) On the server side, CouchDB has Futon to simplify our learning curve and give us a basic GUI to have a sanity check when our code wasn't functioning as expected. Not sure what Realm's server side looks like.
I'm sure I can come up with other reasons. But definitely want to check out Realm sometime. Looks like it is shaping up quite nicely.
I hope one day soon, after already 10 years of iPhone, we finally get a sane development environment again - with stable APIs, responsive UIs, portable code, easy networking and persistence, but at the same time fast to prototype and hook up. All the solutions I know are around 70%-80% there, there's always something that hurts badly.
That's what I meant with cost.
Edit: Btw. thanks for the link.
In terms of architecture we have about 250 tenants with separate Couch databases per each. We're still running Couch 1.6. We have yet to evaluate Couch 2.0.
It's been mostly smooth ride for the most part but this being a very unusual architecture we had to tackle few interesting problems that came along.
1. Load times. Once you get over certain db size the initial load time from clean slate takes ages due to PouchDB being super chatty. I'm talking about 15-30 mins to do initial sync of 20-30mb database. We had to resort to pouch-dump to produce dump files periodically. That helped a lot. I think this issue has been rectified with Couch 2.0 and sync protocol update.
2. Browser limits. Once we hit the inherent capacity of some browsers (namely Safari on iOS, 50mb) we had to get creative. Now we're running 2 CouchDB databases for each tenant where 1 has full data and the other only contains last 7-8 days. Pouch syncs to the latter one. We run filtered replications between the full db and the reduced db and do periodic purging. On the client side if a customer tries to go back more than 7 days we just use the Pouch in online only mode where it acts as a client library to remote couch and doesn't sync locally.
3. Dealing with conflicts. This might matter or it might not depending on the domain but you have to be aware of data conflicts. Because CouchDB/PouchDb is eventually consistent multi-master setup and you will get data conflicts where people update the same entity based on the same source revision. PouchDB has nice hooks to let you deal with this but you have to architect for it.
4. Custom back-end logic. Because Pouch talks directly to Couch you can't exactly execute custom back-end logic when needed. We had to introduce a REST back-channel to make sure our back-end runs extra logic when needed.
5. We had some nasty one-off surprises. Last one was with an object that had 1700 or so revisions in couch and once it synced to PouchDB it would crash the Chrome tab in a matter of seconds. Due to the way PouchDB stores revision tree (lot's of nested arrays) Chrome would choke during JSON.parse() call and eat up memory until crash. We resolved this one by reducing the revision history limit that is kept.
According to your 3rd point on conflicts, could you shed some more light on:
>PouchDB has nice hooks to let you deal with this but you have to architect for it.
I think I remember this issue (I was formally a heavily contributor to PouchDB) I think Nolan ended up writing a non recursive JSON parser to deal with this and there was some debate about whether it made sense to be used as it was significantly slower (though could handle deeply nested structures)
Unfortunately the only way to resolve this without vuvuzela would have been to change the structure of the stored documents which would have required a large migration, so I'm glad to hear that the vuvuzela solution was the right way to go.
For pluggable storage engines our node.js adapter uses leveldown, so any *down backend can be plugged in. https://pouchdb.com/adapters.html has some more information about this.
For p2p yup its been something quite a few people have been using pouchdb for, one of the more notable examples has been http://thaliproject.org/, the core storage format is entirely compatible with p2p (inherited from couchdb)
https://pouchdb.com/faq.html#sync_non_couchdb
Basically, your other data base probably handles conflicts in a different way to Couch that would make syncing this way nonsensical.
If you're happy to handle that yourself, the Couch replication protocol is well documented and there are plenty of libraries written for it.
PouchDB is a CouchDB implementation in the browser. I don't think your sync with other db would be easy and I believe there is no such desire in the project roadmap.
From how I understand it, a local database will persist until the browser clears its cache. What happens in the situation where the cache clearing takes place while you are using the pouchdb database? Can that condition be handled?
I thought pouch might be a good fit for a web app you could use offline - you'd need connectivity to login or whatever, and sync the database initially. Then you'd modify the local one and sync it when it's all done, saving round trips to the server.
But that also raises the question - on the server side, is one database per user feasible? IIRC Couch can only handle 100 or so different databases on one instance. And you can't do views across them.
You could also encrypt data.
This sounds _nuts_ from a traditional database mindset, but works great in Couch.
You can then create another database that replicates from all the user databases in order to perform your aggregate queries on the back end.
Sounds weird coming from my Postgres and MySQL background, but works great in practice, depending on your use case and if you clear out old revisions and unused docs, which for our use case can number into the thousands fairly rapidly.
That sounds horribly space inefficent.
Everything in software is a trade-off. This trades space efficiency for multi master replication with first class offline app experiences.
For many cases, that's a fine trade. Disks are cheap. If that isn't a fine trade for a particularly large dataset, use a different technology that's better at space efficiency and worse at other stuff.
There are actually several layers of "cache" inside of a browser, including the traditional HTTP cache (which is global) as well as the site storage, which includes stuff like IndexedDB, WebSQL, LocalStorage, AppCache, and window.caches (all of which is per origin).
This site storage _can_ be cleared by the browser (it's "temporary" per the spec), but in practice it isn't very frequently cleared unless the machine is running low on space. E.g. Chrome only does it if a site exceeds 20% of total per-origin browser storage (https://developer.chrome.com/apps/offline_storage) whereas Edge is extremely conservative with clearing IDB because it's considered user data (e.g. email drafts in Outlook). In any case, when the browser does clear this storage, it clears everything at once for that origin, so the user essentially has the experience of visiting the site for the first time. This is why it's a good practice to periodically sync your PouchDB data to CouchDB because it can be lost in rare cases.
Also there is a new Storage spec that allows site authors to designate certain buckets of per-origin storage to be persistent, but this typically requires a user permission and isn't widely supported yet: https://storage.spec.whatwg.org/
1. CouchDB 2.0 is still rough around the edges, particularly with its new Fauxton interface (ex: completely broken when proxied into subfolder).
2. CouchDB 2.0 brings the _bulk_get API which has improved sync by an order of magnitude(s).
3. I do have custom logic for logging in overriding the _session API in order to do rate limiting. (I proxy with nginx for IP rate limiting and node.js for failed password attempts rate limiting.) I also have custom logic for provisioning a new CouchDB database per user and setting up permissions.
4. I host on Digital Ocean but I use their new block storage solution so that a growing db does not become unwieldy/expensive.
5. My SaaS subscription system is kind of unique: When your subscription expires you'll simply lose write access to the server (CouchDB), but you can still pull down your data to PouchDB.
A lot of people would like their current data to just be able to sync, but it almost always needs changes in the way data is stored and complementary changes to the application code
Additionally, moving out of Cloudant and into CouchDB with an openresty based reverse proxy has made things even better, and really fun. This is one of those stacks that feels easy and simple at the same time. (Ref:https://www.infoq.com/presentations/Simple-Made-Easy).
Its really openresty + couch that does it for me. The idea of writing security / validations / routing etc right into ngnix combines beautifully with the CouchDB way of thinking.
https://www.ibm.com/blogs/bluemix/2016/09/new-cloudant-lite-...
https://www.ibm.com/blogs/bluemix/2016/09/new-cloudant-lite-...
Stefan Kruger, IBM Cloudant Offering Manager
Main advantage is to have an app that works offline and all sync happening under the hood.
If you are going to give a try, highly recommend watching Nolan Lawson's videos in YouTube.
I was peripherally involved in the part of the services that dealt with user data - basically, massaging data from the address book and call history, dealing with the backend.
This is a collection of things I believe I can say after working on this for almost two years:
- PouchDB + CouchDB works well, as long as you are already used to model your application in terms of CouchApps. No SQL, no E-R representation of your data, etc. If you are not comfortable with the couchDB model and don't know your way around views, you are not going to have a good time.
- This is not an issue with PouchDB, but it if you are planning on having native apps, I'd strongly advise against. All of our native app developers struggled with the change in mindset and the not so mature state of couchbase for Android/iOS. And because they can use SQLite, forcing them to adopt couchbase was not but pain in our team.
- You need to figure out security and authentication: if you take the direct approach of using one couchdb per user and just replicate that, you need to figure out on your own how to secure access. CouchDB only supports OAuth 1 out of the box. We are using oAuth 2 for our services, and I basically had to implement a proxy server that checked oauth tokens before passing requests to our couchdb server.
- CouchDB that does not allow cross-database views, so if you take the "one-database-per-user" approach and you need to check anything that spans more than one database, you are on your own to create more databases/replicate/index/aggregate the data.
- The solution for sync helps, and continuous replication makes for a good demo, but it is not magic. Imagine if you have already accumulated some good amount of data in your database. The moment you start your application on a different browser, PouchDB will start desperately to sync everything you have. You need to know how filtered replication works.
- Very resource intensive, more so if you keep continuous replication. Constant usage of CPU and network I/O can make your app feel sluggish.
As for your app feeling sluggish, my hunch is that you're seeing this in a Chromium or Gecko browser (e.g. Android WebView) in which cases IndexedDB does quite a lot of heavy operations on the UI thread (http://nolanlawson.com/2015/09/29/indexeddb-websql-localstor...), which can be mitigated by moving it to a Web Worker (as Pokedex.org does) or a Service Worker (as HospitalRun.io does).
Your PouchDB application works locally on your device, whether the connection is up and down, and the data is synched with the remote database whenever there is connection.
The alternatives to this PouchDB to CouchDB synching mechanism would have to be either:
- the user checks whether the connection is up, and manually manages the sync, or
- the application programmer saves the user the trouble by adding code that checks whether the connection is up, and automatically manages the sync
However just wanted to clarify that PouchDB doesnt natively use localstorage for storage, its primarily IndexedDB or WebSQL in the browser (leveldb in node).
We do actually use localstorage for cross tab messaging, but thats mostly a hack due to the lack of idb event listeners (that are coming in v2)
localStorage is not indexed, you have to implement some sort of lookup yourself. If you go that way, a tip is to store values as objects with ids as keys, that way you get a sort of hash map as index thing going which alleviates things a bit. PouchDB is indexed and provides a way to query your documents.
The only downside I've found so far is that the PouchDB Inspector on my Chrome browser tends to go rogue from time to time and suck up > 50% of the CPU time and has to be shutdown manually.
Datomic already has a CouchDB-compatible datastore via its support for Couchbase. That would mean Datomic-ClojureScript on PouchDB could run in the browser and sync with Datomic-Clojure backed by CouchDB/Couchbase on the the server -- that would be killer and give Datomic massive reach.
Has the Datomic team considered PouchDB as a possible datastore for Datomic-ClojureScript in the browser?
NB: Just posted the question to the Datomic discussion group: https://groups.google.com/d/msg/datomic/uqBQE4QlnzI/VuZ14pqO...
db.replicate.to('http://example.com/mydb');
IMHO, should be db.replicate(to: 'http://example.com/mydb'); db.replicate
.from('http://example.com/mydb1')
.to('http://example.com/mydb2')
.on('complete', function() {
})
.on('error', function() {
})
.then(function() {
});Unless Pouch has anything required missing I think it's fine though maybe a little unintuitive (like can you do multiple froms? Multiple tos? Etc)
db.replicate
.from('http://example.com/mydb1')
or db.replicate
.from('http://example.com/mydb1')
.from('http://example.com/mydb2')
.to('http://example.com/mydb3')
or db.replicate
.from('http://example.com/mydb1')
.to('http://example.com/mydb2')
.to('http://example.com/mydb3')
? serious questionreplicates from mydb1 into the pouchdb object represented by 'db'
2 & 3:
I'm pretty sure chaining replications doesn't work in that way, although thats a pretty interesting thought.
If you have to chain replications and achieve the structure in 2, you'd have to define the pouch objects as:
db = new PouchDB('localDB');
mydb1 = new PouchDB('http://example.com/mydb1');
etc ..
and then do:
mydb2.replicate.to('mydb3',{live:true});
mydb2.replicate.to('mydb1',{live:true});
mydb1.replicate.to('db',{live:true});
the 'live' flag, as you'd imagine, makes the replication live/continuous, as opposed to one-time.
db.replicate.to_stdout('http://example.com/mydb');You can also explore pouchdb-replication-stream to build bundles that PouchDB can bootstrap from a little bit faster than a chatty replication.
That said, I've found initial replications of large databases (one I've worked with this week is a 25+ MB CouchDB database full of photos) is quick enough (and mostly bandwidth constrained) that I haven't had much in the way of concern over it.