CouchDB Book Now Available from O'Reilly
oreilly.com
oreilly.com
In the section "Self-Contained Data" the author makes the argument that some data is better stored as a document than a bunch of relations, and uses INVOICES as an example: "Accountants appreciate the simplicity of having everything in one place. And given the choice, programmers appreciate that, too.". Invoices (and other business data) are the classical example of something that belongs in an RDBMS and not in something like CouchDB. It's not that hard to come up with a good examples for "relaxed data", like... say... blog posts, or blog comments, or twitter messages, or search results, or log entries?
Later, it gives incorrect/meaningless definitions for Consisteny, Availability and Partition Tolerance. It then completely confuses the CAP theorem, incorrectly placing Paxos (puts in under CA, should be CP). It also places RDMBSs on the figure, which is silly, "Relational Database Management System" in itself says nothing about how it is distributed, if even.
Part I/1+2 needs major rewriting.
This book should be a blog. It's not worth money.
Blog comments, on the other hand, might benefit from more normalization. Cramming a bunch of comments into a post document could result in contention problems, depending on the engine and how you're handling them. In CouchDB, you'd end up having to update the entire post document for each new/changed comment. Mongo handles it better, since you can atomically push new comments into your doc. Still, I think I'd rather have them in their own documents on a busy site.
In an Invoices you have an Invoice Header, that should refer to multiple entities, buyer, seller, shipment address etc ... . Invoice detail that refer to the Header and represent a one to many relation between the data in the invoice header and the items in the invoice. Which may also have additional relationship attributes, like price, discount or any details specific to the item, for example, maybe your DB support a shipment address per item!
Anyway, in a relation DB an invoice refer to many entities (separate facts) and many relations (also separate facts) each worthy of its own table that refer to each other.
If you believe that Invoices are a good candidate for a document DB, you probably don't believe that the Relational Model is valid in general. Or that the 2 concepts I mentioned at first really help integrity!
The main flaw I see is that the relational model makes is hard to create dynamic models. A good Relational Model practice is that an entity in a Relational DB should represent an entity from your Universe Of Discourse (uod). That is the say, a model is better when tables represent real entities of the problem you are modelling! This is sometimes impossible when you want to store dynamic structures, some argue that this is not the fault of the Relational Model theory, but rather its implementations.
Also, in the case of a simple blog app, I don't think that contention is going to be your biggest worry. The on-disk structures for CouchDB are append-only, so your biggest worry isn't going to be locking, it's going to be the stale document revisions taking up disk space between vacuum operations, and the replication overhead for all the intermediate versions.
I was read the first few chapters as it was being written and thought it was great.
Anybody know the backstory of the title change? Did O'Reilly push for it?