Scalable ACID
dbmsmusings.blogspot.com
dbmsmusings.blogspot.com
As for scalability, as this post shows, I think the best thing to come out of NoSQL is the race for greater scalability it has and will continue to inspire in SQL solutions.
1. The rejection of what some see as an overly complicated and inflexible query language in favor of map-reduce phases or other application-side querying/processing.
2. The rejection of serialized transactions in favor of something like eventual consistency, often with application-specified conflict resolution.
3. Or it could be a rejection of the relational model (especially as it is popularly implemented) in favor of key-value, graph, document, column-oriented, etc. models.
If I read you and the authors right, I think you're thinking more along the lines of 3 and the authors are focussing more on 2.
I think both "sides" are just starting to come to grips with the idea that these design choices can be orthogonal (a relational DB without SQL? an eventually consistent SQL DB? a graph DB with 2-phase commit?) and choosing to reject one part of "the old way" doesn't mean you have to reject all of it. The upshot is, we're going to have a lot more tools to choose from for our own unique problems.
So, I'd say: don't worry about feeling like a curmudgeon! Rigidity will have its place in the beautiful gleaming pluralistic future of datastores that the SQL vs. NoSQL "debate" is building the road towards.
I can't really figure out your foo example though, needs to be a bit more specific.
Most of the popular NoSQL solutions were designed to run as large server clusters, possibly distributed across multiple data centers. They're designed to handle such problems as, "Our fiber backbone is down, and our database has split into a 5,000 machine cluster and a 7,000 machine cluster! And it's three days before Christmas!" (This is more or less the problem Amazon's database is designed to survive.)
For any given "NoSQL" database, you want to figure out how they handle this problem: Do they allow the 5,000 and 7,000 node datacenters to become temporarily inconsistent? Do they block writes to the entire database? Do they block writes to the 5,000 node database (because it doesn't have a quorum)?
Any of these options are valid—you just need to know what tradeoffs you're choosing.
"Over the past few weeks, in my advanced database system implementation class I teach at Yale, I’ve been covering the CAP theorem, its implications, and various scalable NoSQL systems that would appear to be influenced in their design by the constraints of CAP. Over the course of my coverage of this topic, I am convinced that CAP falls far short of giving a complete picture of the engineering tradeoffs behind building scalable, distributed systems."
http://dbmsmusings.blogspot.com/2010/04/problems-with-cap-an...
That is indeed counter-intuitive at first, but they make a good case. I'm impressed with the idea, and hope to follow this more closely. Anyone know any other research on the subject?
"Hence, in an era where disk reads caused wild variation in transaction length, allowing nondeterministic reordering was the only viable option."
In other words, when all your data is in ram (possibly distributed across a cluster of machines connected by a relatively fast network), the performance tradeoffs are very different than when your data is on rotating mechanical disks.
I don't know that I'm convinced that their approach is better, but they've certainly given me something to think about (and re-read in more detail later).
I am not sure I am convinced by their results though. They say their deterministic system seems viable when comparing its performance to traditional systems under short in-memory transactions. That is a special case that is clearly in their favor though. In that situation, the amount of time spent in processing data is greatly reduced so the network overhead becomes much more significant - so the system that does less network communications will obviously win...
I guess it may potentially be good for in-memory database systems for stuff like OLTP apps (e.g. VoltDB and TimesTen), but then I think most OLTP apps are okay with a more relaxed isolation...
Different tools for different cases. Polyglot persistence is the way forward.
They had me up until they called NoSQL "lazy" and claimed NoSQL solutions "give up on ACID". They aren't lazy and a lot of them support ACID, or most of it.
From that point on this veers away from logic and moves towards a self-justified argument. Tools are just tools. If you need immediate consistency then go for an immediately consistent solution, if you don't then you should feel free to take advantage of the performance gain of eventual consistency.
It's hard, but it's doable. I find it quite striking that no one ever seems to mention the two leading high performance parallel databases, HP Neoview and Teradata. Is it that people don't realize they exist?
Any DML that modifies a row in Teradata either locks the whole table if it doesn't know immediately where the row is, or locks the row hash if it does (if you give it the primary index value). Locking the whole table involves sending a command to all of the nodes and waiting for a response. That's why most writes are typically done as large batch loads at infrequent times.
Not to be ornery about this or anything.