NoSQL: If only it was that easy
bjclark.me
bjclark.me
Relational databases, such as MySQL and PostgreSQL, are great for some things -- especially when you need to (put simply) divide and recombine various pieces of information easily. JOINs. UNIONs, etc. are fantastic when you need them, and it's where relational databases excel.
On the other side of the equation are hash and "document" databases, which are mostly key-value stores with some unique functionality (e.g. the indexing feature of Tokyo Cabinet Tables). These are great when you need to store a large volume of data, but don't need to frequently recombine the stored data in many different ways. You can run simple queries with some hash DBs and retrieve specific data, by key, at the drop of a hat; this is where (for the most part) key-value stores win out over relational databases.
It is clear, at least to me, that these two classes of DB serve two different purposes, with some overlap. By that logic, I see no reason why they can't happily co-exist in a single project. In fact, I'm using multiple database formats in a single application without any problem. Sure, there is a little more management and some extra logistics involved, but the point is that each type of database is used with its strengths in mind.
I'm not going to force MySQL to be a giant hash table when something else will do the job better, and I'm not going to force Tokyo Tyrant to try and be a relational database (not that I really could).
Am I wrong in finding a happy medium between multiple technologies?
No, you are correct. The main source of contention lies in people trying to use a particular technology outside of its comfort zone. Many in the NoSQL crowd suggest that this is a common symptom of SQL installations. Unfortunately, a good portion of the arguments that address this point are hyperbolic and seem to suggest that SQL is rarely the right tool. This is what keeps degrading the signal to noise ratio.
As in, when you design a relational DB by the book, you first create a normalized schema with indexes, and subsequently issue SELECTs with WHEREs and JOINs and GROUP BYs and ORDER BYs and LIMITs against them in a declarative manner.
With a KV store you basically have direct access to the btree layer, so you have fine control over what's going on. Just as with Assembly, it is much harder to organize your data/program; most people don't need this kind of low-level access, and will shoot themselves in the leg given the possibility. Currently it only makes sense to use a KV store if you need this kind of low level access for performance/scalability reasons (or you want really strong replication which only exists in Keyspace, which is a KV store).
Hardcore SQL admins will say that one can optimize the RDBMS , and they're right; this is the usual trade-off when your high-level abstraction's default behaviour is too slow and you start to break open the abstraction to tune the underlying layer: in my opinion, it makes sense to just abondon the higher-level abstraction at this point --- but this is a matter of taste since you're loosing many other conveniencies along the way.
The other use-case is sharding. With a KV store it's much more natural to store parts of your tree on different servers and just issue GETs and SETs over the network. This transparency is one of the things you gain if you go down a level in the layers of abstraction.
['] http://scalien.com/keyspace
[''] Prolog really
In NoSQL, I don't have to tell the database layer what I'm stuffing into the value (or even how I'm building the keys). Of course the application is taking on a lot of responsibility that previously belonged in the persistence layer.
I think this is the true draw to the NoSQL philosophy. Especially when you consider the apparent (to me) rise in meta-programming.
Static typing is a good analogy, but the main purpose of static typing is expressing the semantics of data structures in a way that the system itself can reason with them. (OCaml or Haskell's type systems are much better examples of static typing's utility than Java's or C's, FWIW.) If you're still prototyping and/or the data structures are still in flux, then something completely dynamic will avoid a lot of extra work, but once things settle down, having the database itself be smarter about processing the data is worth consideration. (Of course, if your entire data set fits in a Python dictionary, then using a real database is overkill anyway.)
The comparison to dynamic vs static typing doesn't hold much water.
Can you explain how meta-programming relates to this topic at all?
I'm not sure it's a valid comparison, though: OCaml has a statically typed Lisp-style macro engine (camlp4), for example. (In all honesty, though, I've never used it. Lazy evaluation, the packaging system, and other language features cover many of the same use cases.) I think it's a case of assuming the C family's type system is the cutting edge of static typing, when it's actually pretty archaic.
The point about K/V databases having a rigid schema of (key_type -> untyped_value) is a good one, by the way. Of course, association tables (AKA "dictionaries") are pretty versatile as a basic collection type - looking at Lua, Python, Awk, or Javascript. They're not ideal for all cases, but they're a good start for most.
meta-programming relationship: Since you can store anything in the value part of a K/V store. You could (and I do) store the schema definitions for the other values in the data store. And while you are at it, you could store code segments.
So imagine the following key store (in pseudo code):
{ key: 1728273, value: { fields: { first_name: string, last_name: string } } }
{ key: 8274289, value: { type: 1728273, data: { first_name: "Bruce", last_name: "Saunders" }}}
You could stuff code segments in there in a similar way. So I think K/V stores are potentially highly related to meta-programming. Though certainly the two could exist entirely separately of each other.
Grammar nitpickery aside, good article. Is client-side hashing really that big a deal? I don't know from first hand experience because I've never bothered with these things yet.
Is client-side hashing really that big a deal?
No, it's not that big of a deal, but I tend to feel that any layer of logic that can be done on the DB end rather than the implementation end is better off in the former. Voldemort is a good example, as noted in the article.
A few months down the line your load gets so high that you need to add a third server, but your sharding mechanism is based on mod 2 - so you need to hack around this in one of many ways to utilize the third server effectively.
Make sense?
Sometimes, but oftentimes sharding issues are unique enough that they require custom solutions when the backend technology does not support it natively.
Not a long time ago, if you had to handle 1 billion records then you were already big with piles of $ sitting beside you. Then you would just call Oracle or Greenplum (or whatever) and say "Hey, I've got some million $ here, care for it? I just want to be fast(tm)"
Lately it is quite possible to be small and forced to handle millions of records daily or struggle with new technologies that require less man power for development but need scaling a lot earlier (e.g Rails). You have limited resources and if you can get away with Tokyo Cabinet on a single machine instead of a 100k mysql cluster then you get to live longer.
As any new cool technology it'll get some hype and then fall back to being the right tool for the job. It's good to have some alternatives to RDBMS, some problems are really not relational.
I wasn't there when it happened in the 70's, but I think it's essential to understand why relational beat hierarchical then; and to understand what has changed since; because only then can we make an informed prediction.
I think the issues were efficiency and ability to analyze the data. Anybody know what the actual reasons were?
They have never arrived in the mainstream, but they exist as a niche luxury item, mainly because the mainstream use broken languages with broke runtimes that can't persist objects between sessions without getting their pointers in a twist. So they settled for the next best thing: Object Relational Mapping, along with a matching pair of Impedance Mismatch. It's ORMs that are coming and going; at least in the industry hype-machine.
Though you are right that the reason RDBMS became the norm has a lot to do with working with the data (analyzing to use your word).
A key/value store is effectively just an associative array, while damn useful for certain things, I am hesitant to call that a DBMS.
Also: Hierarchial databases are very different from key/value stores. K/Vs are flat (joining, batch consistency checks, etc. are usually done clientside in a faster/safer language than the DB was implemented in, such as PHP), while hierarchial databases are (naturally) hierarchial. If a client searches for data in a directory, it will only see the nodes in that directory, or those found by walking any subdirectories. While this can be very fast (most records won't be in the same directory, so the first operation has already reduced the search space considerably), much data fits poorly in trees, and constantly rearranging subdirectories or adding symlinks is at best a workaround for the massive conceptual mismatch.
OTOH, it's a good fit for some cases - we're communicating on a hierarchial database right now. As moving one comment thread from one post or parent comment to another would be incredibly rare (if it were possible), a lot of DB functionality to allow flexibility in data modeling can be discarded. Modeling the discussion as a tree of nodes (S-exps, in this case) with pointers to parent comments, subthreads, user IDs, etc. is pretty straightforward.
I've been going over its documentation, and it looks like the cat's meow in getting rid of the database scalability bottleneck. Think of it as memcached done right. Coherence has automatic clustering, needs almost no configuration, supports many flexible distributed schemes, easy API. With Clojure and Scala, you don't even have to use Java to use it.
Has no one heard of it? Does it get no love because it's commercial and (probably) expensive?
Yes.
The "versus" mentality of this stupid meme comes from those either proselytizing or stubborn. Plain and simple.
If you're too ignorant to understand what is the appropriate tool for your situation, or offended by the perceived alternatives, then you should withdraw from this topic.
* Among other things, RDBMSs put a lot of resources into ACID transactions -- in the use cases for which they were designed, data is expected to outlive some employees, and just dropping part of an update to e.g. a medical record or insurance claim would be a disaster. Similarly, the databases can do elaborate checks on the data, and any potential changes, to ensure that it's always internally consistent. Losing a youtube comment wouldn't be a big deal, though, and if you're willing to cut corners on transaction guarantees or internal consistency, there are performance benefits. (Being extremely protective of the data is a good default, IMHO.) And, yes, not just relational DMBSs are ACID, but many "NoSQL" dbs have that trade-off in mind.
That's not fair. Everyone starts out as ignorant. I think the "versus" mentality helps less informed people separate two fundamentally different approaches to persistence.
There's a lot of information in this post. Granted, it's just one person's opinion. Take it for what it's worth and move on.
My thoughts at http://hypecycles.wordpress.com/2009/08/05/look-ma-nosql/
SQL has survived because it is reasonably portable, is descriptive and has a rich ecosystem. NoSQL is a response to a problem but it is one that will not get mainstream adoption because it is non-standard.
It is a great stepping stone; what comes out of it is hopefully a standards based scalable data storage and retrieval system.
Can someone recommend any solution? I'd be perfectly happy with a "replicating memcached"...
I just spent some time evaluating several k/v stores for a new project. My conclusion: These things may be useful when I'm ready to optimize. Until then, I'll put everything in postgres.
projects im working on now sometimes have the same data represented differently in different systems to take advantage of what those systems have to offer. many of them mentioned in the article.