Twitter Drops MySQL For Cassandra
informationweek.com
informationweek.com
This and a few other poor explanations make me wince reading this. So I don't particularly recommend this article, although I've seen worse. :)
TFA does at least link the original interview w/ Ryan King of Twitter (http://nosql.mypopescu.com/post/407159447/cassandra-twitter-...) which is much better for people at HN level.
My own article at http://www.rackspacecloud.com/blog/2010/02/25/should-you-swi..., although high level, also has some useful links for those who want to drill down for more details.
As an aside, three sentences doesn't really warrant a "tl;dr".
http://code.google.com/p/scalaris/wiki/FAQ#Is_the_store_pers...?
http://highscalability.com/are-cloud-based-memory-architectu...
http://www.allthingsdistributed.com/2007/10/amazons_dynamo.h...
My understanding is that it came down to Cassandra being the only one that was well tested in multi-datacenter configurations. I think there were other reasons, but that was the big one.
As the expression goes: at the poker table, if you don't know who the sucker is, it's you. Guess who's responsible for those fuddy-duddy, old-fashioned things like querying your data, or making sure that you've written data, or that your database isn't corrupted... it's you.
If you're Twitter, and you're fail-whaling every day, then maybe the work required to make this trade-off work makes sense. But I can't help but feel there's got to be a better way.
A distributed hash table can be built with transactions, isolation, and so forth. Such a system would offer a different set of trade-offs that might satisfy a different set of users.
In practice though, I think if you introduce (multi-'row') ACID into any of today's NoSQL database, you'd just end up with a bad traditional database (a 'relational' database without a strong theoretical grounding, without the ability to do joins and without a powerful query language.) This whole NoSQL movement feels like a re-run of the evolution of the relational database - those that don't know their history are doomed to repeat it.
Strongly disagree, although there is truth to the converse (adding sharding to a traditional database yields a bad nosql implementation :).
There's a sketch of how Google added multirow transactions to bigtable for app engine (via a layer above bigtable called megastore): http://perspectives.mvdirona.com/2008/07/10/GoogleMegastore....
The key idea is that "a per-entity-group transaction log is used. One of the rows that stores the entity group is the entity group’s root. The log is stored with the root, which is replicated like all rows in Big Table."
This is basically a version of the approach advocated by Pat Helland in his paper on "Life Beyond Distributed Transactions" -- http://www.cidrdb.org/cidr2007/papers/cidr07p15.pdf -- namely, that the most sane approach to distributed transactions is to redefine the problem, and restrict transactions to be within "entities" that fit on a single machine. Which, as App Engine demonstrates, turns out to be enough to do an awful lot.
Would you mind explaining why you disagree?
> the most sane approach to distributed transactions is to redefine the problem, and restrict transactions to be within "entities" that fit on a single machine.
How is disallowing distributed transactions a sane approach? There's no harm in allowing distributed transactions using 2PC and letting the developer decide if he wants to incur the performance penalty or wants to colocate the data-structures to the same node to avoid 2PC overhead.
> Would you mind explaining why you disagree?
I did, that's what the rest of the comment is. :)
> How is disallowing distributed transactions a sane approach?
That's the subject of Pat's entire paper, which I linked.
Or as Orwell put it: "A not unblack dog was chasing a not unsmall rabbit across a not ungreen field."
And when combined with a thinly veiled insult, gives the wonderful "back-handed compliment"
"He's not as unintelligent as he looks" "She's not the liar she might have been given her poor upbringing"
my not-unpreferred figure of speech, please. If you want to make it not unpossible to believe you.
(Thanks, really; I've developed a belated soft spot for rhetorical devices.)
I had a conversation with an engineer who works at a pretty well-known company here in SV and their sys admins are dropping Cassandra and pushing all the engineers back to MySql. I don't know the whole story but it seemed to be implied that open-sourced Cassandra had issues and supposedly Facebook had a much different version they were using.
Of course this is all second hand, so I tried to search on the experiences of other people using Cassandra (with decent volume). Unfortunately most of the threads I found had people just like me, at the exploratory stage. Or they hadn't been live with it for long.
If there were any pitfalls or hairy parts with maintain Cassandra, that would be good to know. Also, examples of clients who have decent load and have been using it for a while.
I'm curious what you are thinking of, because I have better picture of companies using Cassandra than most. :)
I do know of one company that fits your description, where some of the mysql DBAs were very anti-cassandra because, frankly, it's not mysql and that's what they were used to. But that has been resolved (the most vocal DBA left) and the Cassandra migration is continuing.
> If there were any pitfalls or hairy parts with maintain Cassandra, that would be good to know
Other than the obvious (it's not a relational database), we've documented the main limitations here: http://wiki.apache.org/cassandra/CassandraLimitations
http://ria101.wordpress.com/2010/02/24/hbase-vs-cassandra-wh...
~10k ops/second per quad-core node. Scales roughly linearly w/ node and core counts. ("Roughly" means, obviously there is network overhead as you move from single node to multiple, that kind of thing.)
> How big can the values be?
2 GB, although you probably don't want to max that out.
Cassandra is mostly used for smaller pieces of data, although I do know at least one person using it as an S3 replacement by chunking files into 64MB chunks; each file is one row consisting of columns that each contain one such chunk.
> I was thinking of using it as a backend for a mail system.
It should work fine for that.
Another, that goes into even more detail (imo, too much for one article, but a good article all the same): http://www.vineetgupta.com/2010/01/nosql-databases-part-1-la...
Benchmarking systems w/ very different data models is difficult to impossible, which is why you don't see that in this kind of survey piece. You're best off by picking one segment and focusing on that. Yahoo did that with the ColumnFamily stores (cassandra, hbase, and one they wrote internally) here: http://www.brianfrankcooper.net/pubs/ycsb-v4.pdf (note that Cassandra 0.5 results are on page 16 and 17, not inline w/ the rest)
Hrm, maybe it's a book more than an article, although I don't think I'd buy it, given that it'd be out of date so soon.
The point being, as a casual observer of these things, I don't yet have a good feel at all for which ones might be good for what.