4 Months with Cassandra at Cloudkick (YC W09)
cloudkick.com
cloudkick.com
Can anyone explain/posit an idea as to why MongoDB and CouchDB have been stealing the thunder for the past year rather than Cassandra? It just seems odd.
Administration is pretty hands off for most tasks. For instance you want three copies of each piece of data? Set the replication factor to 3. Other parts are difficult to "get" if you don't have an understanding of the code, my best advice is to make sure you approach Cassandra from a developer perspective. You can't treat it like a black box quite yet.
Since Cassandra uses Apache Thrift as the default RPC mechanism, exposing the Thrift layer to any non-controlled data can be dangerous. We use firewalls on our nodes to make sure our Thrift ports are only exposed to a very small set of machines, because even just telneting into the port and typing "hello" can cause the JVM to OOM.
--
I use Redis and heavily guard its telnetable port, but it doesn't OOM. This issue should have been fixed before public release, imo. You wouldn't want something as simple and common as a port scan to shutdown your data layer.
You can see the opinion of the Cassandra developers, which is pretty negative towards thrift, in the linked ticket: https://issues.apache.org/jira/browse/THRIFT-601
This is one of the many reasons the Cassandra developers are looking at replacing Thrift with Avro: http://hadoop.apache.org/avro/
It doesn't look like that from the API wiki, but maybe someone knows if that's possible, or planned.
('to' is optional, and technically, so is 'from'.)