The Abandoned Facebook Tech That Now Helps Power Apple
wired.com
wired.com
As for not mentioning Apache.... That is bad. I'm not sure what they could put in that article that would speak to their audience though.
Moreover, FB needed a lot of people who knew the tools. They acqui-hired a team of HBase experts. In 2010 I would have chosen HBase over Cassandra for what they were doing. In 2014, I'd choose C*. But as always, it depends on your application and what you know (or can hire for).
Cassandra has tunable consistency. A lot of people just think "eventually consistent" but it really lets you make tradeoffs yourself about performance, consistency, etc. Plus, the way they do read repair, etc even at CL==ONE makes it a lot better for messaging systems than you'd guess. I'm using it successfully with most stuff at One.
Mostly the limitations people come up with with Cassandra are around cross-row ACID, which if you really need it you can do with LWT, though in practice the better solution tends to be to just denormalize in to one row and/or NOT actually require cross-row ACID. It generally isn't what you really need anyway. Most businesses don't really shut down just because they can't achieve consistency.
In practice, Cassandra gets around this issue by implementing a custom CRDT for counters. But CRDTs don't exist for everything and in general it's not trivial to get things right. I'm not arguing Cassandra is bad, but it's a misconception that you can tune the consistency levels and somehow get Cassandra to behave like a strongly consistent store.
I want to know: Can someone tell me how/where to focus my efforts to learn more? In my day job, I use mostly MongoDB these days, but also some MySQL. I have a few resources that I'm aware of:
- http://aphyr.com/tags/Jepsen - Have browsed this and found it interesting.
- MySQL reference manual http://dev.mysql.com/doc/#manual
- MongoDB Manual docs.mongodb.org/manual/
- High Performance MySQL - I purchased this book but never got past the first chapter. Not to say that I can't...
- use-the-index-luke.com/ - Rarely read this, but have given it a half-hearted start from time to time.
- http://www.mysqlperformanceblog.com/ - ditto
I don't know when to use a relational database and when to use NoSQL. I don't know the differences between the various NoSQL databases. I don't know the ins-and-outs of any of the systems and I don't know how to determine what path to go for high availability, geo-distribution, replication, etc.
I find it all very overwhelming and would like to learn more. I'd like to be able to pull plugs in things and see my systems stay up (which, I'm sure would require intimate knowledge of at least one system). Does anybody have any suggestions? How valuable do you think this stuff is?
I originally skipped over the first chapter on Postgres but went back and read it and learned a few things (too bad there's no mention of the recently added json and jsonb data types).
It sounds like you're conflating three (or more) problems -- knowing if your data is relational in its nature, understanding how to efficiently retrieve the data you're looking for and understanding system administration/architecture tasks for a particular system.
The reason why I bring this up: my first job out of college was as a (very) junior Oracle DBA. At that particular job, that meant knowing some stuff about data retrieval but knowing much more about performance tuning and the architectural idiosyncrasies of Oracle and lots of Linux stuff.
Later I moved into startupland and necessarily learned quite a bit more about the data storage and retrieval patterns popular in web applications.
My point is: you might want to pick a particular aspect of databases that you want to learn about and really focus on that -- just like you would for a new programming language or framework. The way I usually do it is by finding a project that seems just a little too hard to be easily in reach and learning everything I can to make it happen.
Anyway, good luck and happy learning :-).
a good method of learning is to use some really large dataset and come up with some questions you want answered about the set, then write sql to get that information out. if your queries take a long time, begin indexing and looking at how the queries can be rewritten to return results faster.
i cannot recommend a good, large interconnected/relational data source, but something like zipcode databases [1] can get you started. you can also play with geo-spatial indexing to find zipcodes within x miles of each other. [2]
[1] http://download.geonames.org/export/zip/
[2] http://www.mysqlperformanceblog.com/2013/10/21/using-the-new...
Perhaps this is now out of date?
We switched to Cassandra because of the control it offers over how your data is laid out on disk, and it's ability to not totally fall apart under load (in fact performing really well). It's also a lot easier to deal from an ops perspective.
I did a post about the migration here: http://blakeeggleston.com/migrating-databases-with-zero-down...
http://www.fullcontact.com/blog/mongo-to-cassandra-migration...
http://relistan.com/cassandra-vs-mongo/
The main selling points for Cassandra over Mongo are that it scales to more machines easier for handling larger workloads, and the write performance is better.
MongoDB is a document store and so if you data is structured that way it is unrivalled. But if isn't (more than likely) than it will suffer greatly as you start to increase the number of joins.
Just anyone who has run in to scaling problems with MongoDB. ;-)