First off, most use cases don't actually fall within the "big data" family of use cases. Three reasons for that (a) You will often work with datasets that are so small, that the sheer amount of data and efficiency of processing won't present a difficulty anway. (b) You will often be in a situation where it is workable to do random sampling of your real dataset very early on in the data processing pipeline which will allow you to still obtain valid estimates of the statistics you're interested in while reducing the size of the datasets that need to be juggled. (c) You will often be able to do preaggregation. (For example: instead of each observation being one record, a record might represent a combination of properties, plus a count of how many observations have that combination of properties).
My strategy will be roughly as follows: A "database object" that's tabular in nature is, by default, a CSV file. An object that's a collection of structured documents is, by default, a YAML file. The data analysis is split up into processing steps, each turning one or more input files into one output file. Each processing step is a Python-script or Perl-script or whatever. You can get pretty far with just Python, but, say there's one processing step where you need to make a computation where there's a great library to do it in Java, but not in python; feel free to drop a java programme into the data analysis pipeline that otherwise consists of Python. Then you tie the data processing pipeline together with a Makefile.
This general design pattern has several things to recommend it:
(1) Everything is files, which is great! If I work in a big dysfunctional corporate environment, I might be faced with this scenario: I have a bunch of databases at my disposal, like an enterprise-scale Oracle server. But waiting for signoff from managers and waiting for database admins to provision a table space, create a schema, make some grants, would take so long, that in the same amount of time, I can be half-finished implementing the whole thing with a file-based system, since I need no one's permission and no one's cooperation to just create a file on some machine. A bit further into the project, I might face a tough deadline for producing a report that depends on a heavy payload of data crunching. Going with a shared Oracle, I might find myself in a tough spot where the database server is completely hammered by what other people are doing with it, and there will be zero I can do about it to get my payload finished in time. With a file-based solution I can usually work in a compute environment where resources are less contended company-wide than on a central database server and I can easily move from one server to another if I should need to for capacity-reasons. It sounds like I'm incapable of system-level thinking and being a teamplayer. But it's just the way real life is. DBAs are unsung heroes who don't get the resources they need. Shared resources fall victim to the tagedy of the commons. But at the same time, going into a meeting saying "I don't have the numbers today, because Oracle was slow" sounds like "the dog ate my homework" and will reflect poorly on me personally, rather than the organization, so I try not to put myself in that situation.
(2) I like to equip my scripts with the ability do progress report and extrapolate an ETA for the computation to finish. If, 1% into the computation, it becomes apparent that it takes too long, I cancel the job and think about ways to optimize. I'm not saying it's impossible to do that with a SQL database, but your SQL-fu needs to be pretty damned good to make sense of query plans and track the progress of their execution etc etc. If you have a csv file, it might be as simple as saying "if rownum % 10000 == 0: print( rownum/total_rows )" then control-c if necessary. In practice, doing things with a database often means that you send off a SQL query with no idea of how long it's going to take and if it's still running after a few hours you start to investigate. But that's a few hours of lost productivity. -- Things are particularly painful when the scenarios described under (1) and (2) combine. You might be used to a certain query taking, let's say, 4 hours. Today, for some reason, it's been running for 8 hours and is still not finished. You start suspecting that the database is busy with other people's payloads and give it a few more hours, but it's still not finished. Only now do you start investigating what's happening. But this sort of lost productivity is often the difference between making a deadline on reporting some numbers or something or missing it. (Think about the scenario where it's a "daily batch", and you need to go into that all-important meeting, reporting on TODAY's numbers, not yesterday's.)
(3) "Make" is a great tool for doing datascience, but in order for it to be able to work its magic you mostly have to stick to the "one data object equals one file" equation. You want parallel processing? No problem. You've already told Make about the dependencies in your data processing. So wherever there isn't dependency, there's an opportunity for parallelization. Just go "make -j8" and make will do its best to keep 8 cpus busy at all times. You want robust error handling? No problem. Just make sure your scripts have exit code zero on success and nonzero on failure. "make -k" will "keep going". So when there's an error it will still work off the parts of the dependency-graph where there wasn't an error. Say you run something over the weekend. You can come back monday, inspect the errors, fix them, hit "make" again, and it will continue and not need to redo the parts of the work that were error-free. Etc. Etc. Etc.
Now, after this whole prelude around my philosophy of doing data crunching pipelines in a datascience context, we finally get to the point about KyotoCabinet.
Even though you usually find that CSV or YAML is fine for MOST of the data objects in your pipeline, there will almost always be SOME where you can't be so laissez-faire about the computational side of things. Say you have one CSV file which you've already managed down to a manageable size (1M rows, let's say) through random sampling. But it contains an ID that you need to use as a join-criterion. Let's say the table that the ID resolves to is biggish (100M rows, let's say). You can't really apply any of the above "tricks" to manage down the size of the second file. By random-sampling you'd end up throwing away most of the rows, and your join will, for most rows, not produce a hit, even though there would have been one to begin with which you've just decided to throw away, which would be pretty bad. So, for that file, you can't get around having it sitting around in its entirety to be able to do the join, and CSV is not the way to go. You can't have each of the 1M rows on the left-hand side of the join trigger a linear search through 100M rows in the CSV.
Your two options for the right-hand side would be to load it into memory and join against the in-memory data-structure. Or use something like KyotoCabinet. The latter is preferable for a number of reasons.
(a) Scalability. A datascience project usually has a tendency for the size of data objects to get bigger over time through additional feature requests being added to the project. If you get to a point where a computation that you've initially implemented in-memory exceeds the size where this is no longer feasible you're in trouble. If you go to a pointy-haired-boss and tell them "I can't acommodate this additional feature request without first doing some refactoring. So I'll do the refactoring this week, then start working on your feature request next week", it sounds in their ears like "I'm going to do NOTHING this week". So, I have a strong bias against doing things in-memory so as to not put me in that situation.
(b) By making the datastructure persistent, it means that the computational effort that goes into producing the data structure doesn't have to be expended over and over again, as you go through development cycles, fixing errors etc on the payload-side of the computation.
It may not even be slower in terms of performance, thanks to the MMAPed IO that KyotoCabinet and other key-values stores do.
...so this is roughly where I'm coming from as a KyotoCabinet frequent-flyer.