This seems like an interesting project that tackles some of the data versioning stuff. However, I believe that, at least in data science, we need data versioning closely tied to the analyses themselves for complete reproducibility.
That is, we need the versioning tied to the inputs/outputs of data pipeline stages, such that we can reproduce pipeline runs at any time and incrementally improve and run pipelines based on diffs in data.
As mentioned elsewhere in the comments, Pachyderm (http://pachyderm.io/) does exactly this. Working both as git for data, but also enabling data pipelining and analyses with the data versioning.
One way to think of Noms is that is an index optimized for computing diffs.
Here's a screencast that shows off Noms diffing things fast: https://www.youtube.com/watch?v=Zeg9CY3BMes
* Convert the versioned data to tab-separated values
* COPY it into Postgres every time
* Hope Postgres can act immutable enough even though it wasn't designed to be
The closest I've come to improving this situation was Kyoto Cabinet (unusable license) and rolling my own damn hashtable (it worked okay but adding new kinds of indexes was just unmaintainable, there's a reason databases should be made by experts).
Pachyderm is git for data. We work hard to make sure we can store data of different types (binary, text, json) efficiently. We also work hard to give you good mechanisms to read the data in a distributed way. I'd be curious how this suits your purposes.
While I can see how a git-based filesystem can help with some use cases, does it do any kind of indexing at all? I see that the FAQ recommends exporting the data from Pachyderm into PostgreSQL, which leaves me where I am now.
The nice thing about git is that it doesn't really require much of the server. I run my own "git hosting" with linode, apt-get, and ssh.
Most of the solutions for immutable big data don't have that level of convenience, as far as I know. Anything which requires a lot of sysadmin and ops work will eventually become a commercial cloud service... git on the other hand is useful by itself, and good companies can obviously be built on top of it as well.
The difference between this and a version control system is a major part of why the UI is so awful.