Dat – A git-like tool for large datasets
dat-data.com
dat-data.com
Now, it is possible that this happened by chance: for instance, perhaps the author has a friend who is a graphic designer, and this friend got really excited about the project and put together all of this as a favor. But perhaps it is a deliberate strategy, in which case I wonder: what are the benefits of front-loading the PR work like this? Did it help in securing grants, for instance?
> a slick web page - this is very often (but not always) a sign of a library that put all of its time into slick marketing but has overly-broad scope and is very likely to become abandoned
http://www.reddit.com/r/javascript/comments/2378xo/the_best_...
I would find helpful if author would outline how PR is done.
Why spend two months building a product if you don't know it will be well received? To me, this is startup development done right: reducing the unknown, reducing the risk in the most efficient way. Even if it means showing a mockup of your service.
iRODS is a data storage system that provides replication (e.g. on local filesystems, S3, and HPSS. Deployments at different organizations can be linked, to make (part of) the namespace of one company be visible to that of another company. It has a permission system on namespaces.
But one of the nicest features is that it has a built-in C-like trigger language. You can specify triggers for certain namespaces. E.g., for research, it is often necessary to pointers to data (URLs) that are permanent and can be published in e.g. papers. For this reason EUDat[1] has triggers on namespaces that automatically create PIDs[2].
But since the trigger language is very complete, you can do nearly anything with it.
Note: I am not involved in iRODS anyway, I just know a couple of colleagues who use it for safe replication, sharing, etc.
I could be missing something here so anybody more knowledgeable should chime in.
The second sentence on their website appears to state otherwise: "As a team we have a bias towards supporting scientific + research data use cases."
Edit: Well, on further reading they do go on to say, "What we're building isn't quite ready for prime time yet, but if you want to play around...", so perhaps I should not jump the gun :)
1. one that commits treeish diffs into objects, and then stuffs them into arbitrary content-addressable storage;
2. and one that manipulates (local or remote) mutable tables of refs to content-hash-URNs -- or "git repositories" for short.
The key component would be a facility (either a library or an OS service) for resolving content-hash-URNs into file-descriptors in some content-addressable storage system. The critical point is that this would not be a git component -- the committing tool wouldn't be in charge of where its objects go, and the repository tool wouldn't be in charge of resolving their URNs. Instead, both would just be relying on the host to decide how to get objects from/put objects into storage, and it's the host that would have a distribution strategy set up for all of its content-addressable objects, not just the ones that happen to be attached to git.
The remote server can obviously be running an object-storage node itself, and so can your own local git repo -- the point is that you can have objects in more places than just your machine and the repo it's pulling from.
See http://www.fancybeans.com/blog/2012/08/24/how-to-use-s3-as-a...
(and if you like that article, consider subscribing to LWN!)
Currently, the site is just a bare-bones version that lets you submit the URL of a publicly available dataset, rather than including any actual dat integration. I'm planning on changing that though once dat becomes just a bit more stable.
[0] source: https://github.com/kcorbitt/datrepo [1] current site: https://datrepo.com/about
[1]: https://github.com/maxogden/dat
[2]: https://github.com/maxogden/dat/blob/master/usage.md
[3]: https://github.com/maxogden/dat/blob/master/what-is-dat.md
We use git to push configuration data to (near) real-time systems that keep this data in local memory mapped key-value stores. It's essentially a fully replicated, eventually consistent key-value store, and git plays a staring role.
We regularly push multi-gigabyte files through this system with ease, with only a few tweaks on git's configuration. Git has some major advantages. It's fast, can handle large files, and you can piggyback on the versioning system to ensure that multiple writers aren't competing. Also, git is highly configurable and has a lot of out-of-the-box options for easily setting up git servers.