Camlistore – open-source personal storage system for life
camlistore.org
camlistore.org
>Your data should be alive in 80 years, especially if you are
Is there any special significane for 80 years? And exactly how is it safe? The "Download" sections have me installing the server on my box, and my hdd certainly won't survive 80 years.
Under storage, it says that
>Implementations are trivial and exist for local disk, Amazon S3, Google Storage, etc.
How does encryption work? I can't seem to find out at which step the encryption is done (client side, or server side?). The home page mentions it's "private by default", considering that I don't know that much about security and cryptography in general, how safe would I be in using this, for whatever purposes that it could be used for.
Mention under potential use cases: filesystems backups and Document management CMS. That's ... interesting.
I believe Private by default just indicates that objects are not exposed until they are shared: like a file system is private by default and a Github repo is not.
What's cool is that you can abstract where it's stored... want to use S3? Point your data there. Google Drive? Point your data there. Instead of the user interface and middleware of syncing and storing being directly tied to a storage backend (a la Dropbox, Google Drive), Camlistore allows you to separate it.
At least that's how I think of it... I might be wrong. :)
I imagine that description getting better when they start wanting more end users to use it.
Right now they have a bunch of pages which are html with 0 css styling, cause it ain't done yet. Despite these things, I find camlistore extremely cool, and look forward to it's growth.
I personally think this is one of the most hugely important ideas/protocols to come out of the last decade. Even if Camlistore doesn't do it, it's hard to imagine software programmers of the future not agreeing on a shared protocol to fill this huge, huge need.
I actually wrote a proposal for a Mozilla grant recently outlining a couple of the reasons why decoupling storage from user interface is a fundamentally good thing for society[1].
[1]https://www.newschallenge.org/challenge/2014/submissions/sto...
I get what they seem to think they're doing.
But "storing arbitrary" data is a problem everyone in the world is trying to solve. It's what a filesystem is. It's what backup tools do. It's a pretty well attacked problem.
A bunch of SHA1 hash addressed data is almost as useless as a raw disk with no filesystem.
Similarly, I'm seeing a lot of JSON. Who says JSON will be remembering in 80 years? It's ASCII! But that's not much of a guarantee.
There's a lot of claims here, which don't seem to translate into something which seems actually useful. For example, "immutable" sounds good till you run out of disk space because you have 1000 copies of a sightly different VM image in the system. It's pretty easy to build a system which stores "everything". It's a lot harder to build one which is useful.
Camlistore has always been a project that I'd like to try out someday but every time I've looked it's seemed more hypothetical than real. Kenneth Reitz (of python requests fame, among other things) put together a much smaller thing called Elephant[1] that I've been tempted to explore as well, and it's sort of in the same vein.
>> Suddenly, your data becomes as durable as S3
Seems slightly less ambitious.
Anyone know if these guys are working with existing archival groups and standards? It would be a shame if they're reinventing the wheel.
Very nicely written description of something useful. IMO much better than the words on that website.
What I don't like about the website is the "jargon". These words are from just the first paragraph:
formats
protocols
modeling
synchronizing
post-PC
objects
FUSE
Huh?The website text was probably written by an engineer. Reminds me of a great quote:
"engineers are all basically high-functioning
autistics who have no idea how normal people
do stuff" - Cory DoctorowMy general counter argument to jargon filled websites is the great ending of Trading Places:
What about lunch?
The lobster or the cracked crab?
What do you think?
Can't we have both?
Why can't a website have both a clear explanation for normal people, followed by all the jargon necessary for the target audience?IMO there's a sine-qua-non that most websites should have, and that unfortunately far too few do have. Well known companies like Apple and Google don't need it, but most others do. It's what people have called an elevator pitch. Here's how Wiki puts it [1]:
An elevator pitch, elevator speech, or elevator
statement is a short summary used to quickly and
simply define a person, profession, product,
service, organization or event and its value
proposition.
That information should be at the top of a website. It's what tells people, immediately, what the product or site does and how it could be useful for them.And, not coincidentally, the words I quoted from Wiki are the totality of the first paragraph for that Wiki entry. Simple, clear, easy to understand.
There's nothing wrong with having details on a company's website. But that jargon should not be the totality of the first paragraph on the site.
As a tool it's a bit rough yet, but it's going to be awesome.
Disclaimer : I have a man crush with Brad Fitzpatrick. Well, mostly with his code.
Where I struggle is that either the definition of "your data" is narrow, or I shouldn't be using it for all my data.
Back in 1999 when I first learned about MP3, I started ripping my CDs. I have several thousand CDs, this took a lot of time. Before I completed the task, at a rate of a few CDs each evening, FLACs came into my life and I started back at the beginning. I deleted the MP3s as I replaced them with FLACs.
I really don't ever need to keep some data. But maybe it's not the kind of data that I should be putting in Camlistore? I think of it as my data, after all these are my CDs.
I struggle with the concept of Camlistore as I have an 18TB NAS in RAID6, 12TB usable... and it's 80% full. If I had history I'd have a storage problem today.
I'm perhaps an outlier, I chose to self-host my data locally rather than rely on cloud based things. And I chose to keep everything... photos, documents, email, video, music. And everything I keep is in the highest possible quality: FLACs, DVD VOBs, raw photos, etc.
But then... who is Camlistore aimed at if not the people who like to store and have control over their own data?
I guess I just find delete too valuable a feature for the larger data I store.
And perhaps I'm just wrong on the use-case, maybe it's really "for all your data (that you cannot re-acquire)". I just don't want to ever rip those CDs again. But if I do, those old versions are dead to me.
However, even git will delete data if you delete the "tree" metadata, ie you nuke some branch that has no downstream dependencies because you never merged it or there are no branches off of it. In that case, if the blobs aren't reachable by any tree/graph, git can garbage collect those blobs.
Camlistore does the same thing: if you delete all pointers to the data, those blobs might eventually be reclaimed. As a matter of implementation, camlistore doesn't do that today, but it's not the case that camlistore can't or won't let you delete data.
In addition to that, I think that there is not such a thing as space limits, at least, generally (not always, of course).
Specifically, one would imagine that the general attitude is 'I have to store such and such - where do I find the space'?
Instead, I think that in general, it is 'Oh, I have such space for free/cheap price... let"s store stuff!'. Especially with the advent of the Terabytes order of magnitude, I guess most of the storage is simply composed of movies, even if they're never watched, either once or more times.
Again, pay much attention to the bias. As an amateur photographer, I'm tempted to think that "raw is the law", but in fact, for the vast majority of people, it simply isn't.
I've only met one OCaml programmer in my life, and they were a graduate student and rather strange.
OCaml users: http://ocaml.org/learn/companies.html
FB Hack: http://cufp.org/2013/julien-verlaguet-facebook-analyzing-php...
JaneStreet: https://blogs.janestreet.com/category/ocaml/
git-annex stores things content addressed, gives me different views into the data (tags, etc), and supports different back-ends (S3, remote rsync, local filesystem, external disks, etc). Isn't this exactly what is described here?
[1] http://tent.io
If it's the site, that's on me and I have a bunch of work queued up to better represent the tools. Specific feedback would be welcome.
If it's about the tools, the first UI is the command line as we expect these to be components that developers can use. In terms of the initial applications we refer to, they would effectively be CardDAV and CalDAV servers so you'd hook in your existing apps to them.
http://webcache.googleusercontent.com/search?q=cache:P_L5IWO...
What is this doing that solves so many problems that apparently they can't outline what it is actually doing on the frontpage of the website in clear language?
What does "camli" means?
Content Addressable: What things are named depends on their content. Two identical things have the same name. For example, the "name" or "key" for the data is the SHA-1 for the data, ala git.
Multi-Layer: The whole storage stack is built out of several layers. The blob store sits on the bottom, and only knows about bytes, and access is via the SHA-1 of those bytes. Things that you might store (Files, directories, sets, collections of tweets, social graphs, etc) build on top of the blob store by additional blobs that hold pointers to data blobs. Again, it's sort of like git. A front-end might sit on top of that abstraction.
Indexed: blobs of JSON that have a few special attributes are recognized and indexed. So, you might have a bunch of blobs with these special attributes (ie, "tag") and be able to ask the indexer "Give me all blobs with tag equal to foo", rather than having to search through the blobs directly.