An introduction to IPFS
medium.com
medium.com
The claim is that IPFS could replace HTTP, the web, and so on. The only thing I see, however, is a distributed filesystem, which is only one part of the puzzle. Real-world applications require backend systems with access control, mutable data, certain information being kept secret, and so on - something that seems fundamentally at odds with the design of IPFS.
How would IPFS cover all of these cases? As it stands, it essentially looks to me like a Tahoe-LAFS variant that's more tailored towards large networks - but for it to "replace HTTP", it will have to not only cover every existing HTTP usecase, but also do so without introducing significant complexity for developers.
Seriously, I'd like to see an answer to this, regardless of whether it's a practical solution or a remark along the lines of "this isn't possible". I'm getting fairly tired of the hype around IPFS, with seemingly none of its users understanding the limitations or how it fits (or doesn't fit) into existing workflows. I can't really take it seriously until I see technical arguments.
From just a quick look (I'm not IPFS expert):
IPNS is used as the mutable state driving 'dynamic' content. The dynamic content is still published as a normal immutable reference, but is referenced by an IPNS name/path/hash. To me this appears to be like a CNAME in dns. IPNS provides the pointer to the immutible hash, and for dynamic content you would only ever use the IPNS.
For authorization, this could be done via PKI. Take some data, encrypt it with one or more user's public keys, then post it to IPFS. Then the recipient, and only the recipient, would be able to access the content.
I can't really take it seriously until I see technical arguments.
Well, to be fair, the technical arguments appear to exist (after all there is a IPFS draft document which covers the stuff I mention above), and while something like replacing HTTP is obviously extreme, those arguments appear (on their face) to not be completely crazy. So maybe if there are specific technical issues that aren't covered by the draft documents and other documentation on the IPFS site, github faq page, etc, then maybe bring them up and people can help you answer those questions.
In this context, IPFS is part of the toolkit for building decentralized applications. The architecture of these applications is fundamentally different from today's apps. Users control the data, not the application developer. Private data can live on the user's machine or some cloud service they trust to hold the data for them. IPFS probably isn't a good fit for private data, but public apps like a decentralized Twitter will probably use it for all of their data. Encrypting private data and storing it with IPFS will sometimes make sense.
To deal with mutable data, it depends what your goals are. If users don't need to agree with each other about the value of the data, they can store it however they want and maintain their own view of the world. If they need to agree on data that only one person can change, then you can use IPNS to manage mutable references controlled by one keypair. If users need to agree on data about their interactions with each other, then they need to use a decentralized cloud computer that can enforce rules on a historical record. We call these enforcement clouds "blockchains," and Ethereum is a blockchain that defines a process for defining new rules so application developers don't have to build their own blockchains.
Building decentralized apps with IPFS and Ethereum is the most exciting work I've ever done. It doesn't just improve the apps themselves, it fundamentally changes the economics of our industry. Most businesses are built on network effects, but decentralized apps don't bottle up the value of their networks: they give them to their users.
That's more like the world I want to live in. If you want to be a part of this change, you should join us at ConsenSys, or join the decentralization movement in general.
The future of cryptocurrencies is bright, there's now a large set of tools in the toolbox to use to build resilient, permanent applications for people to use. I'm really excited to see what people build in 2016, the cryptocurrency ecosystem has been growing rapidly in the shadows, and they're finally breaking into the mainstream.
https://www.reddit.com/r/ethereum/comments/46fz8f/ethereum_i...
To decentralize private search, we'll build protocols for publishing your private data to cloud services that represent your interests. For some purposes, you'll publish data to services that keep your data isolated. You might even publish that data to your own machines. This works for email and cloud storage like Dropbox. Other purposes require data to be combined to add value, like recommendation systems and AI doctors. Services will compete with each other for your trust so you'll pick them. The bigger they are, the more useful they will be, so they're vulnerable to collapsing back into a monopoly. Anti-monopolists will give their business to competitors to keep them alive.
That is nice in theory, but not how it works in practice. Searching public data generally involves a significant amount of investment, both financially and effort-wise, to have anything approaching a usable search engine, even for relatively small sets of data.
Easy. Do not make sites and claim they are apps, when clearly they are acting as graphical terminals to a remote service.
An app should be something that runs locally, stores the result locally, and optionally stores things remotely as a backup measure.
The whole "lets do all the heavy lifting out there" approach is a throwback to the leased terminal era. And as best i recall, the personal computer became popular because accountants would rather run their spreadsheet (even if slower and with severe data limits) locally than have to grovel to the admins for mainframe time.
> IPFS probably isn't a good fit for private data, but public apps like a decentralized Twitter will probably use it for all of their data. Encrypting private data and storing it with IPFS will sometimes make sense.
This is actually a great example of where I can't see the IPFS model working, because not everything is black-and-white private-or-not.
For Twitter, for example, you have private accounts - this doesn't mean that nobody can see it, it just means that you need to have permission to see it. How do you implement an access control system that:
- Is mutable (ie. allowing new users to see your feed)
- Is private (ie. there is no public list of who can see your feed)
- Is revokable (ie. you can remove people's access to your feed)
... without having to re-encrypt everything using a secret set of keypairs every time the ACL changes, which is an unreasonable burden on resources?
I understand that you build applications on a different kind of architecture, and I'm all for usable decentralized primitives, but there have to be explicit solutions for these kind of 'business requirements' (whether commercial or not), for it to ever be able to "take over".
With such a database system, it would be trivial to build a decentralized forum, for example.
The tooling merely provides a means of viewing Merkle DAGs (the base data structure) as a file hierarchy because it's a pretty clean 1:1 mapping that's familiar to many people.
If you've ever worked with stateless servers using something like JWT, you will have run into the issue of 'stale data', essentially the cache invalidation problem. Trying to solve ACLs with cryptography is prone to the exact same issues, as well as overhead problems. The crypto itself is immutable.
EDIT: In fact, Git is a great example of how you can't really do access control very easily in such a setup. Have you ever tried restricting access to a specific branch in a repository?
wish wolfram was open source
I don't think that's at odds with the design.
If you want to keep something secret, you encrypt it. If you want to manage access control, you sign it. A client application can discard information that cannot be decrypted, or isn't signed by an authorised key.
You don't get atomicity guarantees of course, but you can often get away without them. I don't see a system like IPFS as a replacement for all of the web, just a sizeable percentage of it.
This is illogical, as evidenced by the illogical replies you are getting below. You are asking in a "serious" way for the IPFS community, developers or Juan to explain how all current use cases would work on IPFS.
That's a pretty tall order for an Open Source project that (seriously) owes you nothing.
I would suggest you attempt to run an IPFS node yourself and write a few applications and tell us how it goes.
EDIT: Also, please get out with the "owes you nothing" argument, I'm really growing sick of it. If I see something generating undeserved hype with potentially bad consequences for 'society at large', then I will damn well call it out. Don't like it, don't reply to it. It has precisely nothing to do with anybody owing me anything.
- https://github.com/ipfs/faq/issues
- https://github.com/ipfs/notes/issues
- https://github.com/ipfs/apps/issues
The technical answers exist. Please look for them! I'd love to write personalized answers for every person that asks, but i better spend my time developing. Also you should look at the rest of the discussions too, not just my compressed answers. be part of the discussion, raise your concerns, and have them answered.
For sake of saving you time, some short / compressed answers:
- mutable data: look into ipns (in the ipfs paper or repos linked above) and look into CRDTs. CRDTs can layer cleanly on top of ipfs for distributed mutable state. if you haven't seen CRDTs: http://hal.upmc.fr/inria-00555588/document
- access controls / certain info remaining secret: data encryption and capabilities. you mention Tahoe-LAFS, yep!, we can (and will) build a cap system similar to it on top of raw ipfs objects (ideally actually collaborating with the excellent Tahoe-LAFS team directly). --- oh and many users want easy selective disclosure for encrypted multiparty records. we've not built this in yet because we have other priorities right now. if you'd like to contribute to this we'd love the help!
- "it will have to not only cover every existing HTTP usecase" -- we don't aim to cover _every_ HTTP facility, instead we can power a nicer model for distributed data, where you replicate datastructures directly. This looks a lot closer to the data models people use in apps today (think single page app models, {backbone, react, meteor, ...}. these are layered on top of APIs like REST, but can trivially layer over IPFS too)
- pub/sub: you didnt mention it, but it is necessary for fast (subsecond) updates to mutable data in the large (>millions of nodes). we're working on these designs now, and impls coming in a few weeks/months. (some simple non-scalable versions already proposed and implemented to get started, but the long term solutions coming later).
I ask you please take a deeper look at the comm channels of our community, and that you voice your concerns + questions in our FAQs and other relevant repos. That way we can improve our models and systems based on your input. And if you have time to help us implement some ideas, join us! :)
more links at this answer https://news.ycombinator.com/item?id=11137313
> you mention Tahoe-LAFS, yep!, we can (and will) build a cap system similar to it on top of raw ipfs objects (ideally actually collaborating with the excellent Tahoe-LAFS team directly).
The cap system that Tahoe-LAFS uses is great, but only serves a limited subset of usecases. It is not possible to revoke access, for example.
Depending on implementation it may be possible to accomplish this (to a degree) by having a separate mutable pointer for each person who has access to the data - and simply not updating their pointer to newer versions once their access has been revoked - but this still doesn't cover the "oops, accidentally gave them access, let's hope they didn't notice and revert it" usecase.
(EDIT: It's not even possible to delete files on demand in Tahoe-LAFS, at the moment.)
> This looks a lot closer to the data models people use in apps today (think single page app models, {backbone, react, meteor, ...}. these are layered on top of APIs like REST, but can trivially layer over IPFS too)
That architecture has actually turned out to be flawed in quite a few ways - in many implementations, it has turned out to be considerably harder to build on top of them than on more 'traditional' architectures. SPAs are common primarily due to hype, not because they are the superior technical solution.
They also come with some pretty serious (and in my view, unacceptable) trade-offs, like requiring JavaScript to be able to use it.
IPFS is definitely not a filesystem in the classical (POSIX-y) sense. I think the naming is unfortunate, because it implies that IPFS has things a global root directory, permission bits, users, ACLs, immutable human-meaningful names that resolve to mutable content, a POSIX-y consistency model, etc. IPFS does none of these things--it defers them to layers above it. All IPFS is concerned with is providing a Merkle DAG abstraction, and the necessary network protocols to walk it and fetch the underlying content.
My biggest concern is that in the end IPFS isn't even really "permanent" in the way I understand it. Objects added to IPFS still need someone to in a sense "seed" them for that content to be available. What advantages does that give over just hosting the internet over static torrents?
- ipfs-cluster - discussion on a tool we'll build to replicate archives to many nodes and have a RAID-like cluster configuration - https://github.com/ipfs/notes/issues/58
- ipfs-persistence-consortium - https://github.com/pipermerriam/ipfs-persistence-consortium
- pincoop - https://github.com/victorbjelkholm/pincoop
- http://filecoin.io (long term solution)
Authenticity over P2P is complex, that's the cost of having no server. But if such a cost can save you hardware and server hosting, it's worth it.
> What advantages does that give over just hosting the internet over static torrents?
No need for DNS. That is pretty huge.
The vast majority of this discussion is miles above my head, but this made my Spidey Sense tingle. Someone somewhere has to host the data, and seemingly in more than one place. There is no such thing as a free lunch.
While the IPFS whitepaper and spec does utilize a lot of modern technology terms, it does not mean they were assembled without serious consideration and purpose. What gives you that impression?
All of those pieces have important properties which create the essence of what IPFS is. When you say Blockchain, Git, BitTorrent, I hear: Directed acyclic graphs (like git) with hashed hierarchical checkpoints (like a blockchain or merkle trees) distributed peer to peer (like bit torrent). This is literally what IPFS is, and using those terms is one way to describe it.
This was funny. Suppose you wanted to build a node that linked to itself. You'd have to find a fixed point in the combination of functions that adds other data to the link and hashes it. Finding a fixed point of a hashing function is hard.
Using the big-step, little-step cycle detection algorithm to avoid using gigantic amounts of memory, you're then looking at an average of 1.5 * 2^129 iterations of updating your graph of 256-bit cryptographic hashes in order to discover you've hit a periodic point.
Offhand, I don't know the probability that there's a fixed point for a given starting point for a random mapping of 256-bit values to 256-bit values, but my intuition is that it's vanishingly small. If anyone has an elegant derivation of the probability, I'd love to see it.
This doesn't tell you anything about concentration bounds or whatever, but it's a neat fact nonetheless.
Unfortunately, it's a random mapping, not necessarily a random permutation. An ideal block cipher would be modeled as a random permutation. Though, in this particular case the domain and range are the same, so unless I'm missing something, the expected number of fixed points comes out to 1 by the same reasoning.
Note though that a correct block cipher is necessarily a permutation, because it's invertible (by definition, a permutation is just an invertible mapping with domain and range equal). A hash function on the other hand needn't be a permutation even when you restrict the domain to inputs of the same bit length as the hash output.
The only way we are able to productively use git is because there is a convention to have some state in a non content-addressable location (.git/refs, .git/HEAD, etc...).
Saying that IPFS could replace the web means either: 1) Introducing shared mutable state; or 2) full knowledge of everything on the network.
I'm guessing that the existing web is what provides that layer right now. Is there any work going on for novel IPFS-based content discovery mechanisms?
Another thought: Given the content-addressable, immutable nature of this graph, how does one discover that a new version of something is available without a central authority? How could we discover the tip of a blockchain with IPFS alone?
IPFS needs not be a kitchen sink: it's very easy to layer entirely application-layer overlay networks that do other useful things.
git, hg, monotone, ..., all offer(ed) similar feature sets -- dvcs. yet they're very different tools.
Immutable is an interesting idea - it's a lot less interesting when 100 different copies of the same slightly changed RAW file from my digital camera are using up 100s of gigabytes. Or I misclick and something goes into the public store which shouldn't - I might not be able to get rid of all of it, but I should be able to undo it a little.
Regarding misclicks, for camlistore everything is private by default; in order to make something public you have to build an authorization, which you give to someone. You also have the possibility to remove that authorization, and the data won't be publicly accessible anymore.
Some more links for people to check out:
## (upcoming) IPLD "merkleized JSON" format:
- improves upon our basic format to make it much more pleasant to build things on top of ipfs.
- JSON meets CBOR meets Merkle-linking
- mini-spec: https://github.com/ipfs/specs/blob/master/merkledag/ipld.md
## answers to some common questions i've read on this page:
- content model / replication: https://github.com/ipfs/faq/issues/47
- how resolution works: https://github.com/ipfs/faq/issues/48#issuecomment-152917088
- how IPNS / mutable linking works: https://github.com/ipfs/faq/issues/16
- this is a very poor answer, sorry, i'll write up a post or paper on it.
- for now if interested, see the QConf slides below, specifically slides ~110 to ~130 -- the DNS, IPRS, SFS/Mazieres linking, IPNS parts.
## These repos have interesting "lab notebook" style discussions:
- https://github.com/ipfs/notes/issues
- https://github.com/ipfs/apps/issues
## deep dive talk at stanford:
- video: https://www.youtube.com/watch?v=HUVmypx9HGI
- slides: lmk if you want them, i'll pdf them up
## talk at ethereum's devcon1 covering blockchain uses
- video: https://www.youtube.com/watch?v=ewpIi1y_KDc
- slides (interesting bits start at slide ~70): https://ipfs.io/ipfs/QmUgRq7QfmRbPw5kXqwSs1TRtPDBXMoDNiYwJQg...
## talk at qconf sf (similar to above)
- in this talk i discuss a bunch of datastructure stuff, including using IPFS for PKI, for arbitrary dns-like records, for name systems, for CRDTs, and so on.
- unfort video will be released in march: https://qconsf.com/video-schedule
- slides (intersting bits starts at slide 80): https://ipfs.io/ipfs/QmPpYmdSEKspjgXxVyGK9UMHV54fKZS8MwJjppg...
storj is for paying people to store your personal data and ipfs is more for public data.
Of course you can use ipfs for private data too
"testing 123\n" isn't anywhere, and "Hello World" (and its hash) is pictured twice. I'm sure that the testing.txt arrow should just be pointing to a node with a different hash and content.