Filecoin: Proof of Storage Systems
blog.coinlist.co
blog.coinlist.co
S3 storage pricing, for most users is just a rounding error compared to e.g egress or compute cost. And it has proven to be tremendously reliable, both in durability and access. No sane person would point a CDN against a Filecoin backed origin, right?
For kinda-censorship resistant distribution of stuff like Wikileak dumps we already have torrents right now, where the seeders are aware of the content they spread.
For uses where S3 is not competitive, like storage of my media rips, spinning rust at places like Hetzner exist, or Backblaze et. al for more casual backup uses.
The only point I can see for decentralised "storage" of this kind is for ultra high durability, small data size, almost never read, almost some sort of "stakes". One could imagine something like storing contracts/land deeds, similar to Gwern's timestamping of URL's. The files should be small, the storage model optimised for maximal durability (every node store's almost everything?), and prices high enough to ensure durability.
[0]: Just look at this PR page https://filecoin.io/store/ [1]: https://www.gwern.net/Timestamping
> S3 storage pricing, for most users is just a rounding error compared to e.g egress or compute cost. And it has proven to be tremendously reliable, both in durability and access.
It's not like you can separate S3 storage costs from egress costs, though. I wish I could use S3 for storage but pay for my own bandwidth.
Even if you'll to go all the way to their network and plug a network cable in, well... They charge $7.40 a day for a gigabit port, and $54 a day for 10gig. That's on par with what many companies charge for transit to the actual internet, but fair enough. Oh, wait, I neglected the per-byte cost on top. The per byte cost that, for a mostly saturated port, is $210 or $2100 per day respectively.
And that's the cheap option.
But I guess it also depends a bit on your use case? E.g if you upload big datasets for heavy computation and then only download small results (≈ machine learning) or have a cheap(er) CDN in front of S3 with a favourable usage pattern, the storage and egress would be somewhat separate.
Still, something like Filecoin is obviously not a resonable solution for the run of the mill "cheaper than S3 file hosting and network" problem?
(I've been on the lookout for something like that, for personal use where I don't all those nines and other S3 benefits. The best thing yet seem to be Wasabi [0])
when somebody will push into Filecoin some "bad" files (anything, from docs needed by the mafia to children being sexually abused, to fake videos/claims/whatever about you, to any other bad things you can think of), how can that be detected and then deleted? (ignoring here the theme related to accountability)
Question based on https://docs.filecoin.io/introduction/why-filecoin/ :
Filecoin resists censorship because there is no central provider that can be coerced into deleting files or withholding service. The network is made up of many different computers run by many different people and organizations. Faulty or malicious actors are noticed by the network and removed automatically.
In this context I'm not a believer of "automatically".
It should be solved by allowing an "expensive" (a) "vote" (b) to override, or delete, or hide for a number of years.
(a): expensive should discourage the option in general (b): proof-of-stake, proof-of-ownership, proof-of-work are some of the many ways in which you can define voting rights.
Unfortunately, it doesn't work easily in practice. Example:
You created a file (say, a CV, or a newspaper article in PDF) to discredit me, but it's a lie. To me, not having that file publicly accessible would be worth $1,000.
You could blackmail me - give me $1,000 or I'll upload the file! But I think you might do it anyway, so I don't negotiate (never negotiate!!). The file is 1 MB, and it costs $1/year to keep it up and running on Filecoin.
I should have an option to ask a voting pool to remove that file for $100. The pool agrees, and the money gets burned (I shouldn't create any incentive for someone to cash in the amount).
However, you can upload the file again; or a slightly different version, and we are back to square one.
It then goes to the "discoverability" of the file. That part is probably going to be decentralized too, so no court could "order" Filecoin to make a file not discoverable.
How do we solve it? I think there will emerge a search system that provides services to censor, obfuscate, etc, a number of files, based on reasonable requests. Whoever builds that company, I bet it will not be poor.
Not sure...
But overall the root problem would still exist and the solutions would have to gather a majority - e.g. asking a "pool" to remove a video would mean for the "pool" to watch that video (to ensure that the claim is valid) and after the first experience (if it's child porn) we would all probably have to call a psychiatrist to help us overcome what we've seen => not feasible, it would actually destroy the pool's mental stability which is, in this concept, needed to evaluate the contents (videos in this example).
Edit:
Being ignored by some search engine would still be a no-go for me. Knowing that some personal data (general files, pics, videos, whatever) exists somewhere would make me feel absolutely NOK.
I don't think Filecoin can be MORE efficient than some of the larger, current systems.
I actually think this statement is misleading. It should say that Filecoin's goal is to offer storage options that are distributed (hence redundant), protected from censorship, and possibly removed from absolute control of a single large corporation.
This is an oversimplification: it assumes that discovery is very low cost and that there's a peer with a copy enough closer on the network that it's faster than talking to server-class hardware in a data center with a high quality network connection. Given the number of assumptions which need to be true for that to be a net-positive I'm skeptical that it'd be easy to hit anywhere close to the best-case theoretical scenario.
This is admittedly a non-trivial task but perhaps an example can shed light on way it can be more efficient in practice.
Consider the Apple Photos app. It's on your phone, it's on your desktop, it's on the cloud (iCloud storage). Say you have so many photos they don't fit on your phone. When you try to access a photo that is not on your phone it gets it from icloud (these are literally features that are currently offered by photos/iCloud). Now suppose your desktop has lots of storage, you usually look at your photos while connected to your local network, why cant you get the data from there?
The key to the efficiency of file coin is not some magic cryptographic token, it's not even "blockchain", the key is the underlying protocols being unified for synchronizing devices across multiple tiers of networking. Filecoin mostly just makes sure the data doesn't disappear and gives you an optional substitute for cloud storage. Content addressing, p2p communication by default, flexible/future-proof standards for negotiated what data is and how it is formatted, and in modular plug-and-play utilization of many different communication/synchronization techniques is the key to applications being built in more effective and efficient way than conventional client/server architecture.
The word that comes to mind is "flexibility". Want you store to be S3 buckets? Great! Want to host a private network, only your devices can access? Great! Want to joint the public network? Great! Want to ensure persistence of data via a token? Great! Want to just use the address mechanism? Great! Want to run the software in the cloud, on you local machine, in a sandboxed browser tab and have them all seamlessly forming a p2p network over a multitude of network standards (tcp, udp, web-sockets, etc)? Great!
I worked extensively last year with various ipfs techs utilized in a private p2p network consisting of ARM nodes capture time-series measurement devices and cloud based persistent nodes managing the pinset of data and persisting long-term storage. There are definitely some rough edges but again it's open-source built is such a modular way it's almost difficult to develop your self into a corner.
The software coming out of protocol labs is overwhelmingly not new technology. Protocol lab large focuses on utilization of battle-tested existing standards and tech (many of which are relatively ancient like TCP, ssh, git, json, etc) unified, without breaking existing stands, via a common addressing mechanism.
tl;dr Filecoin/crypto-tokens just a side note! The vast majority of software coming out of protocol labs is independently useful and largely focuses on bridging together widely-used battle-tested tech.
Yes, what Dropbox, Crashplan, etc. offered a decade ago — it's a neat sounding idea but there are two reasons why this isn't a huge win in practice: people are mobile so the number of times where you need an uncached file and happen to be on the same network is relatively low and increasingly few people have an always-on device with a ton of local storage free (phones are really good at generating high volumes of data so this is a non-trivial problem).
Working on things like this is really interesting but it involves a lot of work to handle unreliable clients or networks and bitrot (hashes don't solve this if your client helpfully replicates the bad sector from your desktop over the pristine copy on your phone). That overhead makes it a lot harder to beat conventional services, especially in cases where the time investment is greater than the possible savings.
The idea that ipfs/related-project are trying to beat conventional services is big misconception. The overwhelming majority of the tech does not preclude usage in conventional services. For example, IPFS can 100% be configured as server/client offering roughly the same costs/reliability many are used to. The advantage is interoperability with many different "services", conventional and alternative (e.g. filecoin) alike.
An "address" is also self validating. Hash addressing, where a hash of the data is the address of the data, are use to accomplish this. This is useful in many contexts and an [extra few bits][1] on the the front allow an address to represent much much more.
"location" is by default a distributed hash table enabling the fulling distributed routing of content and by uses [Kademlia][0]. But the software could easily be configured to with fixed values for the hash table mapping any "address" to the same location which happens to be a cloud provider.
Could you elaborate on what you mean by "anyone who keeps their API locked down is unlikely to adopt standard IPFS."?
[1]: https://en.wikipedia.org/wiki/Kademlia [2]: https://github.com/multiformats/cid#how-does-it-work
Secondary markets for drives may appear to close the gap a bit though.
I remember a while ago hearing about both trying to decentralize storage but never kept up with it enough to really suss out the difference
If you wanted to build a decentralized SoundCloud you couldn't use Sia but IPFS/Filecoin is the primary building block of such a dapp.
If you do want something like Dropbox that is built on Sia, there is Filebase[4].
I am not associated with Sia at all. I've just been following their twitter for a while.
[1]: https://siasky.net/
[2]: https://nebulouslabs.github.io/skynet-docs/#introduction
[3]: https://twitter.com/SiaTechHQ/status/1291441690008592384
Filecoin seems to be taking an approach more like Holochain stuff. Meaning they are shipping early and often but only internally and you'll get to see it when its ready™.
Filecoin raised over $200m and hasn't shipped anything that I'm aware of in 3 years. Maybe it will be revolutionary, but the crypto industry has long moved on from white whale projects that don't ship.
But to build a team around a project? You can hack a lot of stuff together than only makes sense to you, crippling any of your would-be collaborators. Certain kinds of doubts should in fact stop you, or at least slow you down. It's a difficult thing to balance, and most of us struggle to pull it off.
I tried to help out with Freenet years ago. The code was in a language I had quite a lot of deep knowledge about, I should have been able to help quite a bit, but I ended up sticking to specs and architecture.
Bluntly, the code was not the output of an organized mind. Even though most of it was written by one person. It ping-ponged all over. The prototype was already full of esoteric optimizations that sacrificed clarity (and thus correctness) for a little more speed. Major rafts of functionality were still being designed, and it was already work-hardened.
When I see a team ramp up and get wedged, my first questions are about code quality, practices, and team dynamics. About cohesion. Too often the old guard can't delegate, even if they want to. With enough tribal knowledge, your coworkers just start to get interesting, and then they move on to a new job, and you replace them with another person who struggles to accomplish even simple things.
Or even worse, they made a promise that Information Theory says they can't keep, and they've been bargaining the entire time. It's a terrible place to be.
The other is some sort of authentication. Encryption of the data itself, IMO, is not good enough. In 10 years time I don't want a nation state or megacorporation with enough entangled qubits to be able to pwn all of my data. But distributed authentication as I understand it is a hard problem, because you probably need to make your authentication proof public, which creates a similar set of problems on the authentication itself.
One is that frequently verifying availability is going to get expensive quick with iops for coldish storage systems.
Secondly, iirc symmetric encryption isn't really vulnerable to quantum algorithms, as far as we know (and maybe even proven not to be?). It's asymmetric crypto that's at risk.
By erasure coding your data over many providers, you minimize the amount of harm that a couple of bad apples can cause.
For example, the phrase "decentralized marketplace of storage providers and choose the one" was just glossed over -- through what mechanism does one choose a storage provider? Is there an API that (for example) a container storage interface provider could consume? like that kind of "choose a provider"?
I also would have enjoyed more hyperlinks for terms such as zk-SNARK, and "a storage miner" which is similarly glossed over
(Submitted title was "Deep Dive into Filecoin".)