Filecoin: A Decentralized Storage Network [pdf]
filecoin.io
filecoin.io
There are other schemes that don't require a copy of each shard, allowing you to reconstruct the file if any subset of the nodes is online.
Any subset? Surely there's some minimum number of nodes (>1) that need to be online. Unless each node has a complete copy.
I think Sia's method would prove to be extremely reliable, but again, it depends on the hosters. If the largest single hoster owns less than a third of the capacity of the network, even if they suddenly go offline, you should be able to recover your file since you rely only on a third of the 30 hosts who hold your file to be online.
Also, I think that the number of pieces you need to recover a file can be tuned, so if 33% would prove to not be reliable enough, Sia/Filecoin can tune it down to increase reliability.
Filecoin (and other decentralized storage systems) provide monetary rewards and penalties for storing or failing to store files. This is introduced in Section 4.1.1 of this paper.
The bad thing is that you would not be able to hold them responsible.
And on private torrent sites, where there is an incentive (usually non-monetary) to seed torrents, torrents really don't die all that often.
To go even further, if you orchestrate several miners carefully (maybe computationally) they can together perhaps stop people from accessing certain files.
It'd blame Apple & Google for not being p2p seeds.
For example, with a decentralized system files could be broken up into shards and distributed across multiple nodes, so unless a significant percentage of computers on the planet all melt down at once your data would remain accessible (i.e. zero downtime). That would also allow downloads to be parallelized across multiple connections like with BitTorrent, meaning bandwidth wouldn't be an issue either.
And that's not even considering the effect that a commoditized pricing system would have on the costs of a distributed storage solution. I imagine it'd be significantly cheaper than AWS too, as storage would be sold at a price very close to costs.
On the other hand, you could spend some money and guarantee it instead: https://aws.amazon.com/directconnect/
As for price, AWS has some of the cheapest operating costs due to scale, something which smaller players will have trouble competing with. Moving data within AWS is fast and sometimes free.
Smaller players also pay the highest cost for bandwidth. If some plans don't have bandwidth caps yet (many already do), they will get caps as soon as serving bandwidth as an individual becomes profitable.
Pretty much all distributed systems have the following in common: You pay for resilience with overhead. If it wasn't that way, everything would become as distributed as possible over time.
Isn't it the other way around? Small providers offer low traffic costs while the big cloud providers (Google, AWS, Azure) charge significantly for bandwidth.
You can setup any old distributed database to do this... If there was a benefit to breaking down a file rather than replicating it, AWS would be using this efficiency to provide a better service.
> the effect that a commoditized pricing system would have on the costs of a distributed storage solution
It'll always be cheaper for someone to sell their hard drive than to sell access to their hard drive, just like it is cheaper to buy bitcoin instead of mining it.
Here is why: let's say Alice is profit-driven and greedy and is making $1 profit for 1gb of space. She can use those profits to buy more hardware, make more profit, buy more hardware, and so on and so forth until she has 100,000gb of space being offered. At that point, because of economies of scale, it will cost Alice less to maintain 100,000gb then the total cost of 100,000 people hosting 1gb each (eg. reduce total electricity cost because of bulk purchase).
Now, Alice can grow her operation so big (say, 100 petabytes) that it actually becomes no longer profitable for Bob to host 1gb. Alice has so much lower costs per gb that she was able to drive the price per gb down to a point where the price is lower than Bob's costs.
This is basic economy of scale so far, right? And it's the same reason why data centers filled with ASICs make it useless for you to mine Bitcoin on your laptop.
OK now that Alice is hosting say 100petabytes and has priced out all the individuals on the network, she will be subject to data policy rules and will have to obey law enforcement to kick out users that are using her hard drives for illegal activity.
Alice is now AWS.
Are you suggesting that AWS is doing every possible thing that would increase efficiency, or in other words, that AWS' services are optimally efficient? That seems very unlikely.
If FileCoin actually worked as a better technology, Google or some big name would say they will start using it to beat AWS in the data storage game.
I'm not saying FileCoin is useless, it is useful to store 3D printed gun designs and illegal media, because it is outside the reach of the law. But to think it is a more efficient technical solution to data storage is a bit naive...
Not that I'm saying Filecoin or a decentralised storage coin is in any way the way forward or even a competitor.. but saying that is how capitalism works is such rubbish.
In an ideal system you are right, but how often have you worked for a company that is raking in the cash by marketing well and being the "established" hand in the market?
There are so many companies kicking around that are the biggest player in their field that continue to make bank because accountants and upper management decide they are the proven, safe pick in the market and have signed up for long, entrenched software packages.
Also, if this upended the entrenched business model companies were using, they would not be able to quickly switch to it.
There are plenty of ways this could be the a better solution and AWS wouldn't switch to it, saying capitalism proves this is not a good argument.
It is precisely because these companies are not entrenched that my argument holds. But also consider the technical downside to FileCoin (slow connections to unreliable laptops VS fast connections to redundant data centers) and the economic downsides to FileCoin (it won't be worth hosting due to economies of scale).
I don't know how big the Filecoin network is, but it's very telling that they're not sharing this information. I expect that the amount of data stored is not large. Siacoin makes this available but total contracts are currently a bit under 200TB.
For those of you keeping track at home, that's a single medium-size storage appliance.
For example, see https://partysha.re
Of course. That's why when someone proposed a potential mechanism for increasing efficiency, it doesn't make sense to say "that couldn't possibly work, because if it could work then AWS would already be doing it."
2. Filecoin enables a transparent market for storage with bids and asks for storage, so you'll know what you're paying compared to AWS, Azure, etc. and you can decide what to do.
3. I'm on Comcast; if some guy two towns over from me is running his Filecoin-enabled IPFS node also on Comcast, there's going to be less latency than hitting an AWS or Google data center thousands of miles away.
4. My DigitalOcean droplet is ~200 miles away from me; trace route says it's 10 hops away. I can see the network name for Comcast two towns over from that's just 4 hops away. The other guys may be big but they can't change the speed of light.
5. Because IPFS is content addressed, I don't need to know where the content I want is located; thanks to the distributed hash table, any node that has that data can respond to my request.
6. IPFS and therefore Filecoin still work even if you can't reach the internet backbone… you can still get stuff via your neighborhood mesh network.
However it neglects to include Backblaze's B2 storage which is only $5/TB/month - https://www.backblaze.com/b2/cloud-storage-pricing.html
That is maybe a not a big enough price saving to convince conservative corporations to switch. Imagine going to your accounts department to request they purchase a cryptocurrency so you can use it to pay for data storage.
Amazon S3 also offers "glacier" for longterm storage with few accesses, which is $4 per TB.
If you're just looking for storage, S3 is certainly not the cheapest provider. Glacier is a bit different but clearly just for archives.
As for bandwidth cost: I don't believe Sia can be successful and stay that cheap. Why would providers of bandwidth for Sia be able to offer it magnitudes cheaper than the biggest tech companies in the world? Answer: It's offered by a bunch of individuals with no caps on their data plans. If Sia takes off and lots of people start using terabytes of bandwidth, the ISPs will put an end to it.
It's best not to use Glacier as a price comparison because it's really misleading.
* From https://www.backblaze.com/blog/hard-drive-cost-per-gigabyte/ 1TB of storage cost around $25. * I'll assume 3 year life-span for a drive. * Add 50% costs for the infrastructure and electricity to support the storage. * Add another 17.6% due to redundancy (17 data shards + 3 parity) - https://www.backblaze.com/blog/vault-cloud-storage-architect...
Gives a minimum cost (before employee costs, marketing etc) of $1.22/TB/month.
If we do the same with Siacoin (which I believe stores 3 copies of your data) we get $3.12/TB/month.
What do you think they will answer?
Option A: cross-region replication, regular backups, and secondary service provider failover plan
Option B: decentralize everything on people's laptops
But hey, your argument is totally valid, which is why everyone is still programming in Cobalt!
Argument to the Future. You can't provide evidence of a future event and asking for it is a joke. Maybe Sia/File/Storj are the future maybe it's something else, but asking for "evidence" of a future event is a joke.
Don't have to be choosing either "Businesses To Big To Fail(tm)" or "Average Joe" when you can have "Many businesses"
(Bear in mind I know what they are - I want you to sell me on them).
Say... if you were running a torrent site? Or you just wanted redundancy for your cat videos.
The encryption doesn't bother me. The way I encrypt it ensures it won't be unencrypted without actually knowing the passcode. What bothers me is that when you finally need to access it, it might not be there.
2. Price. There are many computers with a lot of free unused storage space and unlimited internet plans. Basically storage and transfer is 0-cost for them. So they'll compete creating very affordable pricing. AWS can't compete with that, because they have real operational costs.
http://blog.dshr.org/2017/07/is-decentralized-storage-sustai...
At least there is a benefit to being close to the requestor in terms of bandwidth/latency so as long as the network rewards that storage should be more decentralized than mining.
On the other hand, if I wanted to create a dedicated machine just to farm filecoins, that _would_ cost me extra for hardware and electricity.
This is why decentralized storage has an advantage in this regard. Unlike mining, storing files incurs almost no extra costs for the average home user, but does incur costs for dedicated storage systems.
The cost can be also in concurrent access. Unless you apply a heavy QoS on your network, serving filecoin network may degrade your other tasks.
Now if you don't have caps and want to keep it that way, I wouldn't be excited about filecoin. Unless you're actually paying for guaranteed bandwidth and it's part of your contract, your ISP is overcommitting. The moment people actually start relying on that fact, the ISP can do one of: raise prices and buy more pipes, lower the speed, reintroduce caps. The first one is a limited resource that takes time to apply. The other two are quick and known solutions.
It sounds lose-lose-lose situation: keep wasting your space or saturate your bandwidth or waste power on keeping HDDs spinning all the time due to accessing them every few minutes. It may vary with your usage patterns (if these disks are already always spinning and you have REALLY good bandwidth or two separate lines) but still, seems barely worth the effort.
And why couldn't it be optimized in a purpose built machine? Cheap Tibetan hydro and free cooling (altitude), low power CPUs, fast internet, fast switches and routers, 10 gigabit lan, many HDDs and SSDs (dug up from the Western trash piles even, it's not like failing once a month is a big deal). Chinese already have ASIC for bitcoin and this seems easier to make than that.
[1] I have no idea if I'm generous or stingy on USA power prices and network speed.
It's worth noting though that Sia (another decentralized storage network) is estimating their current storage costs at $2/TB-Month. That's half the price of Amazon Glacier. Whether that's the result of users farming from their home PCs or giant Chinese storage farms though I have no clue.
With a decentralized trustless system you have to assume a higher failure rate for nodes than with a central trust based system.
If you assume a higher failure rate you have to replicate data more extensively to achieve reliability.
If you have to replicate data more extensively then decentralized will always be more expensive than centralized.
If nodes are not trusted then replication has to play a major roll.
It doesn't matter how you mix up the data. You still have to store it somewhere and if that somewhere is a random untrusted stranger you will need to copy your data to a lot of random untrusted strangers in order to have kind of reliability.
Another way of thinking about this is to realize that the centralization of all data storage in China has not happened already. Why not? A motivated company (one as big as Amazon or Google) could move their datacenters to China if price was the only consideration. But they haven't, because latency (among other things) matters.
I'm not sure if this rebuttal applies in the case of Filecoin, since it's not clear to me to what degree renters are able to choose the hosts that store their data.
It is also accompanied by an announcement of the Filecoin Token Sale: https://protocol.ai/blog/ann-filecoin-token-sale-and-new-pap... which will begin on July 27th.
Filecoin miners can do the same by (a) registering as service providers and (b) complying with blacklists and takedown notices as mandated by their legal jurisdiction. Note: Filecoin miners don't just store arbitrary files assigned by the network; they sign contracts with specific clients to store specific files. The only difference from Amazon is that the network itself enforces these contracts.
That would be a huge overhead to be replicated across participants, driving out smaller players, lowering the "distributed" aspect of the system, creating larger points of failure and increasing the price.
No, the difference is that Amazon's contacts can identify the buyer and they're actual, legal contacts. They also have terms and conditions. Miners have contacts in terms of what the network is doing, which have nothing to do with contracts as understood by law. These are two meanings which just happen to use the same word.
https://en.wikipedia.org/wiki/Section_230_of_the_Communicati...
Seems like important sections of this paper were not ready in time for them to perform their ICO.
> FileCoin is using proof of storage
Yeah, Proof of Replication and Proof of Spacetime, to be precise.
It is always trivial to simulate storing client data that you don't have. So you'd have to invent a system where simulating client data is equally expensive to storing real client data. But, simulating client data can always be done deterministically using a seed, which means you don't actually have to store the data.
So you have to prove that generating client data is equally expensive to storing it, if not more expensive. Given that legitimate client data can literally be all zeros, I don't see how you can achieve this without some overhead.
That overhead will make you less competitive than Sia, which doesn't have the overhead.
I haven't yet gotten to the point where I fully understand proof-of-spacetime, but that is hard when the paper is literally referencing unpublished work. But I'm skeptical they have solved the generation / simulation attack.
My main concern at this stage is the replication setup time, which needs to be computationally expensive for the strategy to prevent generation attacks. That can possibly be amortized over long-term storage, but it sounds like it might be a long amortization period.
Can somebody explain to me how in theory could this decentralized storage beat something like AWS in efficiency, performance and uptime? I just think big beefy servers in dedicated data centers with fast and robust uplinks/downlinks should be better than consumer hardware on unreliable slow networks with shaky uptime?
It's also more likely to survive a major war if it's redundant enough.
Is it better for serving hot assets, no. Is it better for storing a massive amount of raw data that you need occasional random, indexed access to? Maybe.
Edit:
Also, less regulatory concerns since in theory nobody knows who is paying for it. This is probably it's biggest weakness, to be honest, since it could be used to spread child porn.
If there is a major war, international Internet will be one of the first things to go. In that scenario, having your files spread out across the world is a major negative. If you want the data to survive a major war, bury an external hard drive.
Most nation states would probably block outside internet access to avoid propaganda during time of war. I think I would have bigger concerns than whether I could recover my files stored online.
For cases like this I'd much rather trust just a simple USB stick or external drive which I could carry with me in my backpack and backup important files there so I don't have to rely on internet.
Second, hardware-level reliability doesn't actually matter for any distributed storage, reliability is achieved on the higher level with software. It's all about software. Decentralized global internet network is also better than any centralized equivalent, it has orders of magnitude more capacity and resilience. Ever noticed how bittorrent can saturate any link and how even google global caches sometimes have problems serving youtube videos for like months?
Other than that I don't believe decentralized laptop/desktop storage could even exist, for the same reasons why people don't mine cryptocurrencies on excess capacity on their laptops/desktops in the background. It can only be attractive to run it on dedicated storage nodes working 24/7 with good bandwidth.
You could build a Dropbox clone where users have the option of contributing storage to get rewarded with real currency. Why would anyone want to use the decentralized version instead, both as a user and as a contributor of storage?
The decentralized version is going to be much harder to get right for the developers, and why should the users believe that the technology is trustworthy? I just don't see the advantage, but I would be truly interested in hearing what I may be missing.
They do value simplicity of use above all and they do trust large companies (for better or worse).
What makes this system better than the other systems that fulfill essentially the same purpose? What does "getting it right" mean? Are they getting it right and how so?
I'm not saying it's not viable at all, but the distributed aspect alone requires a certain amount of popularity to sustain it.
Welcome to branding / advertising / marketing.
Does the blockchain have some way to prevent miners from double-spending their storage space? Is payment only collected upon retrieval, allowing Filecoin to be abused for storing backups that are rarely retrieved?
From the perspective of the network, there is no difference. When you agree to dedicate some storage to the network, you post a collateral in FIlecoin. If you lose something you've agreed to store (fail to prove that you're storing it when asked to do so by the network), you pay a penalty out of your collateral.
> Is payment only collected upon retrieval
No. Storage miners continuously prove (probabilistically) to the network that they're storing the files they've agreed to store.
Maidsafe made some noise in early 2014 but I haven't heard much from them since.
1. http://www.investopedia.com/terms/a/accreditedinvestor.asp
I do recommend anyone to also check out Swarm, similar but built entirely on top of Ethereum.
It seems to be just a document describing the protocol of Filecoin, which is fine. The problem is that all these "whitepapers" coming from these cryptocurrency communities seem to be promoted as scientific papers, but in reality do not stand a chance in any science/academic-related setting.
I'd wish they'd give more attention towards the science, if there's any. Otherwise, hey, sure it's just a PDF on some site - but don't call them "papers".
More than the academic presentation though, I'm just baffled by the lack of science that these projects with huge followings have.
Also, even if they don't pass an academia setting, what is your point? So many papers actually pass academia setting but does it do anything worthwhile? Filecoin is a project in existence for the past 3yrs working on top of IPFS, a popular file sharing protocol, and being built by a venture funded startup. Its founder has a vision of a decentralized future which makes intuitive sense.
don't call them papers
This is ridiculous and elitist. Papers outside of peer reviewed academia can and do exist. If anything, academics ought to qualify what they mean when they say "paper".
Also coming from academia, I'm surprised with the amount of crap that makes it through "peer review". And there's nothing wrong with concept papers as long as they don't purport to be more than that. Lastly, I don't think the review stage is ever truly double-blind. In a niche field, you almost always have some idea of who the paper's authors are.
In engineering (although I worked in neuroscience, I worked with many academic engineers and read way too many of their papers), whitepapers and concept papers are also often published - there's nothing to indicate here that this was a reviewed journal article. It's just a formatted PDF.
On the other hand, this is massively better to what normal corporate types call a "whitepaper" (usually just a glossy brochure that tells you nothing new.)
> Where are the experiments?
The experiments are in the markets. :) I.e. "here's a new coin, will it make money?"
That level of centralization is unacceptable.