For example, with a decentralized system files could be broken up into shards and distributed across multiple nodes, so unless a significant percentage of computers on the planet all melt down at once your data would remain accessible (i.e. zero downtime). That would also allow downloads to be parallelized across multiple connections like with BitTorrent, meaning bandwidth wouldn't be an issue either.
And that's not even considering the effect that a commoditized pricing system would have on the costs of a distributed storage solution. I imagine it'd be significantly cheaper than AWS too, as storage would be sold at a price very close to costs.
On the other hand, you could spend some money and guarantee it instead: https://aws.amazon.com/directconnect/
As for price, AWS has some of the cheapest operating costs due to scale, something which smaller players will have trouble competing with. Moving data within AWS is fast and sometimes free.
Smaller players also pay the highest cost for bandwidth. If some plans don't have bandwidth caps yet (many already do), they will get caps as soon as serving bandwidth as an individual becomes profitable.
Pretty much all distributed systems have the following in common: You pay for resilience with overhead. If it wasn't that way, everything would become as distributed as possible over time.
Isn't it the other way around? Small providers offer low traffic costs while the big cloud providers (Google, AWS, Azure) charge significantly for bandwidth.
You can setup any old distributed database to do this... If there was a benefit to breaking down a file rather than replicating it, AWS would be using this efficiency to provide a better service.
> the effect that a commoditized pricing system would have on the costs of a distributed storage solution
It'll always be cheaper for someone to sell their hard drive than to sell access to their hard drive, just like it is cheaper to buy bitcoin instead of mining it.
Here is why: let's say Alice is profit-driven and greedy and is making $1 profit for 1gb of space. She can use those profits to buy more hardware, make more profit, buy more hardware, and so on and so forth until she has 100,000gb of space being offered. At that point, because of economies of scale, it will cost Alice less to maintain 100,000gb then the total cost of 100,000 people hosting 1gb each (eg. reduce total electricity cost because of bulk purchase).
Now, Alice can grow her operation so big (say, 100 petabytes) that it actually becomes no longer profitable for Bob to host 1gb. Alice has so much lower costs per gb that she was able to drive the price per gb down to a point where the price is lower than Bob's costs.
This is basic economy of scale so far, right? And it's the same reason why data centers filled with ASICs make it useless for you to mine Bitcoin on your laptop.
OK now that Alice is hosting say 100petabytes and has priced out all the individuals on the network, she will be subject to data policy rules and will have to obey law enforcement to kick out users that are using her hard drives for illegal activity.
Alice is now AWS.
Are you suggesting that AWS is doing every possible thing that would increase efficiency, or in other words, that AWS' services are optimally efficient? That seems very unlikely.
If FileCoin actually worked as a better technology, Google or some big name would say they will start using it to beat AWS in the data storage game.
I'm not saying FileCoin is useless, it is useful to store 3D printed gun designs and illegal media, because it is outside the reach of the law. But to think it is a more efficient technical solution to data storage is a bit naive...
Not that I'm saying Filecoin or a decentralised storage coin is in any way the way forward or even a competitor.. but saying that is how capitalism works is such rubbish.
In an ideal system you are right, but how often have you worked for a company that is raking in the cash by marketing well and being the "established" hand in the market?
There are so many companies kicking around that are the biggest player in their field that continue to make bank because accountants and upper management decide they are the proven, safe pick in the market and have signed up for long, entrenched software packages.
Also, if this upended the entrenched business model companies were using, they would not be able to quickly switch to it.
There are plenty of ways this could be the a better solution and AWS wouldn't switch to it, saying capitalism proves this is not a good argument.
It is precisely because these companies are not entrenched that my argument holds. But also consider the technical downside to FileCoin (slow connections to unreliable laptops VS fast connections to redundant data centers) and the economic downsides to FileCoin (it won't be worth hosting due to economies of scale).
I don't know how big the Filecoin network is, but it's very telling that they're not sharing this information. I expect that the amount of data stored is not large. Siacoin makes this available but total contracts are currently a bit under 200TB.
For those of you keeping track at home, that's a single medium-size storage appliance.
For example, see https://partysha.re
Of course. That's why when someone proposed a potential mechanism for increasing efficiency, it doesn't make sense to say "that couldn't possibly work, because if it could work then AWS would already be doing it."
2. Filecoin enables a transparent market for storage with bids and asks for storage, so you'll know what you're paying compared to AWS, Azure, etc. and you can decide what to do.
3. I'm on Comcast; if some guy two towns over from me is running his Filecoin-enabled IPFS node also on Comcast, there's going to be less latency than hitting an AWS or Google data center thousands of miles away.
4. My DigitalOcean droplet is ~200 miles away from me; trace route says it's 10 hops away. I can see the network name for Comcast two towns over from that's just 4 hops away. The other guys may be big but they can't change the speed of light.
5. Because IPFS is content addressed, I don't need to know where the content I want is located; thanks to the distributed hash table, any node that has that data can respond to my request.
6. IPFS and therefore Filecoin still work even if you can't reach the internet backbone… you can still get stuff via your neighborhood mesh network.
However it neglects to include Backblaze's B2 storage which is only $5/TB/month - https://www.backblaze.com/b2/cloud-storage-pricing.html
That is maybe a not a big enough price saving to convince conservative corporations to switch. Imagine going to your accounts department to request they purchase a cryptocurrency so you can use it to pay for data storage.
Amazon S3 also offers "glacier" for longterm storage with few accesses, which is $4 per TB.
If you're just looking for storage, S3 is certainly not the cheapest provider. Glacier is a bit different but clearly just for archives.
As for bandwidth cost: I don't believe Sia can be successful and stay that cheap. Why would providers of bandwidth for Sia be able to offer it magnitudes cheaper than the biggest tech companies in the world? Answer: It's offered by a bunch of individuals with no caps on their data plans. If Sia takes off and lots of people start using terabytes of bandwidth, the ISPs will put an end to it.
It's best not to use Glacier as a price comparison because it's really misleading.
* From https://www.backblaze.com/blog/hard-drive-cost-per-gigabyte/ 1TB of storage cost around $25. * I'll assume 3 year life-span for a drive. * Add 50% costs for the infrastructure and electricity to support the storage. * Add another 17.6% due to redundancy (17 data shards + 3 parity) - https://www.backblaze.com/blog/vault-cloud-storage-architect...
Gives a minimum cost (before employee costs, marketing etc) of $1.22/TB/month.
If we do the same with Siacoin (which I believe stores 3 copies of your data) we get $3.12/TB/month.
What do you think they will answer?
Option A: cross-region replication, regular backups, and secondary service provider failover plan
Option B: decentralize everything on people's laptops
But hey, your argument is totally valid, which is why everyone is still programming in Cobalt!
Argument to the Future. You can't provide evidence of a future event and asking for it is a joke. Maybe Sia/File/Storj are the future maybe it's something else, but asking for "evidence" of a future event is a joke.
Don't have to be choosing either "Businesses To Big To Fail(tm)" or "Average Joe" when you can have "Many businesses"
(Bear in mind I know what they are - I want you to sell me on them).
There are other schemes that don't require a copy of each shard, allowing you to reconstruct the file if any subset of the nodes is online.
Any subset? Surely there's some minimum number of nodes (>1) that need to be online. Unless each node has a complete copy.
I think Sia's method would prove to be extremely reliable, but again, it depends on the hosters. If the largest single hoster owns less than a third of the capacity of the network, even if they suddenly go offline, you should be able to recover your file since you rely only on a third of the 30 hosts who hold your file to be online.
Also, I think that the number of pieces you need to recover a file can be tuned, so if 33% would prove to not be reliable enough, Sia/Filecoin can tune it down to increase reliability.
Filecoin (and other decentralized storage systems) provide monetary rewards and penalties for storing or failing to store files. This is introduced in Section 4.1.1 of this paper.
The bad thing is that you would not be able to hold them responsible.
And on private torrent sites, where there is an incentive (usually non-monetary) to seed torrents, torrents really don't die all that often.
To go even further, if you orchestrate several miners carefully (maybe computationally) they can together perhaps stop people from accessing certain files.
It'd blame Apple & Google for not being p2p seeds.
2. Price. There are many computers with a lot of free unused storage space and unlimited internet plans. Basically storage and transfer is 0-cost for them. So they'll compete creating very affordable pricing. AWS can't compete with that, because they have real operational costs.
Say... if you were running a torrent site? Or you just wanted redundancy for your cat videos.
The encryption doesn't bother me. The way I encrypt it ensures it won't be unencrypted without actually knowing the passcode. What bothers me is that when you finally need to access it, it might not be there.