Distributed is not necessarily more scalable than centralized
muratbuffalo.blogspot.com
muratbuffalo.blogspot.com
Riak is a similar type system.
Dropbox is "centralized" in the sense that it is one service, but it's not the opposite of distributed which would mean "running all on one computer."
Edit: I said "hash of the data's key" but really it's a hash of the key plus the bucket.
So it seems, distributed systems are not defined in terms of spatial or logical distribution of "stuff being done" (because everything is distributed in this sense), but rather by assumption of (un)reliability of the links between the components, and the choice of components. And this may depend on your vantage point, too.
So if that reliability is good enough, you don't need theory of distributed systems, and it all makes the system more efficient. In that sense, Dropbox is more efficient than P2P solution because you have hidden assumption (from the user perspective) that Dropbox servers will always be available.
Edit: Also reliability of components plays a role. Thinking about it more, it really seems to be a question of hierarchy. Distributed systems are less hierarchical than the centralized counterparts; there are less bottlenecks (that may fail) but more coordination required. So the OP is basically arguing that some hierarchy scales better than no hierarchy.
The student is correct. Lets ignore the fact that Dropbox is actually distributed and say it is centralized because all nodes of the system belong to one provider. The only way Dropbox could have scaled to 200m users was tons of cash. In a distributed solution where each node is a provider themselves, each additional user could potentially increase the performance of the system. The distributed alternative scales much more gracefully without running into the bottleneck of needing more cash to buy more machines/storage/bandwidth. In this particular frame, distributed is most definitely always more scalable than centralized unless you have unlimited cash.
Wouldn't each additional user have to potentially buy more hardware to increase the performance/capacity? There's still a cost, there's still a need to buy more capacity to increase capacity -- you've just "distributed" the cost to the (organizational) nodes. Which, okay, can be useful sometimes, but clearly there's a huge market of people who would rather pay someone else to take care of it, than spend that same money (or likely more) on being a "provider themslves".
The type of distributed storage system that I'm speaking of relies on each node/user being a provider themselves.
For an example checkout http://storj.io/
There's a point that you might be missing. Similar to how you have people purchasing and maintaining powerful machines to mine bitcoin, you also have people using their spare HDD space or potentially buying dedicated storage to support the "distributed dropbox" system. The reason is that there is an incentive for doing so. You get paid for the storage/resources you provide. This adds to its scalability because any addition to the system is not only paid for but profitable.
2) This article doesn't actually make any argument about why a centralized system can scale as well as a distributed one.
The point is to compare a more centralized architecture to a more/fully decentralized architecture.
Ok am I missing anything. So we are employing Paxos to replicate the centralized server. Are we replicating it to itself? Because if we are not, we got ourselves a "distributed" system.
I suspect the student thinks that distributing his/her files among his/her friends and/or multiple services (bittorrent-style) will allow his/her to increase throughput -- however, I suspect it will merely increase complexity (and possibly also cost) without actually making syncing/back-up faster.
This is really an industry-wide problem begging for a neat solution. Software eats middle management! (Devops => Devmangops? Mmm... mangoes...) Perhaps the world needs an open source tool in the organizational management/risk space that models business-level risk based upon commercial as well as technical infrastructure.
Perhaps the best model for developing such a capacity is a generic exchange protocol with plugins for risk management? My start brainstorming @ http://www.ifex-project.org/our-proposals/ifex
Distributed tends to produce higher availability than centralized systems and often that is worth the cost.
http://www.allthingsdistributed.com/2007/10/amazons_dynamo.h...