This is the start of a scaling path that winds down Distributed Systems Avenue, and eventually leads to a place called Hadoop.
(Replication and consensus are remarkably difficult problems that Hadoop solves).
This is the start of a scaling path that winds down Distributed Systems Avenue, and eventually leads to a place called Hadoop.
(Replication and consensus are remarkably difficult problems that Hadoop solves).
Certainly, and unfortunately, the exact point at which Hadoop becomes the better option over big iron is generally an ongoing debate and shifting target. But there's no doubt that such a point actually exists.
But if it does, then it's a pretty big chunk of data, and a very fast network.
As such, it is actually "Bane's Rule" which states, "you don't understand a distributed computing problem until you can get it to fit on a single machine first."
(Thanks to nekopa, who also referenced it further down in this thread.)