525 karma · joined September 13, 2008
Are you an engineer? I ask because I find it very hard to believe an engineer would write that.
What happens between the time the money is debited from one account and it shows up in the other? Hint: it's not a transaction in the ACID sense of the term.
I would address the rest of what you say, but I'm afraid I cannot help you on your reading comprehension. I encourage you to read the article. All the words please.
As for credit and debit charges, my bank always has two transaction histories. Cleared transactions and pending transactions. Pending transactions are merely funds that have been reserved against your account and they may or may not clear after a period of time. Sounds like that fits pretty well with eventual consistency.
Once one scales beyond a single node the system becomes distributed. Then by definition one must deal with distributed systems problems in order to achieve scale beyond the capabilities of a single node.
> When you decide to give up consistency, your application can no longer assume consistency ever. Giving up availability in case of network partition means few extra minutes of downtime a year.
Depends on the application. Giving up availability might mean cascading failures throughout your entire application. For instance if the datastore is unavailable for writes then any kind of queueing systems built around the DB (a common design pattern) run the risk of overflow during the downtime.
And I would make the argument that once an application scales beyond a single datacenter it cannot help but give up strict consistency under error conditions.
> I don't think he completely misunderstood distributed systems - I think he decided to completely side-step the entire field.
If he didn't misunderstand them then he is purposefully ignoring the hard problems. Which is worse?
8=============D~~~
When dealing with remote processes a monitor is very quick to detect a process crash, however it can take a lot longer than 5 seconds to detect a node crash or other form of netsplit. Setting an application specific timeout in this case can be a prudent strategy for setting up a firebreak so that a netsplit does not cause cascading failures through the whole distributed system.
A case can also be made that just using monitors is not sufficient when talking with local processes. Monitors do not detect those situations when a process becomes unresponsive due to overloads like queue overrun. Once again it is a good idea to have a timeout in the call so that one may avoid a single misbehaving process crashing the whole system.
In fact, I find it difficult to imagine a situation where one would not want to have a timeout on a gen:call. That is why timeouts are set by default in the gen call facilities. Building highly concurrent systems is difficult. Erlang makes a lot of the typical pain points in building these systems go away. But the need to tune defaults and stress test a system remains.