You could even get all fancy by using openvswitch to configure your own virtual network topology. It just doesn't seem so complicated what they're trying to accomplish.
You could even get all fancy by using openvswitch to configure your own virtual network topology. It just doesn't seem so complicated what they're trying to accomplish.
Cheap hardware crashes. It crashes all the time. Furthermore, the application's needs themselves change all the time, depending on traffic curves and processing needs. The dynamic needs of Google's (and presumably Twitter's) completely heterogeneous application stacks don't lend themselves to simple virtualization and over-the-counter software. This is an incredibly tricky bin packing problem that was never quite solved in my 5 years at Google.
I don't know a lot about the "distributed frameworks" you mentioned, so perhaps they do this too. I kind of doubt it. If it were as easy as you think, I'm sure my friends at Twitter would be using what you mentioned.
Add in spinning jobs up and down quickly on demand rather than making them long-running. (This matters for a map-reduce, for instance.)
Add in making it easy to have dev/production running side by side with configurations that are as constant as possible, and not getting into each other's way.
Add in automatic failover from machine to machine, or from data center to data center as needed.
Add in making it easy to find/access applications set up and configured by someone else so that you can set rpcs their way.
After you keep adding things in, you understand why people think of this as a "data center operating system". And you realize that solving the basic problem - get jobs to run on a distributed set of machines - is only the visible first step in a long list of features that you want.
(Why would it be insane? Because with dedicated machines your utilization levels are going to be crap, reconfiguring and rebalancing the machine allocations will be painful, and you'd need massive over-provisioning to achieve a sufficient level of fault-tolerance.)
Another takeout from this, that virtualization is not Web-scale, it's more appropriate to IT and public cloud workloads, large companies like Google prefer bar-metal.