Jupiter Rising: A Decade of Google’s Datacenter Network [pdf]
conferences.sigcomm.org
conferences.sigcomm.org
The secrecy was at first a universally-agreed necessity; in simplistic terms, we didn't want MSFT knowing how much money to throw at the problem, or even what the problem was. This was true (and surely still true to an extent) for all of platforms: it was always amusing to see public photo shoots at "google datacenters", which in reality were little more than a stack of google search appliances at a corp location.
The level of detail here is great and really sums up 10 years of a lot of hard-earned findings. I'm thrilled pictures of the hardware are even included, that team is just top notch.
For anyone who isn't a Google or Facebook but has large or growing colo/dc network needs, check out Cumulus Networks. It applies a lot of the SDN ideas (think ssh/puppet-driven config mgmt for your switches as just the start) and topology possibilities seen here; doesn't hurt that JR did a brief stint on firehose :)
- nolan
It's like finally reading how my car is an "intern-al comb-usti-on eng-ine" whatever that means.
Having full cross-sectional bandwidth between any pair of hosts means the bin packing problem is a lot easier. You don't need to (say) make sure your map reduce job is scheduled with one shard per rack because racks only have so much bandwidth. Any host on any rack will do. You can forget racks even exist.
This makes overall utilization of clusters more efficient (tighter bin packing), and the corollary is you don't need as many clusters and machines. (Not that has ever stopped Google from building more :)
This paper is weird--like maybe they switched around the names of some things and didn't mention others.
Just an FYI.