A Rare Tour of Microsoft’s Hyperscale Datacenters
nextplatform.com
nextplatform.com
This was one of the big secrets that Google learned early on. Every bit of air you cool that isn't going into a computer is wasted. The issue is that co-location facilities have to be ready for any kind of user equipment, but if you own the entire data center and all the equipment inside you can design it differently and much more efficiently. It stops being computers in a building and starts being a building sized computer.
Datacenters are sensitive facilities. It's not surprising that they block entrance to many people.
From what I've seen in Europe, I can remember a couple of places in military, datacenters and national research centers where it's clearly written "European Only" on the jobs.
"A Green Card (Permanent Resident Card): Gives you official immigration status in the United States. Entitles you to certain rights and responsibilities. Is required if you wish to naturalize as a U.S. Citizen"
At the moment, a few providers have the scale and skills to run datacenters much more efficiently. But I'm guessing that within a few years there will be some generic datacenter-in-a-container available, with efficiency not much inferior to the big four.
At that point, we go back to the hosting market of 15 years ago. Everybody can offer a datacenter without deep technical knowledge, and sell compute cycles on an open cloud market.
It's like all the tiny hosting providers, except that it requires more capital. So it becomes financialised -- if you can get cheap power, low temperatures and good connectivity, and borrow a few million dollars cheaply, then you're in business. But margins collapse precisely because nobody can do it.
In the end, once we get over the transition period of this move to cloud everything, datacenters end up like utilities
A while? More like forever. Short of corporations abandoning the notion of shared physical workspaces (which isn't likely to happen in the foreseeable future), corporate MDF, IDF, and edge gear will always be around in some form or another.
No. It's just going to look different, with the most successful vendors selling flexible, nearly white-box hardware with good long-term maintenance terms.
> if you think a lot of infrastructure will be moving over to the cloud (because of cost pressures) over the next few years
There's certainly going to be a TON of infrastructure moving to 'the cloud' in the next 5-10 years, but cost isn't necessarily the strongest driving factor pushing cloud adoption. Big, established Enterprise-with-a-capital-E type businesses are often quite capable of continuing to run infrastructure in-house with better bang-for-buck in terms of raw capacity than what is presently capable with public/hosted private cloud solutions.
In the underlay it's all still very much basic, old fashioned networking using existing protocols: BGP in the DC at FB and MSFT, and BGP, ISIS, MPLS in the WANs. (AMZN likes to pretend they're special so I won't comment on them.)
Google is a bit different, in part I think this is due to their culture and early scale they were forced to drive a lot of developments wrt merchant silicon use, and this naturally led them to their own network OS solutions, with a simplified semi-centralized IGP solution for their DCs, and centralized-TE solution for their inter-DC WAN. It's not entirely clear to me if they would still feel the need to do this if they were starting again.
I took a class on design for low PUE implementation and some comments from government data center technicians who said, in order to comply with the federally mandated PUE requirements, that people were leaving on or turning on zombie boxes to up their IT load.
AFAIK, the DCOI sets a target PUE of 1.5 or less, so running unnecessary workloads to meet the target PUE doesn't make much sense. I would bet there was some other kind of tomfoolery going on (hiding the fact that the DC overspent on efficiency when building out/upgrading the facility, or something along those lines).
But the PUE measurement system doesn't know how the various severs are spending power, just how much they are spending. So a busy wait, computing pi, dynamic language hash table lookups, or doing useful work all look the same.
My OpEx with the new datacenters is that I have to change the filters, and that is really the only maintenance I have. And we have moved to a resiliency configuration where I put more servers in each box than I need and if one breaks, I just turn it off and wait for the next refresh cycle. The whole OpEx changes with the delivery model of the white box. So we learned quite a bit there, but now we have got to really scale.”
I've never heard of any serious non-research uses of it, and I've spent a good bit of time looking. Every time I've heard a rumor, it had turned out to be false.
I was wondering about that sentence, isn't humidifying air a bit problematic in a datacenter? Computers and water usually don't get along that well..
An infrequently-used app also looks better by this metric than a frequently used app. The service behind the mobile weather app you use might look more efficient than Facebook, just because you use it once a day instead if a dozen.
Disclosure: Microsoft employee, not involved in our data center designs.
(That's not entirely fair since locality also requires some overhead. One machine that supported a million users might be great if all of those people were in one city. But usually that's not the case, so at the very least you need a box in the top 100 cities (by whatever measure, the simplest being population), at least.)
It's an extremely bad, resource-intensive architecture ;)
The number of servers in private companies might be around 1 per citizen, and the number of processors per human around 100x (incl. mobile phone, tv, smart lamps). Given a proc has 5m transistors and humans have 100m neurons... our architecture is so bad that we're already outnumbered by machines by a factor of 20 at least.