Under the hood: Facebook’s cold storage system
code.facebook.com
code.facebook.com
If Facebook (and I believe it's similar with other big players like Google and Amazon) are not using "enterprise grade" hardware (because it doesn't make economic sense at scale), who is? And more importantly, why? Why do "enterprise grade" products even exist if the largest and most deep-pocketed corporations don't find them to be good value propositions?
No "ordinary" enterprise (basically everyone else) can just say "oh, mounting and unmounting is an issue? let's use a raw disk instead!" -- that is just well beyond their capabilities.
Also, only Facebook / Google / Microsoft / Amazon have the scale that designing your own cold storage makes sense.
I guess that there are an awful lot of companies that, out of desire to save a few bucks, will go with home-NAS-types.
Now if you are Google or Facebook you want 100 Petabytes of storage but instead of paying Netapp $100 million you employee 20 really smart guys to build you something, you put it on cheap hardware and you build exactly what your software will talk to and what it needs.
It costs you $10 million up front but the hardware cost is just $15 million because you use cheaper drives and build (thousands) of your own servers.
Of course this doesn't mean that "Enterprise Storage" isn't going to be the next premium product hit by Open Source type solutions.
Google/Facebook/AWS (and maybe a Backblaze) have architected their data center setup to not require the functionality that Netapp provides. If you buy a NetAPP and plug it in your netork it is because you cant architect a datacenter from the ground up...
Facebook was a founding member of the Open Datacenter Alliance. They have a cool githup repo and put out interesting whitepapers like:
http://www.opendatacenteralliance.org/docs/architecting_clou...
It is not just cheaper but architected better. They would use Cisco and other enterprisey stuff if it were better but it is a case of "il y a moins bien mais c'est plus cher" as the French Linux motto goes...
Which is not to say NetAPP or Cisco is bad - just that you can engineer the datacenter to make it work better without enterprise functionality.
As an object lesson in failure, I had an experience with my first buildout of Hadoop back in 2011. Not so successful because I bought all the wrong components. Raid cards, 10G networking, HP and Bladenetworking rack switches, dual power supplies.
It was hard to unlearn how I was used to setting up servers.
I was a sysadmin and did the datacenter buildouts for the company. The task was to take the dev setup in Softlayer (on bare metal servers) and reproduce it in our racks.
The error was to take this somewhat literally. The bare metal servers that we leased from Softlayer were 2u boxes with redundant power supplies, Raid controllers and 10g networking.
I found a Supermicro configuration that almost identically matched and ordered 30 datanodes; 2 name nodes (this was back when Hadoop used the curiously named primary and secondary name nodes); and a server to launch jobs from.
We already had a Netezza in place and since that used Bladenetwork switches I decided to use the same as TOR switches with a beefier on as the Aggregation switch.
We decided to use Ubuntu as the OS and the Cloudera packages - but without the support or the console.
Every single one of those choices were mistakes. It is remarkable in hindsight that it worked at all.
The mistakes:
1) starting with the power supplies. Since they were redundant it dictated an A/B power setup in the racks. What this means is that you cannot use more than 50% of the power density because the entire rack is set to fail over. Each PDU has to be able ot keep the whole rack up so it alarms at 40% capacity. My using redundant powersupplies I was more than halving the amount of power i had at my disposal.
2) RAID. I hadnt yet read the excellent Hadoop Operations O'Reilly book which tells you why RAID is a bad idea for hdfs. Further buildouts included ripping out the RIAD card and JBODing the drives allowing Hadoop to properly use the raw disks.
3) 10G networking - because more is better right? Later after we hired talented networking folks I learned a bit about how important caching is in switches - particulary for Hadoop. My expensive monster Bladenetwork switches were falling over because the bursty traffic would saturate the switches. In addition I had the networking setup for a much more conventional network and we were doing the switch aware stuff in the config but had a wide open /16 as the network. I learned about properly segmenting to avoid uneccesary cross talk.
4) Ubuntu - never listen to developers. ;) Actually it was our default OS. CentOS works better for Hadoop I would later find out.
5) Lack of support. Cloudera is expensive but - unless you are already an expert - worth it. We should have gotten help early on.
6) Hardware choices. Supermicro was a bad choice. Penny wise / Pound foolish. They changed hardware config and it was hard to get replacemnts, etc. Cant say enough good things about working with PSSClabs for hardware though. I already mentioned the switches. I learned about Arista switches which are amazing for this application.
By the time I left that company, I learned an enormous amount from really smart engineers that knew gobs more about Hadoop and netwrking than I did.
I think the number of datanodes was up over 200 and we had a lean 1u datanode from PSSCLabs with single powersupplies 12 drives JBOD with onboard flash for the OS. Bonded 1G networking running up to Arista TORs and mutliple Arita Agg switches in a leaf / spine topology.
It was a thing of beauty.
But I think the real question is should you pay $100K to Netapp or $3K/month to Amazon and store your data on S3.
At least newer tech companies seem to strongly prefer the latter.
That's exactly what some companies are doing. There are some startup needs where you need at least an of magnitude reduction in OTS or cloud costs.
Enterprise hardware aims to be predictable: large businesses have to minimise risks so they value stability and clarity in performance, costs, and technical know-how. Enterprisey products are never cutting edge or best-in-breed or efficient, but they are predictable and that helps large businesses to have a reliable foundation for planning and executing IT projects over 3-5 years with known budgets, risks, support contracts etc. This predictability is important for businesses that are not based on IT (for example, making frozen foods or assembling tractors).
Facebook/Google/Amazon and other software-centered businesses depend on agility and finely-tuned machines - they need their "custom-spec" equivalent, and they have the skills required to manage it. But many other companies have known needs and are happy with standard enterprise products - they are the equivalent of an office worker that needs a machine to run MS Office, and a standard machine from Dell/IBM/HP is a more appropriate choice for them (rather than tinker with building their own).
Both consumer- and enterprise-grade hard drives use standard interfaces (SATA, etc) so all this "drivers" stuff you talked about doesn't apply to this particular hardware. Since even enterprise-grade hard drives will fail one day, one needs to have a proper backup strategy, which would backstop the needs of supporting of consumer-grade HDDs. So in light of this, why buy enterprise HDDs?
(I'm not saying one shouldn't, I'm just not buying the explanation above :])
Enterprise hard drives use SAS connectors to allow redundant paths and a full scsi command set. Enterprise hard drives have a bit error rate an entire order of magnitude better than consumer hard drives. The rated powered on hours and mean time between failure will be better, but not more cost effective with enterprise gear unless you take into account the labor cost and downtime cost of replacing that hardware.
It's not about the fact that the hardware looks and interfaces the same, it's that it lasts longer, performs reliably and predictably and comes with support. Even if you design for failure, it's much better to have 1 drive fail than 3 in a system.
The enterprise stuff will just be more durable and usually more performant in every way so it's worth the cost, especially when most companies are not tech giants with all the engineering talent to build something like cold storage from the ground up. FB/Google/Amazon/Microsoft/Apple are very specialized high-tech companies who know how to do this and not representative of most "enterprise" customers.
The short answer is "everyone in between"
I realized this before with regards to the ridiculously overpriced top-top-of-the-line CPUs. Let's say you're running an existing application at, say, a moderately small scale: 20 servers. Each server is only handling a couple dozen requests per minute, so the load is not very high, but the requests still take too long. You've hired a whole team of engineers to work on optimizing the code to make it run faster, and that costs you, say $10 million a year. You'd be happy to buy another hundred servers if that made it run any faster, but that's just not how it works -- the application is all single-threaded code, and maybe you can make that better, but that's what your $10 million engineering team is working on (maybe). So in the meantime, you spend whatever amount of money you can on getting the absolute fastest hardware you can find to throw in those servers -- fast CPUs, fast hard drives, fast network... and even if those servers cost $50k each, that's still just a fraction of what you're spending on the software engineering work.
I think they mean drives that are the quality of the WD black drives, it looks like them in the one picture. The types of drives that you might build a Hadoop cluster with knowing that you will be replacing 1-2 drives per 100 every month.
These drives (unless I misread) do not need to be hotswappable nor any other specialized setup. I would bet they are JBOD and directly attached with minimal controllers. If one drive per tray is powered on at a time then there is no RAID running across that tray...
As a point of comparison an expensive enterprise drive that you might see in a Netezza appliance costs many times more and has specialized firmware to work in a specific application, with certain RAID controllers at specific firmware setings. From experience I do not think that they are much more reliable - they are specific to the appliance, or in some cases like Aerospike required SSDs specific to a software requirement.
This storage design is stripped bare - the HDDs are off for most of their lifetime. I guess the biggest stress on them is spinning up and spinning down. They must focus on precise temperature/humidity control as it might be as much of a problem to have the drives too cold as too hot.
But more conventionally, there is a different thought process in enterprise IT and what I call industrial IT. Enterprise IT implements systems around minimizing failure events and create bespoke environments to meet solution requirements. Industrial orgs acknowledge that these events are a fact of life build process and software to deal with them. They also tend to limit the options available to consume IT.
Also, "at scale" cannot be achieved by an single entity outside of US Federal .gov. I work for a massive organization... 150k+ people. Facebook serves 10000x more users, and have the engineering resources to do amazing things as a result.
Also impressive how they shaved off all that extra power requirements and built-in expected unreliability. And really, utility is pretty darn reliable anyway, mixed with non-customer facing storage. Good move.
Then blow your mind even further by looking up the principles of the McEliece cryptosystem (but don't use it as originally specified - modern developments have weakened and refined it a lot).
I warn you: this is one of those down-the-rabbit-hole subjects that has a deliciously large amount of literature to soak up, you could find yourselves with bookshelves full of it. (An alarming number of patents, too, for the unwary.)
Although I don't split across too many failure domains, it still protects against isolated bitrot.
That's just fun.
In some way I'm amazed that nobody of the smart people at FB tough of the weight of this stuff. But then again, I probably made sillier mistakes myself on a much smaller scale.
My guess is that they were definitely aware of the weight tolerance of the floor and the tiles but looks like the weak spot was the canister wheels on the bottom of the rack. I wonder how they eventually got it out. They probably used one of those rack jacks to get a stronger dolly underneath... Or removed the HDDs and moved it.
This made me wonder about something else - the power density. They say the layout:
"can support up to one exabyte (1,000 PB) per data hall. Since storage density generally increases as technology advances, this was the baseline from which we started. In other words, there's plenty of room to grow."
I wonder how they account for increased energy consumption as they put in larger powersupplies to deal with the increase in power to run and cooling for the larger HDDs. I imagine in this layout everything is in proportion. CPU / RAM / Disk Space. It must be hard to accurately target growth with in the constraints of a limited power footprint.
The hard disk power management part is really fascinating in terms of keeping the storage cluster software alive and well to be able to respond to requests while keep the hard disks off. A low power SSD cache tier on the front may be part of the solution.
Another comparison with Google would be regarding the erasure coding technique described in this blog post. Google offhandedly mentioned using it in a 2010 presentation. Is this really Facebook's first deployment of erasure coding?
When will SSD drives catch up in storage capacity and price to traditional spinning platters? A SSD consumes very little power so this type of setup will not be as necessary.
Also what about the life cycle of drives that get powered on and off repeatedly? On one hand you have backblaze running consumer off the shelf drives 24/7. On the other you have FB/Google/Amazon powering on and off these drives.
Never. Someone did calculate that if you can get flash that lasts 15 years it would be cheaper than disk, though (because the disk has to be replaced). I can't find the links on those.
Also what about the life cycle of drives that get powered on and off repeatedly?
We wrote a paper on that topic where we treated start-stop cycles as a resource to be rationed with a token bucket filter: http://storage-conference.org/2011/Papers/Research/10.Felter...
Pretty likely. One of the post's authors, Kestutis Patiejunas, was an architect working on Glacier a few years ago. (src: https://www.linkedin.com/in/kestutisp )
The last few years have been weird - hell, ram is literally more expensive, per gigabyte, than it was in 2012. PC software requirements seems to be standing still. A five-year-old PC is still perfectly usable.
This wasn't true at all for most of my life.
I mean, yes, if storage space requirements don't start growing, of course, you are right, because hard drives, while they are big, are slow.
But... hopefully, this is a temporary setback. Someone will figure out how to make use of the surplus transistors in desktop PCs.
I mean, it can't be that hard; from what I saw in the '90s and early aughts, Microsoft seemed to release a new version of word every few years, and it required a new PC, even though I personally couldn't see how it was better.
But... apparently that isn't happening anymore.
My point here is just that if the need for disk space grows as fast as hard drive and ssd size per dollar grow, SSD might never become completely dominant, assuming that hard drives maintain their space per dollar advantage.
Of course, if things keep going where they have been going for the last few years, where system requirements for desktops don't really increase over time, then of course you are right, there would be no reason to have spinning disk.
You can already see this, sort of, in mobile, where storage expectations are dramatically lower, but spinning disk never had a real toehold there. Do you remember those tiny spinning hard drives that were packaged in CF cards? oh man, so cool! and so fragile.
Unless SSD will become better than HDD by all parameters, HDD's are very suitable as part of disk setup for anyone who need more than 200 GB data. And there are millions of people who don't use all that cloud things and prefer to download, store and watch/listen their content locally.
If my home media library gets wiped out by hardware failure, replacing it would be an annoyance but not a disaster. Anything that would be a disaster is covered by at least two cloud services.
I have family members who have had and continue to have user induced data loss due to single copy, but even single copy + cloud because they insist on passwords being b.s. so they don't write them down, don't remember them (due to arcane rules requiring them to pick a password they can't remember without writing down), and then the device has some problem that requires a reset, and then they can't get into their online account because they failed to properly set (and apparently weren't required) the recovery account or phrases.
""" To tackle this, we built a background “anti-entropy” process that detects data aberrations by periodically scanning all data on all the drives and reporting any detected corruptions. Given the inexpensive drives we would be using, we calculated that we should complete a full scan of all drives every 30 days or so to ensure we would be able to re-create any lost data successfully.
Once an error was found and reported, another process would take over to read enough data to reconstruct the missing pieces and write them to new drives elsewhere. This separates the detection and root-cause analysis of the failure from reconstructing and protecting the data at hand. As a result of doing repairs in this distributed fashion, we were able to reduce reconstruction from hours to minutes. """
"the amount of engineering at this scale is just insane. kudos to facebook engineering." - emocakes
Or do they specialize these data centers for ONE specific feature? I wouldn't want to be in the shoes of the people designing the full integration of the system in the latter scenario, imagine your use-case changes... that's a whole new level of refactoring, especially once you go into special-purpose hardware.
The slides from Vault 2015 talk about Sheepdog on low power hardware and SMR drives at Alibaba.
I'm a consumate hoarder, I've got 2 decades of emails saved (but not all emails!) ... even I think perhaps we've gone over the top.
Perhaps it's good for social history. Or maybe in the future for a Black Mirror like reconstruction of people's personalities in an AI.
Doesn't there need to be limits on what we keep?
[Meta: the parents is a valid remark and one that I feel adds to the complexion of the scenario under consideration; perhaps it could have been made better, fleshed out, but still. Wish HN - as a population - would value those that add to the conversation in this way.]
Not invented here..
I don't use Facebook at all, but I appreciate the many investments they've made in open hardware and open source software. They don't give away everything. But they definitely contribute in significant ways.