Storage Pod 6.0: Building a 60 Drive 480TB Storage Server
backblaze.com
backblaze.com
I have a question after reading the Vault overview here:
https://www.backblaze.com/blog/vault-cloud-storage-architect...
I'm curious what the bandwidth demand is when an entire host fails and you have to drop a new replacement host in, since you would have 60 Tomes in the Vault that need to be rebuilt at once. You have at least two other parity copies when a host dies (assuming no other drives in the affected host's Tomes are dead on other hosts in the Vault set), so I'm guessing you can afford to wait the handful of days it would take to rebuild the host. I'm still curious to know what rebuild speed you plan for, since I'm guessing you'd be looking at 40Gbps NICs eventually. I was surprised to see only 2x10Gbps.
It's MUCH better to lose only one drive, because we distribute the rebuild CPU task across all 20 CPUs in the tome and we can get it synced back up in a few hours. All the chattering back and forth still doesn't come anywhere close to maxing out the 10 Gbit.
For uploads, each individual pod has a 10 Gbit/sec connection, so ACTUALLY the vault has an aggregate of 200 Gbit/sec facing towards the internet. However, every byte that a pod accepts it has to ALSO retransmit that internally so that cuts the actual theoretical upload rate to 100 Gbit/sec for a properly parallelized application.
Spelling error: check for 'utlizing' -> 'utilizing'.
Dang it! We will fix, thanks.
Wow. This is pretty shocking TBH. My basic ceph nodes will crush a 40gbps link doing a rebuild/rebalance. What's your bottleneck?
Bluestore and rocks DB really changes the back-end story altogether. No more POSIX filesystem junk with journals and such.
For an example, take a look at Mellanox's cable page which lists a large number of combinations of connectors and cable types.
So in general, to call a random sfp cable/port an ethernet port would be confusing unless both parties already knew it was passing ethernet traffic.
The actual connector and cable are mostly independent of each other. The most common connector for an ethernet cable is RJ-45 (what you're referring to as "ethernet ports"). But it's just as possible to have SFP, coaxial, or USB even.
First I think it is awesome that you guys come on HN to respond when backblaze makes the news. Thank you.
If you bought a prebuilt storage pod (from somewhere like http://www.backuppods.com/) could you start out at less than capacity and add more as you go? Say you started with one drive -- would it need to go in a specific location? When you add more would there be a significant reconfiguration process? Is it possible to hot-expand? I guess location and config has a lot to do with software involved. Has there been a blog on the software side of managing that much storage? Can it be as simple as a spanning partition to treat all drives as a single drive?
Also, are you tracking the cost of SSD drives? Do you expect cost parity soon? It seems that capacity parity is getting close, but cost parity (especially at that capacity) is farther off. Do you have an estimate of when it might be or is it "far off, check again in 5 years". I'm sure you all found the ssd endurance test at techreport interesting (http://techreport.com/review/27909/the-ssd-endurance-experim...).
We ARE tracking SSD costs but thus far it's not quite at the point where the Cost per GB is in our favor. We'd love to play with them once it gets close, it would neat to build an SSD pod and see how it stacks up!
> are you tracking the cost of SSD drives? Do you expect cost parity soon?
Yes we track SSDs, and this is a hotly contested topic inside Backblaze. Personally I'm predicting they "cross over" in cost effectiveness within 2 years. But I don't have any more insight than you do, and many inside Backblaze disagree with me.
When I say "cost effective" that is total cost of ownership, which includes paying for less electricity to power the SSD, possibly the higher density of SSDs, and whatever failure rates of the SSDs forcing us to purchase more (might be better or worse than hard drives). So whatever our little spread sheet kicks out as cheaper -> that's the one we will purchase!
One random nice thing about SSDs is there will be less case vibration, and the drives shouldn't be affected AT ALL by the small amount of fan vibration.
It seems like the workload of Backblaze (like many "big data" workloads) is such that the extra IOPS don't actually help.
With thousands of drive spindles, and a very small portion of the data on on the spindles being needed at any given moment, I don't imagine you are actually limited by the performance of spinning rust.
No way! We are disk I/O limited right now. We think we can approximately double our performance simply by using SSDs instead of traditional 7200 RPM drives. I'm advocating for building an "SSD Pod" built with 1 or 2 TByte SSDs.
> and a very small portion of the data on on the spindles being needed at any given moment
With our original Personal Backup product that was true. Everybody was happy with waiting a few hours for their restores to complete. But with our new product line of B2 (competes with Amazon S3) then suddenly performance can matter much more, because programmers using B2 might be implementing all sorts of different access patterns. If those programmers are using B2 to implement online backup, then yes, still fine. But if the programmers are building hacker news or reddit or Dropbox the access patterns might require faster IOPS to be more responsive.
It's surprising to hear this. I would have thought that your backup workloads would leave enough unused IOPS that you wouldn't be anywhere close to hitting a performance ceiling.
What metrics do you optimize towards? Tail latency?
Sure, but there are tradeoffs. One option is to just concatenate all the drives together, either with mdadm (grouping the drives into a RAID-0 array) or with LVM (adding them all to a single volume group). This lets you add drives very easily and efficiently, but has no redundancy. When (not if) one drive fails, your entire FS is toast.
If you care about your data, then you'll want to use RAID-6 for redundancy. But adding a drive to a RAID-6 array is very time-consuming, because all of your data has to be re-shuffled onto a new drive. And you don't want to do this without a backup -- although the reshape process is designed to be safely interruptible, it's not bug-free. (I speak from experience, having nearly lost a few TB this way not too long ago.)
You can compromise by grouping your drives into multiple RAID-6 arrays. For example, instead of using 58 data drives + 2 parity drives, you could use 6 arrays of (8 data + 2 parity). This gives you much less usable space (80% as opposed to 97%) but it allows you to add 10 drives at a time without having to re-shuffle all your data. And it means you can recover from a drive failure much more quickly.
Of course, you can also just treat the drives as separate block devices, and provide a single redundant filesystem at a higher level, using something like Ceph. I would assume Backblaze is doing something along these lines.
They are using a regular filesystem on each drive and implement redundancy on top of regular files. It's more like Swift [1] than Ceph. Which is something parent should consider too, if he's going to deal with that much data.
[1] http://docs.openstack.org/developer/swift/overview_architect...
But seriously I do hate namespace collisions like this.
Swift (the storage architecture) is something i've been a fan of for years, swift (the programming language) i've dabbled in a bit and its interesting for some use cases for sure (aside from iOS programming which I'm not interested in at all). The cognitive dissonance of switching back-and-forth is annoying. I do hope nobody ever writes or needs to use a swift API connector in swift, that would be painful.
You put 10 of those in a rack and now we're up to 1000 pounds. At what point does this become challenging for the floor?
My favorite example is science labs. We've heard from a lot of them that they used to run experiments, collect the relevant data, and toss the rest. It was less expensive to run the experiment again than storing all of the data. That's changing now with storage being less expensive, and they're able to keep the entire experiment data set intact, and they love that because it means in the future they can take a look at all of the information surrounding their findings, and if the technology or science changes, they'll have the entire data set to play with instead of just a subset. Very cool stuff!
If you build your own datacenter, you are of course responsible for the floor loading...
I've pretty much reconciled I'm going have to buy something from http://serverlift.com at some point if I keep doing this job and stuff keeps getting heavier.
Disclosure: I work on Google Cloud (but not Nearline).
What do they have 300TB of, if it's not a rude question... ;-)
Tell your friends? :)
http://www.storagereview.com/what_is_shingled_magnetic_recor...
A write to one area can be done in a single disc rotation, as long as it's sequential.
SMR disks added what are called ZBC, Zoned Block Commands to the SCSI spec. The disk itself exposes the ability to interact with zones several ways:
- REPORT ZONES - What zones there are, how many, what type, and state of the zone.
- OPEN / CLOSE / FINISH ZONE - zone write state management
So essentially, it's like a drive exposes ~8 files, opened with O_APPEND, and you have to manage how you stripe your writes between those files to handle your workload. Your disk can also have zones that aren't shingled, but those seem to have been fairly unpopular.
I wonder what limitations the 6-month production run requirement put on design - it would seem you would have to be a lot more conservative in your "definition of done" (a bad decision can be expensive when you buy a half-year worht of gear).
If Apple could afford to give free OSX update, free iWork and ilife, Why not free iOS Backup?
Or at least offer a Time Capsule that i can backup my iOS devices without going to the cloud.
No 25/40gig networking?
It's nice to point out the depth issue, as older racks are shallower than the newer versions. This blocks things like the 0U PDU channels and cable management channels. That said, the newer, deeper (and more popular) racks fit everything just fine.
How tall are your racks? How many pods are you getting in there? 3ph 50A power?
One step at a time. :-) We think we can put a few more drives in the motherboard half of the POD without many other changes, and connect them to already existing SATA connectors on the motherboard. To get all the way to 75 drives we will need another row of port multipliers, which then means we need even more SATA cards, etc.
> No 25/40gig networking?
We're currently not able to saturate the 10 Gbit so for us it is wasted money to go faster. Remember that we run these in "vaults of 20" so we have 20 pods EACH with 10 Gbits so the vault hoovers in data at 200 Gbits/sec if you can thread the application.
> How tall are your racks?
Most of our racks are 45U tall where the PDU takes 4U, so we can fit (not necessarily power) 10 pods per rack and still fit a network switch in the 1U left at the top. In the past (with 45 drive pods) we ran 2 circuits of 30A 208V power (three phase). The two circuits are redundant - our power strips can flip over from one to the other if there is a loss of power on the currently used power. But I think they only put 8 or 9 of the 45 drive pods in a cabinet because it slightly exceeds what the datacenter likes. And now that we're moving up to all 60 drive pods, we have to rejigger the whole thing.
It is also worth mentioning that when we bring a 20 pod vault online, we put each of the 20 pods in different racks in different parts of the datacenter. This helps keep the vault completely online if any one component fails. For example, if the power strip feeding the pods in 1 rack power dies, it only takes 1 pod offline out of several different vaults, so ALL the vaults stay online both reading and writing data. (I hope that made sense.)
As to spreading the vaults around the DC, it sounds like that's more to prepare for ToR switch failure -- its sounds like an ATS handles your PDU and power circuit failover.
Yeah it sounds like you're mostly CPU bound calculating EC if you're going to rebalance and replace failed pods so 25 gbps ethernet probably isn't interesting (yet). That said, it's only marginally more expensive than 10gbps twinax. Interesting that you're running 10gig baseT. Interesting from a power and latency perspective, though I'd imagine you're not worried too much about latency.
$ exiftool Build+Book+Backblaze+60+Drive.pdf | tail -n5 | head -n3
Title : 7400-0283 Rev A OMS Assy Backblaze 60 Drive
Producer : Mac OS X 10.11.3 Quartz PDFContext
Creator : PowerPointBackblaze is using 4TB drives, which is the current $/TB leader.
Interesting decision by BackBlaze to stuff an extra row of drives in, lengthening the chassis to 33" instead of just under 29". It is their design, so they're free to do what they need for their own racks, but this does mean anyone trying to install a pod based on the 6.0 design couldn't use many closed-back racks or other kinds of racks.
Also, the 6.0 housing will fit in most racks. If somebody can either measure their racks or include a link where the 33" chassis will no longer fit, we really want to hear about it!
Later in the bullet list of changes:
"Added 3 more backplanes and 1 more SATA card."
But that was back in the late 90's/early 2000's, so things might have changed.
Another guideline in there that applies here is the one about not repeating the name of the company when it's already mentioned in the domain to the right of the title. That's a minor point, of course, but one that has had surprisingly good effects.