The Next Backblaze Storage Pod
backblaze.com
backblaze.com
And now that multiple "storage pod-like" systems exist in the marketplace (not just Dell, but also Supermicro) selling 60-bay or 90-bay 3.5" Hard drive storage servers in 4U rack form factors, there's not much reason to build their own?
At least, that's my assumption. After all, if the server chassis is a commodity now (and it absolutely is), no point making custom small runs to make a hypothetical Storage Pod 7. Economies of scale is too big a benefit (worst case scenario: its now Dell's or Supermicro's problem rather than Backblaze's).
EDIT: I admit that I don't really work in IT, I'm just a programmer. So I don't really know how popular 4U / ~60 HDD servers were before Backblaze Storage Pod 1.0
What wasn't really a thing was servers with hotswap HDD trays on both ends (like the supermicros) and things that were designed with vertical hard drives dropped down from a top-opening lid to achieve even higher density.
It had hot-swap SATA disks (up to 512GB disks initially!) and was actually pretty cool and forward thinking
https://web.archive.org/web/20061128164442/http://www.sun.co...
They were incredibly cool for the time.
At the time I advised my company to skip the storage servers (which were overdesigned IMHO, but were unquestionably stout hardware) but that ZFS was an interesting filesystem to explore.
Unfortunately for them, no one knew how to administrate Solaris, so they... installed Windows. No idea if they actually succeeded in doing anything useful with the systems.
Hardware was great, but no linux/zfs support meant we didn’t stick with it.
A couple of years later we had standardised on supermicros with a similar density and ran linux on them.
Sun lost a fair amount of money from us (7 figures) by choosing not to embrace linux.
> things that were designed with vertical hard drives dropped down from a top-opening lid to achieve even higher density.
the backblaze storage pod system and some 3rd party derived designs have the SATA/power ports on PCBs facing upwards, mounted at bottom interior of rack chassis, and 3.5" HDDs mounted vertically down into the system.
Given the bandwidth of 60+ hard drives (150MB/s per hard drive x 60 == 9GB/s in/out), I'm pretty sure you need a decent CPU just to handle the PCIe traffic. At least PCIe 3.0 x16, just for the hard drives. And then another x16 for network connections (multiple PHY for Fiber in/out that can handle that 9GB/s to a variety of switches).
We're looking at PCIe 3.0 x32 just for HDDs and Networking. Throw down a NVMe-cache or other stuff and I'm not seeing any kind of small system working out here.
---------
Then the math comes in: matrix multiplications over every bit of data to verify checksums and reed-solomon error correction starts to get expensive. Maybe if you had an FPGA or some kind of specialist DSP (lol GPUs maybe, since they're good at matrix multiplication), you can handle the bandwidth. But it seems nontrivial to me.
Server CPU seems to be a cheap and simple answer: get the large number of PCIe I/O lanes plus a beefy CPU to handle the calculations. Maybe a cheap CPU with many I/O lanes going to a GPU / FPGA / ASIC for the error checking math, but... specialized chips cost money. I don't think a cheap low-power CPU would be powerful enough to perform real-time error correction calculations over 9GBps of data.
--------
We can leave Backblaze's workload and think about typical SAN or NAS workloads too. More I/O is needed if you add NVMe storage to cache hard drive reads/writes, tons of RAM is needed if you plan to dedup.
Which means 9GBps downstream (to be processed by the GPU) + 9GBps upstream (GPU is done with the data), or a total bandwidth of 18GBps aggregate to the GPU / FPGA / ASIC / whatever coprocessor you're using.
So that's what? Another 32x lanes of PCIe 3.0? Maybe a 16x PCIE 4.0 GPU can handle that kind of I/O... but you can see that moving all this data around is non-trivial, even if we assume the math is instantaneous.
---------
Practically speaking, it seems like any CPU with enough PCIe bandwidth to handle this traffic is a CPU beefy enough to seemingly run the math.
The issue is that Backblaze has a 2nd layer of error correction codes. This 2nd layer of error correction codes needs to be calculated somewhere. If enough errors come from a drive, the administrators take down the box and replace the hard-drives and resilver the data.
Backblaze physically distributes the data over 20 separate computers in 20 separate racks. Some computer needs to run the math to "Combine" the data (error correction + checksums and all) back into the original data on every single read. So singular hard drive can do this math because the data has been reliably dispersed to so many different computers.
Do those run on this machine? I imagine backblaze has redundancy at the cluster level rather than machine level. That allows them to lose a single machine without any data becoming unavailable. It also means we shouldn't assume the erasure code calculations happen on a machine with 60 drives attached. That's still possible but alternatively the client [1] could do those calculations and the drive machines could simply handle read/write raw chunks. This can mean less network bandwidth [2] and better load balancing (heavier calculations done further from the stateful component).
[1] Meaning a machine handling a user-facing request or re-replication after drive/machine loss.
[2] Assume data is divided into slices that are reconstructed from N/M chunks, such that each chunk is smaller than its slice. [3] On read, the client-side erasure code design means N chunk transfers from drive machine to client. If instead the client queries one of the relevant drive machines, that machine has to receive N-1 chunks from the others and send back a full slice. (Similar for writes.) More network traffic on the drive machine and across the network in total, less on the client.
[3] This assumption might not make sense if they care more about minimizing seeks on read than minimizing bytes stored. Then they might have at least one full copy that doesn't require accessing the others.
Given their scale and goal, it would be pretty wasteful to build it to max the writing speed of all hard drives. Considering you rarely write on the pod, you would be better of getting a fraction of that speed and writing on multiple pods at the same time to get the required peak performance.
In fact actually that makes much more sense to put that math on some ingests server and theses hard drive servers would simply write the resulting data. It makes it much easier and faster to divide it over 20 pods like they currently do.
Like most blade setups, like (random example): https://www.storagereview.com/review/supermicro-x11-microbla...
And no, you want to calculate checksums and fix bit errors right here in the RAM buffers you just read or received, because at such scales hardware is not error-free.
Isn't Backblaze's workload nearly exclusively writes? People back up far, far more often than they restore.
RS is normally used as erasure code: It's used when writing (to compute code blocks), and when reading _only when data is missing_. Checksums are used to detect corrupt data, which is then treated as missing and RS used to reconstruct it. Using RS to detect/correct corrupt data is very inefficient.
Checksums are also normally free (CRC + memcpy on most modern CPUs runs in the same time that memcpy does: it's entirely memory bound).
The generation of code blocks is also fairly cheap: Certainly no large matrix multiplications! This is because the erasure code generally only spans a small number of blocks (e.g. 10 data blocks), so every code byte is only dependent on 10 data bytes. The math for this is reasonably simple, and further simplified with some reasonable sized look-up tables.
That's not to say that there is no CPU needed, but it's really not all that much, certainly nothing that needs acceleration support.
I don't think they actually sold any.
[1] https://www.hardkernel.com/shop/odroid-hc2-home-cloud-two/
There are a small number of high end ARM server boards that could do it, but you’re not saving much money at that point. Might be more expensive due to lack of scale.
Amortize a server-grade CPU and motherboard across 60+ high capacity drives and it’s not really worth pinching pennies at the risk of lower performance.
Also, Dell and Supermicro have storage servers inspired by the BB Pods.
Glad to see this scrappy company hit this amount of scale; a long way from schucking Hard Drives
https://i.dell.com/sites/doccontent/shared-content/data-shee...
So did they do something custom (unlikely at this volume) or did Backblaze change their hardware approach?
It's their PowerEdge MX platform, which allows you to slot in different "sleds" for storage/compute etc. as needed. It can take 7 storage sleds for a total of 112 drivers per chasis.
[1] https://www.delltechnologies.com/asset/en-ae/products/server...
https://www.servethehome.com/dell-emc-poweredge-xe7100-100-d...
It's apparently the "Dell PowerEdge R740xd2 rack server".
We deploy our storage with a 2U "head unit"[1] that has the actual computer and all of the SSDs for boot and SLOG and L2ARC and "zfs special"[2] devices. Then we attach 60-drive JBODs externally to those "head units".
With this dual-front design you could have 24 2.5 drives in the first row and 12 3.5 drives in the second row and that would really be helpful when we sometimes need to quickly spin up an adjacent zpool without attaching another whole JBOD.
Sadly, I do not see any configurations with 2.5" drives on the Dell config page:
https://www.dell.com/en-us/work/shop/dell-poweredge-servers/...
If one row was 24x 2.5 drives and the other was 12x 3.5 drives I think we would start deploying these immediately ...
EDIT: Huh ... even the two, rear, boot drives are 3.5" which is sort of weird and gratuitous ...
[1] https://www.supermicro.com/en/products/chassis/2U/216/SC216B...
[2] Yes, it's actually called that ...
We sell a lot of Dell and for base models, it is very economical compared to self built.
The moment however we add a few high capacity hard drives or memory, all bets are off the table and it's usually 1.75-4x the price of a white box part.
I get not supporting the part itself, but, had them not support a raid card error (corrupt memory) after they saw we had a third party drive.... we only buy a handful of servers a month - I can imagine this possibly being a huge problem for Backblaze though...
Every year they post a summary of what models they're working with and how they perform, which is usually good reading. This is last year's: https://www.backblaze.com/blog/backblaze-hard-drive-stats-fo....
That appeared to depend on whether the vendor imposed massive markups on the drives. However they also mentioned service etc.: If they struck a deal with Dell, then Dell might be perfectly happy to sell the servers at a very modest profit while making their money on the service agreement.
Stating the obvious: Backblaze wants investors to value them like a SaaS company. This blog post suggests they’re more of a logistics and product company— huge capex and depreciating assets on hand. As a customer, I like their product, but they’re no Dropbox. If they would allow personal NAS then I could see them being a software company.
BackBlaze is definitely a SaaS company... though the quality of their offering certainly lags behind Dropbox, both in terms of feature set and user experience. They're also in a very competitive industry. Storage/backup is basically a commodity nowadays.
agreed, but I think s3 is the more realistic comparison than dropbox
We decided to stick with king of storage AWS S3.
may I suggest this might not be a typical usecase?
There's a massive amount of premined chia controlled by the "chia strategic reserve".
It will take a decade for the amount of mined value to equal the pre-mined value.
Note: I went to double check something on the Burstcoin website and realized today, June 24th they changed their name to Signum - https://www.burst-coin.org/
FWIW the Burstcoin community have been very helpful, and they have a Windows client which was nice for a "hobbyist" like me.
[0] https://geizhals.eu/?cat=hde7s&xf=1080_SATA+1.5Gb%2Fs~1080_S...
[1] https://geizhals.eu/toshiba-enterprise-capacity-mg08aca-16tb...
The direct impact on larger drives is entirely dependent on brand (Toshiba's Enterprise drives appear to be less in demand than Seagate's Exos, for instance), recording process (CMR versus SMR), and, to a lesser extent, power consumption. In virtually all instances though, the price has jumped significantly [0]. 16TB Exos more than doubled, for instance [1]. In large parts of Europe, large drives have been back-ordered since the beginning of April; my orders are scheduled for delivery in August, yet I would be surprised if I saw anything before September.
[0] https://www.tomshardware.com/news/analysis-hdd-prices-skyroc...
It's where they make most of their margin.
And then, most of the time, the drives they sell come on custom sleds that they don't sell separately as a form of DRM/lock in.
Then you get a nice little trade on Chinese-made sleds that sort of work, but not for anything recent like hot swap NVMe drives.
I'm sure BB were able to negotiate down a lot (Dell usually come down 50% off the list price if you press them hard enough for one off projects), but... yeah. That's how it generally goes.
It’s like outsourcing the EU telecommunications, which we’ve seen a bunch of articles about recently.
Software is now the only thing that separates services like BB, Dropbox, et al.
Don't cry for BB pods, they did their job and now you can get a higher quality chassis at a similar cost basis from Dell as a result.
I used to work in a data centre doing infrastructure; I’m a metal fabricator by trade and now run and operate a laser cutter full time; I’ve done plenty of industrial spray painting; and heaps of press operating; plenty of CAD / CAM experience.
With trainee under me, I could have built these for you in-house and been useful the rest of the month when fabrication and painting etc was done.
That would have been fun!
This one is particularly interesting as they discuss the logistical challenges of their own success in having to build more and more Storage Pods.
As always, a super fascinating read worth your time.
They're still only buying the assembled hardware from Dell.
The hardware components of the Backblaze Pod are commodities but the entire finished unit is not a commodity. E.g. the rough equivalent from 45drives is not a commodity: https://www.45drives.com/products/storinator-xl60-configurat...
Always cool to get insights into the business and technical challenges at Backblaze!
The article doesn't say that. It says:
> So the question is: Will there ever be a Storage Pod 7.0 and beyond? We want to say yes. We’re still control freaks at heart, meaning we’ll want to make sure we can make our own storage servers so we are not at the mercy of “Big Server Inc.” In addition, we do see ourselves continuing to invest in the platform so we can take advantage of and potentially create new, yet practical ideas in the space (Storage Pod X anyone?). So, no, we don’t think Storage Pods are dead, they’ll just have a diverse group of storage server friends to work with.