Storage Pod 4.0: Direct Wire Drives – Faster, Simpler and Less Expensive
blog.backblaze.com
blog.backblaze.com
I have a drive that I may be away from (at least, my Macbook Air is) for weeks at a time, and I don't want to continually lose version history or have to reupload some of the larger files (home videos and my photos).
CrashPlan doesn't delete your backups after 30 days (they keep everything as long as you're subscribed), but the upload speed is so horrendously slow I couldn't upload the entire photo library plus videos (~300 GB or so) in the month I was subscribed. I'm on a gigabit network over Ethernet, so it's not my connection that's the issue (Backblaze, on the other hand, uploads this data very quickly).
I realize that they're trying to reduce costs by deleting data that doesn't seem to "be in use", but I'm sure a lot of people have "archival" data that they don't look at often, but really value.
Instead, I'm currently using Arq 4 (which is fantastic, by the way) with Amazon Glacier for the photos. It may cost something to retrieve them if my drive dies or is physically destroyed, but I'm not planning for that to happen. I have Arq set to use S3 for my other documents and such so I can restore versions whenever I need to, so it's really the best of both worlds.
I finally ended up just purchasing multiple hard drives and using superduper to clone/backup.
Crashplan wasn't an option - would take 60 days to do a 1 Terabyte backup, and I don't leave my laptop connected to the externals that long.
I'd be willing to pay more money to backblaze to have them retain my external hard drivers for a longer period of time.
The second problem with Backblaze is the recovery process: they decrypt the files on their servers and store unencrypted zip archive for 7 days or until you download it. I would much rather see them create a restore process where encrypted data is downloaded and then decrypted locally.
I’m evaluating Arq, and so far it looks very good. Upload to S3/Glacier is super fast.
This was over a year ago, not sure if it's still the case, just be sure not to rely solely on them.
That problem, at least, seems to have improved. I backed up ~250GB a couple of months back and while I don't recall the exact time - might have been several days - it was certainly well under a month.
I've noticed Supermicro offers a 45-drive 4u chassis* that costs more than the Storage Pod's raw parts cost, but less than if you buy one preassembled by 45Drives. Does anyone have any experience with Supermicro's solution?
* Part: CSE-847E26-RJBOD1 (http://www.supermicro.com/products/chassis/4U/847/SC847E26-R...)
One quick note on this particular Supermicro system - it's slightly apples/oranges as the Backblaze Storage Pod is a complete server and this Supermicro system is a JBOD, meaning it still needs to be plugged into a server to work.
Gleb, Backblaze co-founder
http://www.supermicro.com/products/chassis/4U/?chs=417
(Disclaimer/answer to the parent post: we use the 24 drive unit for the GPU compute nodes in one of our clusters, and the 45 drive JBOD units for storage nodes in the same cluster. We have had a very positive experience with both (to the point that I got the 24 drive one as my home fileserver), as well as the Supermicro customer support for such.)
What happens when a Backblaze pod power-supply fails? I'd love to see a post about that. :)
David, KeepVault CEO
From what I understand, the whole pod becomes unavailable. Which is why you would use a front-end system to have redundancy across pods.
The way I would do it myself, is set up network connections between two pods, and use DRBD, along with clustering software for the iSCSI or NAS (nfs/samba) daemons.
http://www.backblaze.com/petabytes-on-a-budget-how-to-build-...
For replicating a similar architecture, I'd probably look at HekaFS/GlusterFS and/or CEPH.
Now that prices have come down so much on the 3rd party stuff, have you evaluated / considered using any of it?
I have a lot of experience with using the Supermicro chassis for storage. The question to ask yourself when deciding between "Supermicro or Backblaze pod" is: "Is storage density important to me?"
The final costs of a loaded Supermicro chassis vs a loaded Backblaze chassis are pretty close. In fact, the Supermicro may be a better option if storage density is not important to you (e.g home/office or inexepensitve city - http://imgur.com/gallery/QfD6qIw) But if your server is in a location where square-footage is expensive, the storage density is the primary metric you're looking to optimize for, since rent (and electricity) is your primary ongoing cost.
To run a scalable storage solution with the Backblaze pods you need GREAT SOFTWARE. Maybe a really well configured ZFS pool/custom software/and great sysops? As far as I can tell that's Backblaze's secret sauce.
For a company that wants to run a server that's going to be "pretty good" for what they need, using a Supermicro chassis with hardware RAID is definitely going to get you by for up to 96 TB.
At least we don't have to deal with master and slave jumpers anymore.
Let's take a look...
They do have an oldish post on hosting costs: http://blog.backblaze.com/2011/07/20/petabytes-on-a-budget-v...
A rack in 2011 cost them $2,100 per month (it's more now, guaranteed). Their racks look to hold about 11 machines. Each pod is then about $191/mo.
Each rack thus needs 1.375 extra pods at "low density". That's $262.50 a month. Since you now need about 10% more pods, you're paying about $29 more in hosting costs per month per pod.
Your savings are not $688, since you also need to construct 1/8 pod later. A low-density pod will cost about $2719, with is $340 split 8 ways. So real hardware savings is $348. Hosting costs eat that in exactly a year, and they are certainly expecting a longer life than that. (That's 2011 hosting costs, which are certainly lower.)
So it's worthwhile to spend a little extra for the additional drives. (Of course, if they could get the extra five in with more modest hardware, as seems quite possible, that would be a sure win. But sometimes ease of development/assembly/spare parts wins over hardware cost.)
Still, at least 12% extra rack and floor space to get the same storage capacity would probably cost them more in the long run.
Deleted comment
Wouldn't a backplane be more analogous to a SATA card than a single connector? In which case, 40 drives will go down on failure.
Neat stuff though. Would be fun to build one.
On a side note, each of the new SATA cards can support up to 40 drives, but there are only 45 drives in the pod, so one failing would presumably take out 22-23 drives instead of 40. Still worse than losing 5 drives though.
But that is finally going to change. As Google lower their price, i see this as HAMR finally coming. With both Seagate and WD promised to get 20TB HDD by 2020, and with a 60TB version after. We have a clear roadmap of what is coming along. Which means Google could now price their storage by it.
I'm thinking that for small-medium deployments, (with archiving in mind) three pods would be a reasonable minimum, possibly starting them out sparse (single radi6 to each pod, say)?
I'm aware there aren't any silver bullets :-)
Amazing and I hope they continue for a long time!
Obviously, for them it is cheaper to live without PSU redundancy because it is a normal day for them to lose disks and PODs because they've got so many of them, and they designed for this scenario.