Backblaze Vaults: Zettabyte-Scale Cloud Storage Architecture
backblaze.com
backblaze.com
"Right now, Backblaze has only one datacenter, so the short answer is "no". :-)
The longer answer is that for online backup, there is one copy of your data on your laptop, and another copy in the Backblaze datacenter in Sacramento. If a meteor hits our datacenter in Sacramento pulverizing it into atoms, you STILL would not lose one single file, not one - because your laptop is still running just fine where ever you are with your copy of the data. In the case that occurs, we will alert our users they should make another backup of their data."
There are a million one things other than a comet strike that can go wrong in a data centre. I would not trust a backup provider that does not replicate my data across at least two data centers.
If you're really lucky, you siphon off customers from the non-redundant service for this at the same or faster rate as they are signing up for the non-redundant service, allowing you to not have to build that out much for a short while.
I think the MOST IMPORTANT THING is that an online backup provider like Backblaze is totally transparent to their customers about their architecture and what we do and do NOT do. If our design is not reliable enough for your particular needs, then you are then able to make the choice not to use us. I never want to mis-lead customers and hide exactly how durable their backup really is.
Finally - what I recommend to my very closest family members and trusted friends is that for data you feel would be catastrophic to lose, I recommend you have at least three copies including the primary copy. That is two separate backups with TWO SEPARATE VENDORS who did not share any code, hopefully managed by two separate UIs. For bonus points, one backup should be "offsite". For example, many Backblaze customers use Time Machine on the Macintosh for a local backup, and Backblaze for their remote backup, and that is what EVERYBODY should be doing. I can show you many support cases where Time Machine failed to restore a file and Backblaze saved the day, and vice versa. The fact is that users make mistakes, UIs are hard to use, your 14 year old son decided to unplug the Time Machine USB hard drive to free up a USB port, or your 14 year old daughter decided to uninstall the Backblaze agent - in other words, stuff happens!!
I routinely recommend BackBlaze as a quick and easy cloud backup solution to family and friends. I view the risk of partial or complete data loss with BackBlaze extremely low compared to the local backup solutions that some people rely on (i.e. periodic backups to an external USB drive)
For myself and my own work files, I don't put complete trust in any single backup provider, regardless of what their replication policy is.
But I would still prefer not to be hit by any Meteors.
I have a fondness for Tarsnap and Arq, but I find Data Backup 3 from ProsoftEngineering to be one of the best for doing local backups. At the end of a day, I just toss in a 64 GB USB Key (with FileVault Encryption), and no matter how much data I've created, it usually is less than 10 seconds for a full versioned backup. SuperDuper Clone about once every couple weeks.
I originally quit BackBlaze because it fired up my CPU to 100%, no matter how much time I spent trying to tweak it, but now I'm running into the same issue with CrashPlan which, after a couple years, has started constantly scanning my filesystem, forcing me to sudo launchctl unload /Library/LaunchDaemons/com.crashplan.engine.plist, and then remembering to sudo launchctl load /Library/LaunchDaemons/com.crashplan.engine.plist before I go to bed to let it run overnight.
What is it about Crashplan not letting you just shutdown everything from within the application.
This type of service starts at around $175 for 500 GB of SSAE16-audited datacenters, but you get all the things that consumer based backup systems don't - such as Plugins for SQL, Exchange, Hyper-V, NetApp, as well as Backup & DR software licenses for an unlimited number of servers - and, most importantly, 24 x 7 US-based engineer-level support.
I hope I'm not the only one uncomfortable about this. I mean, I understand the need for greater flexibility and features that MDRaid doesn't provide, but this wording stinks of NIH and rebuilding the wheel from scratch "because it was fun", discarding the maturity and reliabilty of established software. And data storage is all about reliability.
I hope I'm just reading too much from this, and it isn't actually representative of Backblaze's engineering practices.
We did write the Reed-Solomon ourselves in a "clean room" so we did not have to pay any licensing fees and we clearly didn't steal anybody else's source code, but that is a very small amount of code. Like 80 lines of Java. Seriously. We referenced the technical papers we read to implement it in that blog post, but here it is again: http://www.cs.cmu.edu/~guyb/realworld/reedsolomon/reed_solom... And we unit tested the living heck out of that code, plus we mathematically verified various parts.
But I'm open to an alternative solution if you can suggest one? Remember our three highest priorities are: 1) reliable, 2) low cost, 3) simple. The "low cost" includes things like we do not want to pay ongoing licensing fees to other companies.
My first thought was that you could've reused the R-S code from mdraid or dm-raid or ZFS, but on second thought 1) it may be too specialized to be reusable, and 2) it's GPL (or CDDL), so you can't just plonk it into your own code.
And yeah, if it's just 80 lines of Java, I'm worrying about the wrong things.
Backblaze is a web service, so for better or worse, the GPL doesn't apply here, since we never have access to their binaries. The AGPL would apply, but that's not the license used.
So everyone uses MD, but why does it suck for Backblaze-like storage? It only provides RAID levels, so you have partition-level redundancy but not necessarily object-level redundancy. You still need another layer to put objects in the right places to achieve the level of redundancy you want, and to re-balance as necessary. And a RAID array is local to one system, so achieving multi-host redundancy means duplicating the data as many times as necessary.
MD can also be finicky. Since it's part of the kernel, kernel upgrades can (rarely) produce weirdness. You can get stuck between a rock and a hard place, where you need to upgrade to fix a vulnerability but upgrading too quickly puts your data at risk.
According to Backblaze, they used to do 3 RAID6 arrays of 15 drives each per pod. This gives an overhead of 1.15:1, and two devices can fail before you lose data. Write performance is not going to be great, nor is rebuild performance. That is probably part of the reason why they could only push 950MB/s per host.
This only provides disk-level redundancy, and not host-level. So now you have to at minimum duplicate your array onto a second host, and your overhead is 2.30:1. A third host brings it to 3.46:1. I'm surprised that they were even using RAID for non-boot devices at all, given the overhead.
Erasure coding allows you to safely store data with far less than a 2:1 overhead. Their current design claims that they can lose 3 storage devices before they risk losing data permanently, with an overhead of 1.17:1. That is pretty compelling from a cost perspective.
(I say this as someone who worked at a company where we rolled our own, but that was ten years ago and we didn't have those options.)
Its no wonder storage companies like EMC are hurting when you have innovators like these guys out there.
[1] Which give a 3x replication system and a sharded or (chunked) file can happen pretty quickly.
At our current scale it is becoming less and less of an open debate, because we now have 7 day a week staffing at our datacenter and the datacenter techs jump right in and replaced failed drives often within an hour or so. A "hot spare" would only save a couple hours of rebuild time. But remember, your mileage will vary - until you reach half our scale you cannot afford even a Monday-Friday datacenter tech, so you might only be able to replace failed drives on Mondays and Wednesday, which widens your exposure.
Something similar to how ceph or swift handles rebuilds? you get rid of the individual disk sitting around as a spare. though it would break the idea of a tome being a specific collection of disks. you would need to be able to identify and move a shard around your cluster into other vaults and a shard would need to be smaller than the raw disk size.
this would increase network overhead as well. (more movement.)
I'm probably just rambling here so you can probably ignore me. (you have awesome tech there though)
:-) Not at all! Don't assume we're some perfect team of scientists that know all the correct solutions before we start coding. We often angst over these decisions and designs, knowing that once we write the code a lot will be set in stone (hard to change) for a numbers of years. The reason it becomes hard to change is we don't have a huge development team that can afford to rewrite the software every year, so we try to get it correct and then go on to work on new things or polishing up corners that need polishing.
http://people.freebsd.org/~gibbs/zfs_doxygenation/html/da/dc...
http://people.freebsd.org/~gibbs/zfs_doxygenation/html/d1/d7...
So better than most, but still not arbitrary. And the limiting factor seems to be write performance.
For those looking to build something similar, check out ceph or gluster.
Is a single file spread across multiple data centers? At the claimed 99.99999% annual durability, doesn't the chance of a natural disaster that could take out the entire data center start being a major factor?
I realize that the customer also has a copy of the data so you don't have to take the same precautions as something like S3, but it'd be sad if a datacenter got taken out by a meteor or airplane crash the same day that the customer's laptop was stolen.
Finally, a question for backblaze devs. In your opinion, how often do you need to scrub a drive to check for problems?
edit -> I ignored your Backblaze dev question, sorry. We have multiple processes running at all times on the pods, and they go shard-by-shard. We're always optimizing, but the short answer is, we're always looking for errors.
What took down the DC for several weeks was the firedepartements investigation. They took their time to figure out the root cause of the fire which is their right but the collateral damage of that was substantial.
So don't just plan for meteors.
And what's luxurious about "backup" as a business is this doesn't bother many customers. As long as we keep communicating to them on the progress, and we assure them they are going to get every solitary bit/byte/jpeg/mp3/movie back - they often tell us to take our time and do it right. For "backup" accurate and durable is about a thousand times more important than "instant gratification".
That's a full rack of storage pods a workday. Some back-of-the-envelope math says that's almost a tractor-trailer worth of hardware a month. Wow.
Essentially raid is dead to me, and ZFS/BTRFS, etc seem to be the only way to go forward, so I hope they gpl the code.
For anyone from BackBlaze reading this, I'm curious though, have you found that backplanes are becoming a primary bottleneck? Because that seems to be the case (Sata 6gb/s hurts after using things like fusion-io or even thunderbolt.) Any insights into the future of backplanes?
If you want to pull a single file, you'll be navigating through a windows 95'esq tree. They store snapshots, but if you want to change the snapshot, you wait a minute for each while it loads. Even going back to a snapshot you were just looking at, you wait the whole load time.
Now if you actually need to restore something. You can download a zip file. They will only let you make the zip file so big, so you have to break up your restore into multiple zip files. You will have to do that manually, there is no way to have BB auto-generate the parts for you. These are zip files and not one of the many archive formats that allow for parts.
Besides that, you will need double the amount of storage to recover this data since you'll need to store the zip and the extracted backup.
The way around this is to pay BB to put the data on a USB drive (flash up to 128GB or external up to 4TB) at $99 and $189 respectively. A 128GB external usb on Amazon, first result is $120. They actually won't give you a 4TB unless your restore needs it, but the price is still $189. Labor I guess? According to their own FAQ, you will wait 2-3 days for them to ship these drives, so hopefully you don't need that restore anytime soon.
I really liked Backblaze up to the point I needed to use it for it's real purpose. It seems like nobody at BB cares about the restore process or that it doesn't sell new subscriptions.
--
Also, maybe someone from BB can explain why the secure.backblaze restore website loads tracking pixels from googleads.g.doubleclick, a.triggit, s.adroll,facebook, ads.yahoo,x.bidswitch, ib.adnxs and idsync.rlcdn. Are you selling my need for a new harddrive or something?
Brian from Backblaze here -> I care! It just keeps getting bumped by something higher priority. I have a spec for how to speed up the restore tree browsing, it's just waiting for us to have a spare moment. For a while the Vaults took precedence.
Part of running Backblaze without VC funding is we can only hire programmers when we can afford it out of profits, and we're up to about 6 programmers (the result of a recent burst of hiring) which handle all of: Windows, Macintosh, iOS, Android, and in the datacenter they built the pods, the Vaults, and the web front end. But we'll get there, I swear.
> or that it doesn't sell new subscriptions.
This is unfortunately the heart of the problem. The most important thing to get smooth as glass is the BACKUP part, that sells new subscriptions. If we have your data safe, we can always hobble through a restore even if it is a little slow and clunky we can get all your files back after your laptop is stolen. If it held up sales, we'd jump over and do the one week of work to speed it up.
Crazy amount of hardware involved, and Backblaze is the "small kid on the block" in relation to FB, Google, and Amazon.
(I've never been good at statistics.)
As a consumer, the type of failure you would experience, and the probability of experiencing that failure, given that Backblaze has suffered a vault failure, depends on how they distribute your data amongst their vaults. They don't explicitly say how they do this, so it's impossible to know for sure, but we can consider the two extreme scenarios.
Scenario 1: Each customer is assigned a single vault, and all your files are on it. In this case, if Backblaze lost a vault, you would either luck out and have your files on another vault and be completely unaffected, or get really screwed and have all your files on the bad vault, and lose them all. They've got 150 PB of storage, and each vault stores 3.6 PB of data, so we can estimate that currently you may have something like a 1 in 40 chance of having your data on any given vault. So under this scenario, you would have a 1 in 400 million chance of losing all your files.
Scenario 2: Each customer's files are uniformly distributed across all vaults. In this case, if Backblaze lost a vault, all customers would lose a fraction of their files. Again, using our estimate that they might have 40 vaults, you would have a 1 in 10 million chance of losing 2.5% of your files.
So up to now, we're basically just doing the math without questioning the assumptions of the model. In reality, I think your practical risk is mostly concentrated in things outside of the model: ie, an event that affects all of their vaults simultaneously, like a fire, earthquake, meteor strike, etc. If I had to make a bet about what that number is, I'd put it in the 1/10,000 to 1/100,000 range. In other words, orders of magnitude higher than losing data because some hard drives failed, or a backblaze employee spilled his coffee, or something like that.
Also, I'm not worried. If that probability only concerns data loss on Backblaze's side, even if it's 1/10,000, then that's still not the probability of actual customer data loss. Because for that to happen there'd have to be a simultaneous loss of data on the customer side as well. That probably extends the durability considerably.
My former boss used to say that 90% of all problems are cabling. His percentage may be off, but the sentiment certainly isn't.
Vaults belong to a cluster - so your backup only puts data on the group of Vaults assigned to that one cluster, and the vaults in that cluster don't contain data from customers on OTHER clusters.
Days, Weeks, Months, or Years?
Any plans for that or even just an API so I can write one?
So you would have to read 17x the data to recreate it. Given disk latency, network bandwidth, etc. I'm guessing it'll take quite a while to recreate a 6TB HDD if it fails.
Sending 102TB of data to the pod that's rebuilding would take forever, this is true.
Instead you have each peer pod be responsible for 1/17th of the parity calculations.
1. 17 pods each read 17 megabytes and sends them across the network.
2. Pod A gets all 17 of megabyte 1. Pod B gets all 17 of megabyte 2. etc.
3. Each pod calculates their megabyte of the replacement drive and send it off.
4. Repeat until 6TB have been processed.
So this way each pod reads 6TB from disk, sends 6TB across the network, receives 6TB across the network, and calculates one nth of the data for the replacement drive.
It scales perfectly. It's no slower than doing a direct copy over the network.
Just make sure your switch can handle the traffic (which it already has to handle for filling the vault in the first place).