All the data is encrypted and spread throughout our datacenters. So we fetch each file and decrypt it (with the password you provided when kicking off the restore).
If this is less than 1 GByte it is instantaneous. But if it is a TByte of data, it takes a few hours to decrypt and add to the ZIP file.
There are some things that can affect it. We have a "pool" of restore servers that do this task. The oldest ones in the pool are 9 year old computers with slow hard drives. The newest restore servers are built on SSDs and are blazingly fast and have newer processors. So based on which restore server was assigned the task, plus other things like the size of each file can speed or slow the restore.
We are in the middle of a project to speed up all restores. Hilariously we figured out that simply by decommissioning the oldest, saddest restore servers the restores averaged FASTER restore times.
1. Why is the data being decrypted server-side instead of client-side? The server should be sending the encrypted data to the client, and the password should never leave the client machine.
2. Why are zip files being created? The client software should be receiving the data and writing the files directly to the selected restore directory.
e.g. this is how restore works with CrashPlan, and it only takes a few seconds to begin the process. Files are written directly to the selected output directory.
> 2. Why are zip files being created?
We actually offer two forms of restore: A) Zip File Download, and B) External USB Hard Drive FedEx'ed to your home.
In the case of the Zip file download, we chose the format of zip because both Mac and PC (the most common desktops) natively understand it with no additional software needed. In other words, if you just lost your computer, go to ANY COMPUTER ANYWHERE and you can fetch your files with a web browser. Zip preserves the file hierarchy and the last modified time, so it's "pretty good".
In the case of the USB Hard Drive there is no zip file, we can correctly place each file with all the correct timestamps in the correct heirarchy on the USB Hard Drive (which is actually an encrypted hard drive) and then FedEx the hard drive to you anywhere in the world in a day or two.
> 1. Why is the data being decrypted server-side instead of client-side?
Short Answer: Ease of use.
Longer Answer: To clarify, Backblaze produces four different products/modes for different customers with different needs and requirements. We want customers to choose what is appropriate for them. One size does not fit all:
1) Online Backup ($5/month) where every file is encrypted on your laptop BEFORE being sent to Backblaze and your backup is secured by your username/password - where you can recover your password if you have access to your email account. (We support two-factor auth which provides an additional optional layer of protection.)
2) Online Backup ($5/month) where every file is encrypted on your laptop BEFORE being sent to Backblaze and your backup is secured by your username/password AND your private encryption key is secured by a "passphrase" that is not recoverable in any way, shape, or form. (Two-factor auth is also optional here.)
3) B2 Object storage (half of 1 cent/GByte/month) where you store your file completely unencrypted, and this can be "private" (only accessible by username/password) or "totally public accessible by knowing the URL". A good application of this is serving up a web page to the public - you really WANT people to see all the contents!
4) B2 Object storage (half of 1 cent/GByte/month) where Backblaze has zero knowledge. You cannot browse your file hierarchy because Backblaze doesn't know your filenames. You cannot preview your images. You cannot recover your passwords. There is no other option other than downloading the encrypted blobs and applying whatever decryption algorithm you decided on (we have no ability to know what that is).
Ok, so I think some (many?) people in the security field think that Backblaze should ONLY offer mode #4 (and maybe #3 to serve up public websites). I happen to disagree and I personally feel that products #1 and #2 are useful and appropriate for some customers. But everybody is welcome to their opinion and we want to be completely open as to what exactly is occurring and what we are offering as a service.
Personally I think #2 is an excellent trade off of security vs convenience. Your data is as impervious to attack as a zero knowledge system in #4 for years upon years. Then one day your laptop is stolen or crashes and you want your files back. You want all 4 TBytes back - so you order one of our free (encrypted) USB hard drives to be FedEx'ed to your home with all your data. To kick this process off FOR THE FIRST TIME EVER you tell us your passphrase (up until this very moment it really has been zero knowledge). At this moment you are opening a window of SLIGHTLY lowered security that slams shut after a few hours. For those few hours of preparing your 4 TByte restore, if an undetected hacker had compromised the one restore server in the Backblaze data center that your job was on, that hacker could possibly get access to your files. But then the reduced security window slams shut, we NEVER write your passphrase to any disk so it has now vaporized and we do not remember it, and if a hacker hacks into our system the following day you are STILL completely impervious.
But again, as long as you fully understand the implications of #3 I am COMPLETELY supportive if you instead choose #4 which is our "Zero Knowledge" offering.
This is the part that security people get hung up on.
We have to trust Backblaze [0] to not hang onto the password longer than you say you will, either intentionally (say, "to improve the user experience") or unintentionally ("oh that memory got written to swap and happened to persist on disk for a long time, whoops").
So that's the "malicious" case (maybe some of the above described cases are "benign malice", doesn't matter). There's also the "subverted" case: if the decryption password enters your infrastructure, we have to trust that you have not been popped by an attacker.
Whereas, if Backblaze has zero knowledge, and only serves encrypted bits, the _only_ concerns customers have are with durability and availability, which it sounds like you have pretty well in hand from your description of how bits are distributed around your datacenter(s).
[0] not just Backblaze but 5-years-from-now Backblaze: people and businesses change.
In a commercial setting, you need to control the infrastructure and security artifacts. Otherwise, old fashioned physical controls are effective, trustworthy and have more legal protections than any of these products at a cost of some convenience.
Nothing personal, but I really don't want y'all seeing my data, even if I need to restore.
(context: I've been a Backblaze user for years and love the hell out of you. I've never really thought of how the restore would work with my encrypted data, though. Turns out the story is ... not what I had hoped)
Bonus points for open sourcing a tiny "backblaze decrypter" command line utility that can do this manually. Though maybe that's a stretch.
This is how CrashPlan worked. (And incidentally, as I'm sure the guys at backblaze are fully aware, CrashPlan home is shutting down and I'm currently struggling to find a replacement that meets my desires. So if backblaze started to offer that #2 + encrypted restore soon they would gain my custom at the least.)
We totally understand, especially about the (very real) security issues.
I would reassure you that Zip file restores are COMPLETELY automated and with thousands of them happening every day no Backblaze employee ever sees your files.
The USB restore drives are placed in an envelope by a human, but they are encrypted when they are placed in the envelope so there is still no danger of any person seeing your files.
Plus we take our customer's privacy VERY seriously and it would be a firing offense for any Backblaze employee to ever look at a single customer file (or file name) without that customer's explicit permission.
> In the case of the Zip file download, we chose the format of zip because both Mac and PC (the most common desktops) natively understand it with no additional software needed. In other words, if you just lost your computer, go to ANY COMPUTER ANYWHERE and you can fetch your files with a web browser. Zip preserves the file hierarchy and the last modified time, so it's "pretty good".
Security issues aside, from a space and time perspective, that's much less convenient for both you and me: it takes time and space for you to prepare the zip file, and space and time for me to save it and extract it. It would be better to install a simple client and restore the files directly into a directory. In the case of a few files, it doesn't matter much, but in the case of gigs of files, it definitely becomes a problem.
Is there any chance of you implementing a more CrashPlan-like restore option in the future?
> Longer Answer: To clarify, Backblaze produces four different products/modes for different customers with different needs and requirements. We want customers to choose what is appropriate for them. One size does not fit all:
I wish this explanation were on the Backblaze web site, because this makes it very clear. :)
Is there any chance of implementing a hybrid of #2 and #4, a client-based, zero-knowledge option? i.e. basically, CrashPlan's client in the private-key mode.
If Backblaze were to implement this, and add a Linux client, I'd seriously consider switching.
Thanks for your time.
> and space and time for me to save it and extract it
One advantage is that for millions of small files, batching them together into one monolithic ZIP file allows for faster downloading. As opposed to a round trip to fetch each file. But there could be a happy medium where we batch 1,000 encrypted files to download to your computer then split them up, decrypt them, and cleanup any temporary files before fetching the next 1,000 encrypted files, etc.
> Is there any chance of implementing a hybrid of #2 and #4,
> a client-based, zero-knowledge option?
Yeah, we get asked for a "client restore option" quite a bit (see other comments in this thread also). Realistically I don't see it happening in the next 6 months, but eventually we really should get it done.
Well, any such software that didn't have pipelining would be poorly designed. ;) Compare to software like rsync, Unison, etc, which transfer large numbers of small files quickly. Seems like a good solution would be to build on librsync.