Our 6 TB Hard Drive Face-Off
backblaze.com
backblaze.com
These guys are probably one of the only qualified places to do this at scale and report real world results. Glad they do!
Other companies do talks and papers too, but the Backblaze info seems really unfiltered and close to the real world.
Cloudflare seems to be a similar company. I recommend them both to anyone somewhat technical who needs a CDN.
My only complaint about Backblaze is no Unix client. I thought about setting up a Windows VM to manage backups, but I was terrified that Gninjas would sneak into my house at night and shave my beard.
But the main reason we buy drives is we deploy 30 new pods a month right now, and it's increasing. So we deploy 1,400 drives a month or so in new pods. Sometimes we buy in 3 month batches (to get a break on price) so then you start talking about a 5,000 hard drive purchase in one shot.
At $260/drive, you can see how we care deeply about EXACT failure rates. Why isn't Google and Amazon publishing failure rates also? Why only Backblaze?!
According to their blog they took until 2011 to hit 10PB, which would be roughly 150 of the original 67TB pods. The 6TB drives are 4 times better, so they'd be able to replace all of those with just 40 of the new pods, I'm sure the savings in electricity and rack space would be too huge to pass up.
Then again, they have their own datacenter space now, so space probably isn't an issue but I'm sure electricity is. They said something around $2,100 a rack (which holds 10 pods) for electricity, space, and data. They were incredibly stoked to cut that in half by moving to 3TB drives with the 2.0 version of the pod so I'm sure cutting that in half excited them as well :)
can we trust them with our data?
https://www.ssllabs.com/ssltest/analyze.html?d=backblaze.com
Nothing unusual about their results.
For anything critical in business (customer list, billing, finance, etc.), I would recommend also using Tarsnap in addition to whatever other (non-AWS) vendor/s are used for "offsite"/"offcloud" backup.
You would need to compare the pricing here to Amazon Glacier, if anything.
Because it is "backup" we are pretty relaxed about shutting down our pods to do maintenance, like replacing a drive. But that only takes 10 minutes, we try not to have machines offline more than that if possible, or it causes support issues from people trying to restore files.
> Are you worried about the stresses induced by spin up/down and thermal variations?
Absolutely. Our datacenter techs are convinced that if they swap a drive QUICKLY the pod has a higher chance of coming back up without any problems, if they let the pod get entirely cool it statistically seems to have more problems when it comes back up - a full cool down and heat up seems to cause issues.
Any thoughts to building Backblaze 20' or 40' cargo containers preloaded with rack-mounted pods and dropping them in locations with cheaper power?
"Let’s review the Seagate and Western Digital drives so far:
Initial reliability (how many drives failed) – No failures.
Running reliability (3 months) – No failures
SMART Stats (3 months) – No error conditions recorded for the 5 stats that we utilize."
I'd like to see similar tests from companies like Amazon and Dropbox (although the latter probably uses the former), in various use cases.
I do not use them but they say "That’s why Backblaze backs up your data automatically and constantly looks for new and changed files to backup. Install Backblaze and never worry about losing a file again."
Unless you are running a database it seems like the most common use case for files is write once, update every once in a while. Document creation (code for us or office docs non-coders) would be a bunch of writes up front then follow the same pattern of update every once in a while
I suppose their reads are much less than a typical user, but I didnt think reads contributed much to eventual drive failure
For files larger than 30 MBytes we break them into 10 MByte chunks and only transmit the chunks that have changed. So the worst case is you insert 1 byte at the very start of the large file - this effectively changes EVERY 10 MByte chunk and we transmit it all. The best case is you append 1 byte to the end of the large file, because then we only have to transmit that one chunk.
It seems like if you're already going through the effort to see if chunks have changed, you can readily do some heuristic where you check the end chunk of both sides, and scan inwards until you find a change; something to that respect, further minimize replicated data? I'd assume you might have CPU cycles to spare, but I'm too far into assuming already to feel comfortable so I'll just hope you answer :P (thanks in advance, it's been great to read what you've written thus far.)
We just never got around to it. In practice, it turns out most well written programs with large data try not to insert one byte at the start of the file. For example, take your large Outlook "pst" file (commonly 1 - 4 GBytes). When you get an email, it seems to append it to the END of the pst file, plus update some internal tables. So a large amount of that won't change.
Also, the worst case is it wastes space in our datacenter (and some bandwidth) for 30 days and then is cleaned up anyway, so you can measure the theoretical amount of money to save and it won't profoundly change our business so we put it off. Not to say ANYTHING is off the table, we're just always swamped with some project or other. :-)
Their knowledge base entry at https://help.backblaze.com/entries/20203731-Can-you-tell-me-... doesn't clarify wether or not the secret private password is transmitted to their servers to decrypt the private keys there which would make the scheme rather weak.
Ideally, the private keys would remain on the user's PCs.
If you are concerned about security, the second mode is you set your own "pass phrase" on the PEM file. But oh my lord, make sure you remember that pass phrase because Backblaze never writes it to disk and if you forget it, NOBODY ON EARTH is getting your files. Not you, not the USA government, not Backblaze, that data is GONE. You cannot recover that password. This is a good security mode if you would be arrested if the NSA got a copy of your data, because we don't think it can be broken.
As brianwski wrote in https://news.ycombinator.com/item?id=8169040:
...[I]f you lose a file, you have to sign into the Backblaze website and provide your passphrase which is ONLY STORED IN RAM for a few seconds and your file is decrypted. Yes, you are now in a "vulnerable state" until you download then "delete" the restore at which point you are back to a secure state.
The reason they would need the private keys should be obvious. They are a data disaster recovery service. In the event of a disaster where all of your drives fail, or are inaccessible you would still need access to your data. So if the only copy of the private key isn't on their servers you would be shit out of luck, not exactly a good situation for a company like theirs.
But yeah, store the private key, but make it encrypted by a secure password only decrypted locally, and update documentation to make this clear.
I think even this is "too secure". I forget my 'secure passwords' regularly, and I'm sure I wouldn't be the only one. It's not a huge improvement over telling a user to print out a private key and keep it somewhere safe.
Given the likely threats to Backblaze, I'd encrypt user backup private keys with a threshold or other multi-key scheme. No individual Backblaze employee's key would be sufficient to decrypt a backup; you'd need the active participation (and private knowledge) of at least a few. That prevents individuals from snooping, and requires compromising several employees to decrypt backup data.
Hitachi/HGST, now owned by WD, also makes really fast drives. Their 7200rpm 2.5" drives are the fastest laptop drives I've ever used, except SSDs.
The service I'm aware of that is closest to what you describe is spideroak[1]. Last I used them (years ago) they were a little too slow, and had some other minor kinks -- but might be worth checking out again now.
It looks like they've also finally made good on delivering Nimbus.io -- both as source and as a service (as I said, it's been years since I looked at Spideroak...):
FYI - our largest customer backs up about 60 TBytes for $5/month. So is it "unlimited"? If you have less than that it will work, we have proof! :-) But we lose money on that customer, we survive on the "average" amount of data people have, and on average a laptop of your average home user has only a few hundred GBytes.
I'd personally love it if they (backblaze) started up a sister-service that catered more to what I want -- but I think the cold reality is that it's hard to compete with Amazon S3 on price -- and if you do, in a big way, they can just dump prices to sink your business. At any rate -- you usually don't want to compete on price unless you're the biggest service around -- you wan't the high-margin part of the market...
The ZFS pool itself provides backup for several other systems, and snapshots are made periodically by a cron job to provide a further line of defense.
similarly, most drive benchmarks are 'artificial' in any number of ways and come with another set of caveats. since drive testing is already part of backblaze's process for putting drives into production, it could be especially interesting to compare and contrast performance data from testing with data from production use.
i pick these nits because real-world storage performance info seems like a valuable public good (particularly for enterprise and enterprise-like usage of consumer drives), and i'm grateful to backblaze for offering it up.
Cache becomes more important as the input data becomes more messy; cache is one tool for turning bursty and random data into constant, sequential data.
Also, SSDs come with 10 year warranties [1], so, which exceeds anything you can get with a normal spinning HD. If anything, I'd have more confidence with my SSD than I would with a spinning HD - particularly as you don't have to worry about the drives failing to spin up over time.
[1] http://www.amazon.com/SanDisk-Extreme-2-5-Inch-Warranty--SDS...
http://www.boston.com/news/science/articles/2010/10/17/scien...
Regardless, this is something the market can correct for very, very easily. As helium supply becomes more scarce, the price will go up, resulting in greater supply. Most of it comes from natural gas, and is so cheap [2] it's not worth capturing. Oil trades for around $100/barrel. Helium trades at $100 for a thousand cubic feet (albeit in gas form)
Some of the problems with Helium is that Physics experiments use a LOT, and previously was so cheap that it wasn't worth trying to conserve. That's changing - [3] A recycling system can recapture about $12,000/year of lost helium for a single scientist.
From reading articles - apparently the problems isn't so much that the cost of helium is increasing - but that it's been so cheap because of the US Natural Reserves making it completely non-competitive to capture - they are basically giving the stuff away for next to nothing.
And, putting it in a $100+ Hard Drive that will last a half decade is a FAR better use than putting 10x that much in a $0.50/balloon that will last 30 minutes. (Or Macy's parade which uses 400,000 cubic feet of helium) In fact, it may turn out to be the best conceivable use of Helium in terms of a value per m^3 equation.
[1] http://www.gazprominfo.com/articles/helium/
[2] http://finance.yahoo.com/news/airgas-increase-prices-helium-... And on Friday, the bureau announced that it was raising the price for a thousand cubic feet for crude helium from $84 to $95
[3] http://www.nature.com/news/united-states-extends-life-of-hel...
http://www.forbes.com/sites/timworstall/2013/12/18/chinese-r...
How? Through transmutation? I can't see that being economically viable.
Helium is a noble gas (generally does not combine into molecules) and tends to leave Earth's atmosphere when released. There's no viable way to produce more of it.
If you throw enough money at a problem, it starts becoming viable/feasible.
Have a look at the Diamond production industry for comparison: http://en.wikipedia.org/wiki/Synthetic_diamond
Note, don't let the "synthetic" part fool you into thinking they're not real diamonds. They are real in every imaginable way except they're not "a girl's best friend".
Up until some point, it was also difficult to manufacture synthetic diamonds at large scale. When the price of the thing desired is right / high enough, technology will be found and whatever costs necessary will become feasible.