Backblaze Storage Pod 5.0
backblaze.com
backblaze.com
Their cost for the pod (which they say includes labor) comes out to 0.044/GB. The cost of redundancy is 3/20 drives, which would place the cost to 0.0517/GB of data. They are planning to charge 0.005/GB on the B2 service this is made for. That's 10-11 months to break even on the initial cost. Add electricity + maintenance costs (of which I have no idea what it is) and you get the break-even number. My guess is that it'd be 1 year. So however long each pod lasts past 1 year is gravy (aside from the electricity + maintenance costs).
Amazon S3 must be making a killing. As a customer you must be paying for the cost of the "pod" every 2 months.
Please excuse me while I buy stock in Amazon/Box/Dropbox/Backblaze/etc....
----
As a aside, as a personal user if you store over 1TB you are getting storage at a cheaper cost than B2 users. If under 1 TB you'd be better of using B2.
Seems to me that the average storage per user must be well under 1 TB on their unlimited plans.
You're also right that 1 TB is the approximate breakeven between Backblaze per-GB cloud storage & unlimited online backup by cost and most people are below that. However, the main reason for people to use the backup service is that we take care of all the backup functions (encryption, dedup, compression, restores, etc.)
> Please excuse me while I buy stock in Amazon/Box/Dropbox/Backblaze/etc....
Alas, Backblaze stock isn't yet available for sale ;-)
BTW, thanks for being a customer!
Gleb
YET...or is it? /runs over to SecondMarket.
For most use cases this doesn't matter and the price saving are well worth the reduced redundancy, but Backblaze B2 isn't well suited for being the canonical store of your data.
[1] - https://azure.microsoft.com/en-us/pricing/details/storage/
In the case of Amazon, I wouldn't be so certain that they're geographically redundant copies. I raised this issue on HN a few months ago and someone replied that within an availability zone the data is "a bunch of (datacenters) that are close to each other". https://news.ycombinator.com/item?id=10231677
I like how for GRS Microsoft specifically says "a second datacenter hundreds of miles away". I have yet to find such clear statements from Amazon.
Microsoft is definitely more straightforward in its claims. O365 for my org was in Midwest, midatlantic, and southwest, giving us a really good continuity story.
Also, the B2 service doesn't come with a backup client, AFAIK.
But there's no explicit license declaration anywhere in the drawings zipfile and in fact some of the drawings claim "PROPRIETARY AND CONFIDENTIAL THE INFORMATION CONTAINED IN THIS DRAWING IS THE SOLE PROPERTY OF BACKBLAZE MANUFACTURING. ANY REPRODUCTION IN PART OR AS A WHOLE WITHOUT THE WRITTEN PERMISSION OF BACKBLAZE MANUFACTURING IS PROHIBITED".
Maybe it's just boilerplate? Or maybe you really don't want to face competition in your specific industry. Regardless, an official open source license (even with commercial restrictions) would be much clearer than this.
It's not under any license, it's just 'free'. We've talked about putting the Storage Pods under a license, and may at some point since we get this question periodically and picking a license that people are familiar with may make it more clear that we really are giving away the design.
If you start selling them as your own product, maybe that would be an issue.
The creative commons and OSI are two excellent places to start looking at licenses. creativecommons.org and opensource.org
To be clear on our philosophy - we don't want to restrict people from building these for their own purposes, modifying them, giving away the design, selling the boxes, etc.
Copyrights here only affect the actual drawings of the product. By default, the author is granted copyrights. That means you can't publish your own book with those drawings, for example.
But to actually build the chassis, go for it.
I agree with the parent though, if I were in busines and wanted to use your drawings, I wouldn't touch them with such a license. Even as a personal hobby project I would bug you for (official & written) permission before using any of the diagrams.
Other than that, it is nice reading about the technical decisions that go into building such racks. Kudos for sharing the knowledge!
It would not be a problem for Back Blaze being acquired.
Serious kudos though are due for the Backblaze blog. You're all at least talking about what you're up to, which is awesome both for education around building large storage systems and establishing transparency with your customer base.
Unfortunately, a lot of the bigger clusters' teams won't/can't talk about what exactly they're doing and how they're doing it for trade secret reasons. If you want to learn about it, you have to come work for one of us. :-)
Also, the orangered hardware looks awesome.
I hope you guys aren't using SMR drives...
Edit: Also, when are you going to move out of AWS? I really would like our/your clients to use IPv6.
Where are you buying 40 TB hard drives? I'd love to get some ;-)
We actually don't claim these are 'state of the art', more 'state of the inexpensive'. You can buy faster performance storage servers, more redundant servers, etc. But I believe these are one of, if not the, lowest cost storage servers per GB. (And of course, open source.)
BTW, I'm a fan of Dropbox (use it personally) and, in general, think there is a ton of value in developers, organizations, and companies sharing what they've learned about storage. The better and lower cost it gets, the more interesting things other people can do with it, the bigger the industry, and the more everyone benefits.
(full disclosure - cofounder of Backblaze here)
> Where are you buying 40 TB hard drives? I'd love to get some ;-)
Me too. :-) Not 40TB, but larger than 4T and quite a few more of them.
> But I believe these are one of, if not the, lowest cost storage servers per GB.
Yep, I meant price per GB as well.
> I'm a fan of Dropbox (use it personally)
Awesome, thanks! I've not used Backblaze personally, but I know some friends who do and they're happy customers. Your guys' blog and rapport with the technical community is amazing.
And they have 2 10GiB ports, though I'm sure you can work with them and get more..
It's disingenuous to call this not "state of the art" when it's really quite close, and obviously is meeting a backblaze design goal.
[0]: http://www.newisys.com/Products/4600.shtml [1]: http://www.aicipc.com/ProductDetail.aspx?ref=RSC-4H
with custom-built 2.5" ssd enclosures you could probably fit more in the same space.
The connectivity to the drives is really quite bad, but now I'm just being picky. Thank you for the link, really appreciate the follow up!
> It's disingenuous to call this not "state of the art" when it's really quite close, and obviously is meeting a backblaze design goal.
So, I never said it didn't meet their design goal; it looks like quite a nice system, and it's better than their previous gen. Looks like they're doing things right, looks like solid engineering.
But I stand by the statement that, in the absolute sense, this is not quite the current benchmark on density or $/GB in the industry. That was the only point I was trying to make to the HN audience in general, and it's not intended as a critique of Backblaze specifically.
But there are very different amounts of resources (I'm assuming) invested in achieving those benchmarks than Backblaze can likely commit.
Ergo, Backblaze probably completely made the right decision for Backblaze. Both of these things are possible.
Come on - most companies reasonably can't. There's a WD 10T smr/helium drive, and seagate is selling that crap 8T video drive. Unless you're talking ssd 3d/TLC density (more $), or the vapor-ware 20T drives that some storage vendors hint at and don't let me file a PO for, I literally don't know what you're referring to (now I'm getting where I can't name names)
The reason I'm complaining is because of you don't state:
- what the current state of the art is that they're not meeting
- what they could do better with their design goals given that information.
You merely state that it's just not state of the art, and to come work at dropbox. This is what I mean by being disingenuous.
That's about all I can do. However...
> This is what I mean by being disingenuous.
I'm don't think disingenuous is what I'm being. Maybe I'm not being as forthcoming as you'd like. Maybe my comment wasn't useful without details I'm not providing. But I'm not lying or misrepresenting anything.
So let me try to be helpful without using proprietary information. Using information people have unearthed in this thread, we can do some math to prove I'm not completely bonkers:
96 unit chassis * 10T SMR = 960T, which is in the > 5x the density in 4U
SMR $/GB is quite good if you work with the right vendor and make the right tradeoffs.If you could manage that 960T with the same compute resources, you'd save a lot on compute, since that's currently a ~30% of their cost.
With custom enclosures and hardware, you can do more still. You can have deep relationships with vendors and have them customize firmware for your particular use case. You can take their discards that didn't quite meet spec. You can buy crazy cheap stock with high failure rates b/c your distributed system is insanely good at repairs. You get access to stuff that's not being sold yet.
However, let me state again--this would take a lot of work. You need to mold your software around your hardware profile with very small tolerances. You might need to move some I/O stuff in-kernel, and/or move the IP stack out of kernel. You might need to reinforce your DC's raised floors to handle the weight (spoiler: you definitely do unless you planned for these types of machines ahead of time). You need great operations, and really thorough hardware quals b/c your failure domain is huge. Network traffic will go through the roof if a box fails, since you're repairing so much data.
Dozens of people on a handfuls of teams will spend time attacking this problem from different angles.
So, again--I don't necessarily think Backblaze did anything wrong here. Doing all this might be a crazy idea for them and not worth it. My statement was: it's not state of the art in terms of $/GB or density.
But for some shops, they spend so much on storage they can afford invest really, really deeply on optimizing storage cost. That's where the state of the art is.
My architecture didn't have an IP stack at all: guards or front-end acted as I.P. gateway while internally using much simpler protocol that still worked on Ethernet. Inspired by Active Messages and FLIR protocols of past. TCP/UDP/IP is unnecessarily complicated with simpler message passing easy in software and FPGA hardware for offloading onto cheap FPGA's. A side note was using one or more I/O coprocessors like mainframe Channel I/O (see Wikipedia) to get huge throughput might be cost-effective depending on hardware selected. Esp if using embedded SOC for I/O offloading.
Note: Octeon II's can do many Gbps of line-rate processing for three digits and much simpler internally than a lot of stuff.
The SASD project for secure, hard disks also suggested putting logic in the HD or SATA controllers rather than the system. Made me wonder about putting I/O mgmt and most storage logic in custom PCI cards with lots of SATA connectors. Basically, embedded systems enough CPU to manage whatever you throw at them, just running what's good for that layer, and less watts/money than whole servers. Seems like at some point one could port that layer to straight hardware for great performance/watt although high, initial NRE. eASIC can do that relatively cheap, though, with the Nextreme's if a design is FPGA-proven.
That was just exploration on paper so I'm not sure how practical it would be. My focus was high security systems and storage which necessitate better isolation. Hence, PCI cards rather than servers. Interesting that there was a lot of overlap, though, between what I thought made sense and what you all ended up doing. Very interesting. Guess I'll keep proven tactics in mind in future designs.
So, candidate for offloading that's really commonly considered is "scrubbing". Essentially, protecting against passive bit flips by checking your data against your own checksums.
If you have some N TB of data under management, it's really expensive to be striping these large serial reads all the way to an application's userland. So finding ways to keep these as low-priority "background" reads (that always yield to interactive reads) in the scheduling sense that a.) notify the daemon in userland if they fail and b.) without requiring that daemon to be in the data plane, is high-reward. Ideally, the daemon can stay in the control plane so it can manage/report accounting on scrub scheduling and time since last-scrub per disk region, or whatever.
You can offload it to a kernel module to avoid `copy_to_user` (if you can't DMA) and/or context switching. Or even offload it to hardware--some custom host adapter, possibly, using custom ATA/SCSI commands to control it and query it. (and the `sg` driver).
"Or even offload it to hardware--some custom host adapter, possibly, using custom ATA/SCSI commands to control it and query it. (and the `sg` driver)."
Now you're thinking. Chips like Octeon already do hardware acceleration of checksums and possibly other I/O boosters. So did mainframe I/O. There's many for FPGA's and ASIC's. Cray's ChipKill technology, implemented in Oracle SPARC CPU's, does something similar for RAM in combination with ECC RAM. So, an ASIC with several SATA connectors and logic for this should be able to do it. Might handle, partially or totally, the protocol for processing the data, too.
Wanted to see if I was really thinking ahead by Googling for academic papers. If it's a good idea, there's usually a few people that have done it for years already. Turned up some hits, including on your Reed-Solomon.
FPGA-accelerated decision-tolerant coding for reliable distributed storage https://dl.acm.org/citation.cfm?id=1763276
FPGA-accelerated, flash store for analytics http://people.csail.mit.edu/wjun/papers/fpga2014-wjun.pptx
DB acceleration http://www.cse.buffalo.edu/~vipin/papers/2010/todd1.pdf
Filesystem in GPU. Shows GPU's or DSP's can be used for this. Probably Adapteva's chips, too, esp since they might negotiate for volume deals. https://people.csail.mit.edu/idish/ftp/silberstein13asplos-g...
Tilera's were used in a 100Gbps NIDS that I recall. I expect your protocol and algorithms to be simpler. ;) http://www.tilera.com/products/
So, goes to show either custom, FPGA, or better-suited chips can handle that part at great performance per watt. Cost varies with implementation strategy. Nonetheless, simple functions with streaming data model greatly benefit from this strategy.
Idea is to move straight-forward, possibly parallelizable algorithms off general-purpose, legacy CPU's onto cheaper, efficient ones more suited to algorithm. It wasn't uncommon for FPGA's to get an eightfold boost in performance on stuff like this in various experiments.
In my field, high assurance INFOSEC, the main advantage is the use in inline, media encryptors. I really just knocked off specs and ideas of NSA's Type 1 IME given it was the best:
http://m.nsa.gov/ia/programs/inline_media_encryptor/
Also, Orange Book required full mediation and even SCOMP system mediated hardware I/O. Copying these lessons from the past made my design immune to DMA, firmware, and SMM attacks that were later discovered. Like above, offloading gets performance boost cuz inline I/O handles interrupt handling, prevents many cache flushes, and supports hardware accelerators. Also, keygen and storage stay off main server. Fun times.
Just to ponder on the available confirmed info - lets add Samsungs 16TB SSD to a 48 disk 1U rack and times that by 4...
I'd say just with a lot of cash and stuff you can buy now you can pass that order of magnitude pretty easily. If you make your own...I can easily see the densities knocking Blackblaze's pod out of the water.
As for costs, well massive ordering, preferred rates, custom designs and symbiotic relationships with manufacturers, coupled with a software architecture that complements all these hardware gains could conceivably bring costs down to levels unheard of outside of the big cloud service providers.
The conditions and technology available right now make what you are saying within the realms of the 'adjacent possible' if you have enough money to throw at the problem. With the kind of clout Dropbox et al have I can well believe you're getting incredible storage densities but at incredible knock down prices.
I'd really like to see some usage stats on your SMR clusters because so far the universal verdict is that they're terrible for anything other than WORM.
Usual tradeoff--more work in software, but far less $.
The power guarantees in the zbc spec (as of r03) seem to not be held by some drives. Writing less than a full zone seems to be a bad idea.
Well, specifics within a hypothetical. Hypotheticals are always game to to discuss. :-)
> Does REPORT ZONES not point to this conventional section always?
Yep, REPORT ZONES includes the conventional zone.
> I cannot wait for a FORMAT ALL ZONES
I agree that would be nice, but there are no indications that anyone but the vendor will control conventional vs. sequential layout, sequential sizes, or maximum # of explicitly open zones, anytime soon (absent asking for special favors for a large batch order).
5x the density
so half an order of magnitude, which is much more believable?
Many years ago I liked to say that, and a friend would call me out. So I would promptly change my statement to "a binary order of magnitude".
5x the density
Maybe he should be saying "a quinary order of magnitude"? :)
> so half an order of magnitude
Almost 70% of an order of magnitude.
Slightly off-topic but could you go into some detail on what you think the drawbacks are for the Seagate 8TB SMR drive? I've been thinking about getting a few for large files while keeping them online-accessible, I've found it difficult to get the thoughts of grizzled IT folks on the matter :-)
I couldn't find any reliable data, but I did find a lot of reviews that said the 8TB drives tend to die within a few months; I'm holding on to the 6TBs for now.
Gonna look more into how they lay the disks out. I wonder what the serviceability of the drives is.
So in short, I can think of improvements but most would come at the cost of higher capex to do so, and that could negatively impact the $ / GB numbers that Backblaze is trying to keep down. Perhaps there's not much consideration for leasing datacenter real estate or power costs or something into the equations but without some reliable numbers it's all speculation to me.
Dropbox wants all this trade secret non sense.
And everyone else give them their data.
That they promise they won't read.
Give me a break. Bittorrent sync and blaze pods in house for the win.
What order of magnitude?
Well, this article describes storage pods with 180TB of storage.HGST have a 10TB hard drive measuring 26.1mm * 101.6mm * 147mm [1] and a 4U chassis measures 178mm * 437mm * 521mm [2]
If you had 4U of hard drives with no space for power, cooling, interconnect etc you cold fit 10 terabytes / (26.1mm * 101.6mm * 147mm) * (178mm * 437mm * 521mm) = 1.04 petabytes into 4U
I don't know whether six times counts a a single order of magnitude, let alone multiple orders of magnitude, and the 10TB hard drives I linked are "enterprise" so there might be a substantial price premium.
Samsung claim to have a 16TB 2.5" SSD in the works [2] - although details are pretty scarce at the moment. If its dimensions match normal 2.5" SSDs at around 100mm * 70mm * 6.9mm that would give 16 terabytes / (100mm * 70mm * 7mm) * (178mm * 437mm * 521mm) = 13.2 petabytes in 4U.
Of course, the assumption you could use 100% of the space in a chassis for drives is pretty unrealistic. And nobody's seen the Samsung SSD working - could be a 16tb sticker on a 1tb ssd for all I know :) Filling it with 4TB SSDs would give you 3.35 PB and it actually seems possible to get those [4].
If you don't care about $/TB, it's probably possible to get density a single order of magnitude higher.
[1] http://www.hgst.com/products/hard-drives/ultrastar-archive-h... [2] http://www.supermicro.co.uk/products/chassis/4U/842/SC842TQ-... [3] http://arstechnica.co.uk/gadgets/2015/08/samsung-unveils-2-5... [4] https://www.sandisk.com/business/datacenter/products/flash-d...
"Archive drive" => shingled writes. Are people really using these things for online workloads? Might make sense for Backblaze's original "write and forget" backup workload but not very well for B2 or Dropbox
I found Backblaze and never looked back.
Blog post or it didn't happen :)
Seriously though, there seems to be some really cool engineering at Dropbox, first with Go usage (https://twitter.com/jamwt/status/629727590782099456) and now with your storage capacity. We're all craving for more infrastructure details, see what you consider as the state of the art and how you use the different pieces of software/language !
So basically you decided to go into this thread, post something that isn't possible for anyone here to corroborate and call the original post "far from state of the art"...to what end? What was the purpose of your posting? Just to toot your own horn? I don't get it.
If you want to talk about how awesome or state of the art your company does things then great do that. But your actual post? It just comes off as childish like someone showing off they can do something and you're that person that has to stand in between them and the people watching them to say you can do it too and better. Surely we've all seen those very awkward moments in sitcoms.
> We're (Dropbox) packing up to an order of magnitude more storage into 4U.
Are you sure? They can pack up to 180TB into 4U so you're saying DropBox can pack up to 1,800TB into 4U?
http://www.supermicro.com/products/chassis/4u/847/sc847de26-...
I've repeatedly tried to state in subsequent children in this thread that it seems as though Backblaze made the right set of tradeoffs given their previous gen, and their time to recoup their investment in this, etc.
I was also (somewhat selfishly, I suppose?) trying to tap into the excitement of a group of readers who where thinking about big storage, and let them know that the possibilities were even great given greater investment.
The responses I've been getting (including yours) lead me to believe that my comment is being seen as a slight against the set of tradeoffs that Backblaze has made, and that really wasn't how I intended it.
Edit: I do disagree about it providing zero value, though. We've had a good subsequent exploratory discussion about ways to optimize storage density and (to some degree) cost. And I hope we're not at the point where everything said on the internet demands "proof"; I'm not sure a blog post would provide that, either!
It's very easy to do what you just did and quite hard to do what Backblaze did. Essentially, if you're not willing to do what Backblaze did here the graceful thing would be to either stay quiet, simply compliment them on their achievements or, and most useful, to be specific about what could be improved.
I'll leave this thread alone because there's some useful technical information in here, and amend with an apology.
I get excited about large storage systems since that's what I've been focusing on for the last few years, and my enthusiasm overran my tact/prudence. Sorry Backblaze.
The best I have is 9 addressable disks plus a ludicrous CPU and GPU in a Silverstone desktop case, not very exciting but still going against the grain.
Love that you're fighting the good fight (along with the Internet Archive) to keep a history of the Internet.
The sinister benefit is that the more folks read the posts, the more it gets shared (presumably) and we hope that can lead to some new folks finding out about our online backup service, and therefore signing up. Which would make our finance department feel warm and fuzzy too!
For what it's worth, I've been a happy customer for years and try to spread the love whenever I can. My mantra is: data that's not backed up doesn't exist, and data that's not backed up off-site is not backed up. Thanks for providing a great service for cheap.
/MarketingOff
*Edit -> man, we have to optimize that page for our new design layout...
/yells at design team.
Also, when we work with contract manufacturers on building storage pods, talk with data centers about space, etc., it's really convenient to point to public posts to explain what we're talking about rather than having to put NDA's in place.
Overall it results in a lot of awareness and goodwill. It's incredibly gratifying after the team spends many late nights putting these together.
Gleb from Backblaze
Does it got any kind of integration with stuff you've uploaded through the regular Backblaze app? (e.g. moving stuff over to B2 once it has been uploaded as a regular BB backup).
I can write some script myself to move over my stuff directly, but even as a programmer I'd rather not :-)
I think you're plan for it is that third parties might take that task. So, maybe Arq will.
From afar it looks very compelling, but unfortunately I can't use it, because no Linux :(
I just found this[1], which looks like what you want. It's not a library, but it is only 245 lines of python, so I doubt it's that complex.
I would like to tell the provider; limit downloads to $500, then reply 509 to further requests.
Is there a need for that large of a boot drive? I would have expected that an SSD boot drive would be preferable.
though i think a PXE or similar boot option would be a smarter move. just enough to bootstrap something to run in ram. (alternatively a minimal kernel and dedicate a section of the disks in the chassis to holding the running software (500MB x 45 with decent levels of erasure coding is a fairly minimal slice.) the kernel could then pull down the code after first boot (or updates)
And a raid1 pair of HDDs for the system disk is more expensive than a small SSD, more fussy, and the SSD is still less likely to totally fail.
[1] https://www.backblaze.com/blog/vault-cloud-storage-architect...
Storage execution seems top-notch. End-user software deserves some love.
I have experience with the Mac client, and it's rock solid, and very well executed.
[edit: spelling]
Example is the motherboard for $500+, the fans for $20 each, on off switch for $25, 8gb of DDR3 ram for $90???? etc.
The saving grace is the hard drives that are decently priced and make up the bulk of the cost.
Gleb
For example, microatx mobo: http://www.newegg.com/Product/Product.aspx?Item=N82E16813157...
some non-ecc DDR4 ram: http://www.newegg.com/Product/Product.aspx?Item=N82E16820231...
and a cpu: http://www.newegg.com/Product/Product.aspx?Item=N82E16819117...
and motherboard: http://www.newegg.com/Product/Product.aspx?Item=N82E16813157...
Once you account for the cost of a 10G NIC and the fact that having it integrated frees up another expansion slot, I'd say it's a pretty good deal.
Also, I'm not sure if there are any non-Xeon processors that support ECC.
in fact do you have the storage located in various locations? if so can I store a paid-copy to a separate location physically, just to be safe.
Gleb from Backblaze
But server-side decryption :(
I have been looking into Backblaze every time it was mentioned on a podcast, and as far as I can tell you still cannot download your backup and then decrypt it on your own machine:
https://en.wikipedia.org/wiki/Comparison_of_online_backup_se...
Such a pity. I'm still stuck with CrashPlan (yuck!) for this one reason. I've been evaluating Arq, but I'm not yet confident enough to move all my/relatives' machines over.
I'm really more interested in a long-term backup storage instead of a folder-sync storage, cold storage is fine but glacier seems a little hard to use and retrieval is expensive.
I'm particularly excited to see how their B2 service will stack up to Glacier for half the cost.
"Is this going to work for static asset hosting such as for low cost FE assets? Can this be set to serve an index file from the root directory? Doe's it provide you a public URL to your (stealing s3 terms) bucket? Just curious as to why I would migrate from S3 for FE assets to Backblaze."
Was wondering if any progress has been made since then, especially relating to serving an index document from root dir.
Thanks!
edit: Awesome blog post btw
I think it'd be pretty interesting to expose it as a Windows filesystem with Dokan [0] and I'm sure Jeff from ExpanDrive [1] is looking at it too!
Are 4TB drives significantly cheaper for you? Or is there some other constraint at play against the 6TB?
We're not really in that space, but we're always hiring :D
Where exactly does one buy this?
Initial sync is a nighmare. The upload is on the slow side. Biggest gripe so far is that I just cannot tell the client that some files are more important than the others. I would love my pdf invoices and changed KDBX files to be synced before uploading the uncompressed blu ray of Guardians of the galaxy. On the other hand there are people that produce a lot of critically important raw video which will probably have the opposite preferences.
I can understand why they have not implemented priority interface. It is hard to get it right.
Also, I would hope no one doing mastering is uploading their raws, in real time, to a service like backblaze. On-site thunderbolt / NAS solutions are almost exclusively used in RAW backups in studios. With offsite copies stored at regular intervals via good old sneaker-net.
BB folder and file type structure makes it relatively hard to back up just some files of given type. And some stuff like tars and zips could be arbitrary in size and importance.
Having said that, we offer the ability to exclude folders and filetypes if you have certain things you want to backup later.
(Not disagreeing with anything you're saying...just fyi.)
Gleb from Backblaze
Backups transfers are also a minor pain point - the error horse (on freshly installed BB on windows) and when, how and why it disappears and allows you to do the backup is still a mystery.
I guess that the BB cloud storage will solve some of the customization problems though - if you are power enough user you could write a custom strategy on top of the API.
What's the point of backing up anything?
> Either you have the source material, in which case its almost always going to be faster to rerip than restore from backup
Sure, it will be faster, because it will be a local operation, because you probably keep the original near the same place you keep the main rip. That also means that the original source and the in-use rip are vulnerable to being destroyed together in a site-affecting event (fire, flood, robbery, etc.), which makes having a remote backup valuable.
WRT priority backup - I ended up just excluding the "big folders" manually (reducing backup size by 90%) then including them when the initial backup of the important files was complete. Also, the client backs up small files first, then moves onto bigger files, which is a nice touch - generally I'd tend to care more about losing a large number of files than a few huge files.
Not affiliated with the company, just love their product!...
In your case, I guess it would be much better to look at other solutions. For example if you have a Mac, get an extra Time Machine drive and put it at your employer's place. Every week, bring your laptop there and back it up.
"As part of our backup process, Backblaze will run a checksum against each file before uploading it. This requires the entire bzfileids.dat file to be loaded into RAM. After a long time, or if you have an extraordinarily large number of files, the bzfileids.dat file can grow large causing the Backblaze directory to appear bloated, and use more RAM (potentially slowing down your system). In your case your bzfileids are extremely large, likely because you have a lot of files and an older backup."
here's my backblaze folder: http://imgur.com/ohqUERr (i have my temporary data drive set to a secondary drive)
What kind of internet connection are you on? I recently reinstalled my mac from scratch (El Capitan upgrade went haywire) and had to redo my backblaze backup. I uploaded the 500 GB to Backblaze in under 24 hours.
Also, could it be that the "re-upload" was able to detect your older copy of the files and reuse them? I can rsync terabytes in minutes over slow connections when the other side already has the files ...