Backblaze is now a Terraform provider
backblaze.com
backblaze.com
First, B2's pricing is pretty amazing, especially compared to S3 and similar competitors: https://www.backblaze.com/b2/cloud-storage-pricing.html
Second, be aware those savings come with some downsides. The major one for us has been their maintenance window every Thursday from 2:00-3:00pm Pacific Time. Usually there's no outage, but sometimes there is. There's no warning - it's just down sometimes during that window. So, if uptime is important for your data, consider the cost of also implementing a fallback solution to cover your production use during those maintenance windows. https://www.backblaze.com/scheduled-maintenance.html
I mainly only use them for backups instead of production data
EDIT: As of 2019 there are multiple datacenters, but it doesn't seem like the data is stored redundantly across them
Durability refers to the idea that your data will still be retrievable (ie no corruption) similar to S3's claim of eleven 9s of durability.
Reliability however would be say; the actual availability of the Web UI or API server that you download your data through. If it were down, that wouldn't impact the actual integrity of the data itself.
For anyone interested, any more than about five 9s of reliability is basically impossible anyway when it comes to human intervention. As an example, 6 nines of reliability would allow you 3 seconds of unavailability a year so 250 milliseconds of unavailability a month.
From a users point of view, being "unavailable" includes everything from going through a tunnel and having your mobile connection drop out to a shark biting one of the undersea cables in the middle of the ocean.
As you might imagine, with a human involved, they couldn't even get acknowledge an alert fast enough to meet that deadline let alone actually going about doing any repairs and diagnosis :)
It could be spread over multiple instances and redundant hardware as well but as with any system being touched by humans, it's near guaranteed that something will go wrong eventually.
This helps, too, as it lets us build out services to be more reliable in combination, rather than less reliable. With retries and fail-over, an outage in an entire region may not necessarily result in any user requests failing.
For scale, pre-pandemic our published figures claimed >100M MAU.
I find https://andrewaylett.github.io/multi-burn-rate-calculator/ helpful for visualising error rates -- largely cribbed from the project it's forked from :) but with the tweakables switched around and the time between alert and error budget exhaustion in the tooltip.
It's worth noting that we only evaluate our alerts at most once a minute.
Said Oracle. Then their DNS was misconfigured and their whole cloud went offline for 2 hours [1].
Shit happens, always, at all scales.
[1] https://ocistatus.oraclecloud.com/incidents/qjxllgkywysj
Edit: typo
But perhaps my point wasn't clearly enough made: a claim of "100% uptime" on a service level isn't particularly _useful_ when our users still only see a 99.9% success rate.
Which is to say, individual subscriber lines can have all sorts of faults, but they only affect that subscriber. Trunks between offices can fail, but they only affect a certain number of circuits. The central call processing ability, of the switch to react appropriately to lines changing state and numbers being dialed, for calls to be completed if both stations are available, is what had to meet that target.
I was skeptical when I first heard this, so next time I was in an office that still had a #1 (actually a 1A, it had been upgraded in the late 80s), I talked to the switchman for a bit, and he showed me the downtime counter. It's a mechanical thing like an odometer, and once a second when the switch is executing its main loop, it touches a register that keeps the counter from incrementing. If the processor halts or isn't processing calls for some reason, the counter starts counting.
The switch was installed in the mid-70s, and from the moment it took over for its crossbar predecessor, including an in-place processor upgrade, it had logged less than an hour. Most of that came a few seconds at a time when swapping between active processors during software upgrades, he said. At the time (this was around 2002 I think) it was slated to be replaced with a DMS100, but the replacement activity hadn't commenced yet. I don't know what sort of reliability numbers the DMS machines achieve, but they'd do well to match their predecessors.
If you get better, awesome, but SLA? Too unrealistic.
What's the SLA penalty vs the extra the customer is willing to pay? If I think I can achieve 6 x 9's on a monthly basis, but probably I'll only achieve it 10 months out of 12, I can offer the customer an SLA of 6 x 9's for 100 USD per month.
My penalty can be 50 USD for failing to meet, and then I as a supplier walk away with 10 x 10 + 2 x 50 USD = 1100 USD for offering something I knew I couldn't achieve (consistently).
Not sure what the qualifier “when it comes to human intervention” means, but if I ignore it then five nines is quite standard in certain sectors. For example phone switch SLA (back when I was in the biz) was measured in minutes per decade (as in “cumulative unavailability under 1/2 minutes per decade”). Large baseline power plants can and must run uninterrupted for decades.
Of course it’s a systems issue not a point solution.
I realise in hindsight that I was rambling from the point of view of offering five nines for software that is inherently flaky/unreliable. Companies where developers cycle through and knowledge is lost, technical debt accumulates and systems are used almost counter to their intended purpose (ie Redis as a database)
In that sense, it's like constantly massaging applications to stay alive or at a reasonable level of service and so hence the assumption that someone will be paged and respond in order to preempt a failure or restore service during an outage.
That's why I feel more comfortable storing data across two "unreliable" providers with a lot of physical space in between them rather than one super reliable provider.
You also have to consider that data loss can result of simple things such as an account that gets blocked for stupid reasons. If you want to be safe you always need to have your data with at least 2 providers.
This pretty much prevents me from using B2 for now.
[1]: https://www.backblaze.com/blog/cloud-storage-durability-vs-a...
Then they relax.
If they're in offices, they group up, maybe in the break room. A rousing game of ping-pong breaks out.
Remote coworkers ping each other on Slack. Maybe a few start a round of Among Us. Bread dough is kneaded. Kids get a little more help with their schoolwork.
Everybody takes a very long lunch.
By 2pm, people realize the entire day is gone. Almost everybody has left or signed off by 3. Some roll out to bars; others go home to their kids, or to their gardens or garages or battlestations. Everybody beats the traffic.
Come Monday, the system is fixed. People are a little stressed out, since there's so much catching-up to do, status reports to be filed, widgets to be tracked and poked. But everybody agrees that was an amazing couple of days, and they got lots of rest, and it sure was nice. And hey, I had this great idea over the weekend—
That seems way cheaper than this.
For backups and large, long term storage, AWS has Glacier, that's really really cheap.
Thank you, I was not aware of this policy.
[1]: https://www.cloudflare.com/bandwidth-alliance/backblaze/
Especially because Cloudflare's pricing is "smoother" and detached from any one service.
In the end, S3 can be cheaper but you have to make a lot of assumptions beforehand. Backblaze is cheap enough to just throw everything in there and work with their lifecycle rules. You don't need to make assumptions about download volumes or storage duration beforehand (esp if you can retrieve via cloudflare).
I’m actually a huge fan of Bunny now. The CDN piece is about as cheap as it gets (for any utility based service), it’s optimizer and other things work well, and it works seamlessly with their storage system too. Which is super cheap itself, allows you to control how much it’s replicated (and where) - just waiting for them to deliver S3 compatibility so all the existing tools that exist work, or some other type of CLI tool.
For Digital Ocean, please look that their pricing is higher in both bandwidth & storage.
As a customer is there any way to opt-into a more proactive notification of an anticipated delay, like an email? I understand such things are necessary sometimes, but "always pay attention to some blog or twitter for a rare occurrence" doesn't seem particularly busy-stressed-admin friendly :).
Otherwise, the maintenance window becomes 24hx365, since "ensuring downtime is avoided during maintenance window" means literally - make a maintenance window have the same uptime as non-maintenance window.
Obviously this was a few years ago, but a backup provider failing at their one job and then blaming the customer left a really bad impression that keeps me from using them.
Backblaze eventually admitted that their dashboards aren't realtime, and they had a bug which was showing us (and their client) files that didn't exist.
Maybe it needs a kind of stochastic automated approach.. a program that finds sufficiently small (vs costs) sample of files on your computer (some old, some recently changed, etc) and tries restoring them and verifies.
At least for the standard Backblaze service you can download for free. For a USB drive you float the cost of the drive (they reimburse you when you return it)--maybe you pay shipping?
To restore it to a drive all you pay out of pocket is return shipping of the drive. The one time I had to use it I was slightly over the 30 days (I was waiting on a repair before I could restore the data) and it wasn't an issue.
I did some medium-intensity benchmarking a while back and decided not to put certain server data on it because I was getting a few 20+ second timeouts per thousand read requests. I can handle server errors, and I have retry logic, but this was something where I needed to be able to access the data within a second or two. Maybe it would have worked better if I set a very aggressive timeout, I'm not sure. Deeper testing is something I'll worry about some other time if the data actually grows past a couple hundred gigabytes.
This was mostly with the S3 API, I don't remember if I ever succeeded in getting the program to use the native one.
(The exception being the recent outages GoDaddy caused for them, but since they've moved to using Cloudflare as their registrar, I don't anticipate further issues there: https://news.ycombinator.com/item?id=26119619 )
But my true dream would be for backblaze to someday offer ZFS as a service.
I want to `zfs send -i my_pool@2020-03-11 | b2zfs recv some_bucket_id`, then be able to view my snapshots and files within in the backblaze web UI, and restore with `b2zfs send -i some_bucket_id/my_pool@2020-03-11 | zfs recv my_pool`.
You can already mount B2 as a FUSE filesystem with something like ExpanDrive, then write ZFS raw file vdevs to the B2 FUSE mount, but it's horrifically slow and probably too janky for any real use.
EDIT: As mentioned below this is for the most expensive (lowest capacity) tier. I’d been comparing for my own home use and so I would be unlikely to exceed 10TB but if you’re looking at higher capacity then maybe the calculus is different.
For large quantities, it is 3x, and actually less since there are no charges for ingress/egress.
Further questions/comments over email, please, since this is BBs HN thread and I don't want to butt in.
"Backblaze now has a Terraform provider" or "Backblaze released a Terraform provider" makes more sense to me.
Using the S3 backend with B2 should work fine since B2 makes a S3 compatible API available, but its made more difficult because the "Amazon Provider Team" is the one that maintains the "s3" backend for Terraform, and they want to do additional validation that matches AWS's expectations.
Do you use a similar technique to https://poweruser.blog/embedding-python-in-go-338c0399f3d5 to embed the Python SDK?
Also, was there a reason it's not against the B2 API? Not a judgment, just curious about the design tradeoffs in a professional-talking-to-professional sense.
And appropriately linked?
I can see why they put it on a separate page to not clutter up the article.
[1] https://help.backblaze.com/hc/en-us/articles/1260803375989
I was less interested in the code sample as a "how-to" and more for skimmability – rather than reading a bunch of words I just wanted to see what it would look like.
Eg, from that link:
terraform {
required_version = ">= 0.13"
required_providers {
b2 = {
source = "Backblaze/b2"
version = "~> 0.2"
}
}
}
provider "b2" {
}
resource "b2_application_key" "example" {
key_name = "test-b2-tfp-0000000000000000000"
capabilities = ["readFiles"]
}
data "b2_application_key" "example" {
key_name = b2_application_key.example.key_name
}
output "application_key" {
value = data.b2_application_key.example
}Maybe it was in the beginning, but Terraform is far more powerful than that now. Terraform is a monad that neatly separates pure declarative configuration from the I/O (side effects) that are factored out into providers. Terraform used at its most powerful is not limited to infrastructure, it also sets up the platforms and applications running on that infrastructure for you, by separating the configuration of the platforms and applications from the generic API calls that apply that configuration. Terraform's dependency graph ensures that the calls are made in the right order, no matter if they are made to infrastructure APIs, platform APIs, or APIs belonging to layers further up the stack
For the most part, I provision things with Terraform and then instantiate the servers with NixOS/Nix itself, and this mostly works. For bonus points you can use Nix to generate the HCL that Terraform reads in (because Nix can write JSON, and HCL is just JSON in a trenchcoat) if you want to put some veneer on it.