Backblaze B2 Cloud Storage Now Has S3 Compatible APIs
backblaze.com
backblaze.com
So you are only paying for Storage. ( Correct me if I am on this one )
I wonder why doesn't ALL non HyperScale Cloud Vendors, like Linode and DO provide one click third party backup to B2. You should always store offsite backup somewhere. And B2 is perfect.
Having a functional video player on your site or in your app (e.g. one where you can skip to arbitrary times without requiring the video be buffered up to that point; or where a video can be "resumed" from the middle if you leave it and come back) already requires that you use MPEG-DASH or HLS; which in turn implies/necessitates pre-chunking, no?
Is there some use-case where people are currently serving 400MB contiguous video files from a CDN? I can't think of one. YouTube doesn't. Netflix doesn't. Even porn sites don't.
I guess Archive.org has some large video files on there in various places, that can be direct-downloaded; but the recommendation Archive.org itself makes, is to consume those via BitTorrent. Presumably they don't have a CDN partner willing to handle their unique workload for cheap.
Browsers are smart. They only buffer a few megabytes at a time and can seek around pretty efficiently.
It's totally possible to "encode for streaming", but it usually results in both an increase in overhead [more keyframes] and a decrease in quality [inability to use predictive interpolation, instead relying only on forward-interpolation.]
Mind you, this streaming-enabled encoding is how things were done on the web, before the advent of MPEG-DASH/HLS; and it's still how e.g. the MP2 encoding of digital cable/satellite video works. But we don't really want to go back to those days. They kind of sucked.
Jumping to random byte offsets in a video also tends to screw with any embedded data streams like subtitles or thumbnails, which tend to just be stored in most media container formats as a single chunk at the beginning/end of the file, rather than being spread or copied across the stream. Again, the kind of captioning done back in the MP2 days is immune to this, but it kind of sucked as well (e.g. it wouldn't trigger if you happened to skip to the millisecond after the instruction for it appeared in the stream, often leaving you with ~30 seconds of untranslated audio.)
My impression is DASH/HLS are mostly useful for adjusting bitrate on the fly.
† I mean, a digital-cable video stream is always "in the middle" unless you're just starting a VOD stream, but still.
Do browsers even support embedded subtitles?
I think what you're saying applies more to a setting where the video is being streamed live, so that you cannot access the start of the file to get keyframe metadata. In that case HLS and MPEG-DASH help.
Serving a 400mb video via a CDN is highly dependent on the CDN. Some will construct a cache entry from a slop of range get requests and translate them to get the missing pieces and work brilliantly and other CDNs should be avoided.
2.8 Limitation on Serving Non-HTML Content The Service is offered primarily as a platform to cache and serve web pages and websites. Unless explicitly included as a part of a Paid Service purchased by you, you agree to use the Service solely for the purpose of serving web pages as viewed through a web browser or other functionally equivalent applications and rendering Hypertext Markup Language (HTML) or other functional equivalents. Use of the Service for serving video (unless purchased separately as a Paid Service) or a disproportionate percentage of pictures, audio files, or other non-HTML content, is prohibited.
Does that mean your video use-case would also be fine? I have no idea. An HN comment from the CEO doesn't seem like it would hold up if Cloudflare suddenly shut down your free account.
I'd love for Cloudfront to officially clarify the limits of the Cloudfront/B2 alliance in terms of external traffic. The confounding issue here is that B2, as a storage service, is not really intended for "serving web pages and websites" — it's for larger files, binaries, etc. — and therefore any traffic from B2 going through Cloudfront is sort of de facto in violation of 2.8.
> B2, as a storage service, is not really intended for "serving web pages and websites" — it's for larger files, binaries, etc
It might be missing a couple features (which is a pet peeve of mine) but we SURELY intend for it to be used for serving web pages. That's one of the largest differences between "Backblaze Personal Backup" (our original product line) and Backblaze B2. The largest parts of the redesign/refit when we originally did B2 was around the concept of what we call "Friendly URLs" (web page names, folder names) instead of just ugly 82 character hexadecimal file names like Backblaze Personal Backup stores all your files in.
For full disclosure, Backblaze B2 isn't a great "hosting" solution for something like WordPress because we lack two or three things, one of which is comically easy to fix and I keep trying to convince everyone to do it. The issue is URLs that end in a "/" (trailing slash) basically need to "guess" that after that is an ".html" or ".php" or whatever. So the URL: https://f001.backblazeb2.com/file/ski-epic-c/full/2015_scotl... does not work, but the URL: https://f001.backblazeb2.com/file/ski-epic-c/full/2015_scotl... does work. All modern web servers do this automatically filling in of the "index.html", but it is missing from Backblaze B2 currently. And it would take just a day or two for one of our developers to fix it. And dang it, I'm going to get it done one of these days.
That said, I would absolutely not consider using B2 without support for index.html, error.html, and Website-Redirect-Location.
That covers 10 million requests, which could each be 10MB chunks, or 512MB cached files, or possibly larger.
(A raw mp4 needs about 3 requests to start plus one per seek.)
But I wouldn't be surprised if that doesn't scale to enormous amounts of data.
Another big asterisk is that this applies to all proxied (DDOS-protected) content, not just the ones that use the CF cache. CF pays for all of their uplink bandwidth out-of-pocket regardless of if it's cached. You can see what happens when you proxy multiple terabytes of content on the free plan in this thread[0] (again all proxied bandwidth costs CF, which is why this user had their zone unproxied).
0: https://community.cloudflare.com/t/the-way-you-handle-bandwi...
Every time we've tried this with other "free" providers there always seems to be "fine print" when you start pushing huge bandwidth on these "unlimited" plans. It seriously is not worth the time at some point.
Storage Download
BackBlaze .5 cents/GB 1 cent/GB
S3 (one zone infrequent access) 1 cent/GB 9 cents/GB
I'll transfer my 250GB of videos and images from S3 to BackBlaze some day.if I use against-the-alliance corp, i will just pay the storage+download from backblaze, just the same as using in-bed-with-alliance corp.
Or does the alliance will ensure i only get charged for 1download per file per billing cycle, no matter if the CDN cache can't hold all the data and they download it many times? Does the alliance ensure (and more importantly, who monitors) it?
I think Cloudflare already was against net neutrality, but people who believe in that principle might need to avoid Backblaze as well if that is the case.
This is about the fact that bandwidth is usually bought by capacity between networks rather than usage and they're passing those savings on to you instead of charging expensive per-GB costs like the major clouds.
If the argument is a semantic one, that net neutrality is so narrowly defined that it doesn't apply to them as a CDN rather than a conventional ISP, it still stands that the situation is not-net-neutral if we generalize that to include CDNs, which it ought to, since that's what the N stands for.
disclaimer: ex-cloudflare
The core concept, in my mind, is that all packets should be treated equally, and should cost the customer the same amount per byte.
The problem is that different packets take different network paths and have different costs to your provider. The bandwidth alliance partners are exchanging traffic via peering links, so there's only the fixed costs of equipment and any port charges/interconnection fees to the peering facility. Packets sent to other destinations may pass through a paid transit link and cost the provider. If your provider charges you the same amount for both types of packets, they're not being transparent about their network costs which is bad for you, and is also bad for them, because you may be able to adjust your traffic so more is settlement free and less is paid --- but you won't do that without an incentive.
Of course, if you don't like how Backblaze manages its network, there's a bunch of other network storage vendors who you could switch to (many of which are members of the Bandwidth Alliance). Or, it's not too hard to build a storage box or two and get them into a colocation somewhere.
For your residential ISP though, for most people, if you don't like the network policies, you might have another option, but they will likely have similar policies. You don't have meaningful choice in a single residence, and moving residences to get better choices isn't a meaningful option either. So, network neutrality is a policy to (attempt to) regulate residential ISP behavior, in order to provide a reasonable policy for users. It's still a problem that it doesn't match reality, and it doesn't provide user choice, but it has gotten a lot of support.
I would rather see mandatory line sharing for residential ISPs, it's a lot easier to define and with proper regulation offers a path towards consumer choice, and gives people a way to take action if their provider has poor network policies --- if your provider runs an acceptable last mile service, but provides poor interconnection to the rest of the world, you could become a line sharing partner and provide better interconnection to the world without having to build out an overlay last-mile network (which is incredibly capital intensive and difficult)
Reword it to be generic, and dress it up in more polite language, but that's what it is.
Comcast, nor any other ISP, doesn't get to shake down Netflix or Google or Facebook nor any service providers, big or small, on the Internet.
Sure, behind the scenes there are backbone providers and peering arrangements and BGP, but those are technical details. Network neutrality means as a consumer I pay an ISP and get connected to the Internet.
Whereas Comcast is a cable TV company, the game between companies of consumer chicken is not theoretical, eg
https://www.thedenverchannel.com/news/local-news/why-you-sti...
https://mynbc15.com/news/local/att-and-directv-have-removed-...
https://www.washingtonpost.com/business/cbs-blackout-on-dire...
I agree, that it's bullshit, but please consider:
A telephone company charges me to have a phone line and per minute to place and receive calls, and charges whoemever I'm conversing with to have a phone line, and to call me.
If Comcast wasn't large enough for Netflix to want peering, Netflix would probably use paid transit to get there (they might even use paid transit that Comcast would also have to pay for).
So, the reason it's bullshit can't be because Netflix has to pay for a service that Comcast is providing; because paying for network services is the norm. The reason also can't be because Comcast is charging both ends of the connection for the service it is providing to both, namely connecting the two parties; because both sides paying for connectivity is the norm in telecommunications.
There's some other reason it's bullshit. A big factor, of course, is that internet video services compete with the ISPs video services, and there's some implicit unfairness there. I would say it's because the residential customer is stuck with a very limited set of choices, and can't pick an ISP that distinguishes itself with open peering. It might be because the norm for internet peering is if there is a significant imbalance in traffic, the party sending the most pays; so the residential ISPs should be paying their residential customers; this one doesn't quite work, because it's not actually a peering relationship, and the norm is transit customers pay the transit provider for traffic in either direction.
> Sure, behind the scenes there are backbone providers and peering arrangements and BGP, but those are technical details. Network neutrality means as a consumer I pay an ISP and get connected to the Internet.
The problem here is there is no "the Internet" to connect to. An ISP needs to connect to substantially all of the other networks; and while a small ISP may simply use a single upstream ISP, where it would be easy for the small ISP to be neutral (bits in = bits out), a larger ISP is going to have a diversity of connections, and that's going to lead to defacto non-neutrality as some connections are bigger than others, and some connections are longer than others, and some connections are more expensive than others.
I fully understand the desire to restrain residential ISPs (and mobile ISPs) from anti-competitive behavior; I guess network neutrality might work for that.
But, when you apply it to other networks, it doesn't seem to make sense. In this case, you have a group of networks that would like to lower costs, both for themselves, and their customers, and they've agreed to settlement free peering, and to also not charge customers for bandwidth on the settlement free links. Network neutrality doesn't really speak towards the interconnection, but says that customers should be charged the same rate for all bandwidth, even when the underlying costs are different. I don't understand why that's a good thing in this case?
If a neighbor got from a rival DSL provider, your Internet was liable to go out until your ISP was able to get a truck out to fix the damage to your connection that the competition's technicians did.
Those laws are still around, but unfortunately, (A)DSL tops out at single-digit megabits. An upgrade from dial-up, but not competitive with todays broadband market for those with other options.
Pricing was a big issue in the past --- where the 'wholesale' per-line tariff charged to the competitive carrier was more than the retail price charged to incumbent customers.
I think, in most cases, new install breaks other user types of issues are related to poor records of which line is used by which customer; that happens within the incumbent as well though (when I got ADSL2 installed, one of my neighbor's connections went out, and pretty soon there were three AT&T trucks on the street to work everything out). That particular issue could probably be solved by allowing the incumbent to manage all the installs and monitoring for install-time customer steering.
The federal mandate no longer covers line sharing, it only covers line leasing for copper telephone service; wherever your premises is directly connected to, if there's space for competitive equipment, the incumbent must make it available. Of course, most telephone companies have moved customers to remote terminals for better speeds, and there's no room for competitive equipment in the remote terminals; and a lot of telephone companies are replacing copper networks with optical networks, and those aren't covered either. The whole concept got submarined by the insistence that it apply to telephone networks and not cable or other "new technology" networks, and then deciding not to apply it anywhere in light of court cases that applying to one and not the other was unfair.
The charges are from the cloud (storage) providers. And most of them charge for bandwidth usage, but they don't discriminate based on what data you send. It doesn't matter if it's source code, photos, or linux binaries so there's no net-neutrality issue here.
You can potentially argue that Backblaze is discounting traffic based on the upstream network but that's not active discrimination. There are hard costs to internet transit and some companies can offer better pricing depending on where the traffic goes. For example, downloading from Australia is more expensive than downloading in America because of the infrastructure costs, and this isn't considered a net-neutrality violation.
It is just a property of networks. As an end user you still get to use all providers, but two of them may use a dedicated connection to reduce internal prices (the price benefit of which they may or may not give back to the user).
Neutrality from the point of the user is not affected when it comes to service being available.
Now if one airline alliance said you are not allowed to get on their flight if you hoped off from a rival alliance, then we have an issue.
Cloudflare dont charge per GB. They are more like per "feature" CDN.
But this is also a sad realisation that storage tech aren't getting cost reduction any further.
( Of course Blackblaze can always prove me wrong :D )
So may be it is cheaper, but not really comparable.
It was.. sort of cheaper. They didn't actually build the servers, and as described the server wouldn't work (onboard SATA didn't support port multipliers, lack of ECC would probably cause problems in practice, bit hand-wavey on power/space/network/manpower costs, etc). The goal of the article was to get other people to build cheap storage and put it up for rent on their network. They do have some amount of storage space available for very cheap on the network now, but personally I suspect it's people who figured "what the heck, I'll give it a try!" as opposed to people actually building storage servers and making a profit renting them out.
I was honestly pretty disappointed - I'd hoped they'd found a cheap motherboard with ECC and support for port multipliers, but nope.
Edit: The motherboard they picked does support ECC memory, anyway. In general ASRock models do.
You're right though - since the client's doing the work and they have a lot of redundancy/diversity in the storage it's not as big of a deal for them as it would be for us. I'd be a bit wary because the client-only verification does mean that there's no verification-with-ECC step in the entire chain, but I'm not sure that's significantly worse in terms of actual risk.
Glad to see they've since fixed that, and with this update are clearly continuing to improve ergonomics. I'll have to give B2 a fresh look.
[1]: https://github.com/Backblaze/B2_Command_Line_Tool/issues/175
On the other hand tools that don't view the object storage as a file system have way less gotchas.
In my experience B2 + restic works really well.
I used to be able to trigger my iPhone -> ZFS script when I plugged my iPhone into my Ubuntu desktop using udev (also had to wrap it in flock[1] because it would trigger multiple times for some reason), but at some point that stopped working and I've been too lazy to figure out why.
It's far from perfect but for me it works alright. In this scenario I prefer straightforward and slightly kludgy compared to something with hidden complexity that could go wrong in so many ways. Could you imagine if you used a tool like restic or borg and the pack encoding format changed, or if the tool sources are simply gone when your relatives have to figure out how to get at the files in 10, 15, 20 years - I don't want my relatives playing code detective or archaeologist!
Which reminds me of a downside to using tools like restic, borg, and the like I forgot to mention. When I evaluated them for my hundreds of GB of family pics+videos, there is a "dedupe" step that all these tools want to perform. When I tested them a couple years ago they were dog slow for my files, because pics + video are already highly compressed and there is very little "deduping" you're going to wring out of them unless you have multiple copies of the same files. IIRC borg took several hours to run at at the end it reported 0.01% or less deduping efficiency. Also as I recall there was no way to opt-out of the dedupe step due to the way borg stores "packs". Very annoying!
You'd need a cloud storage provider that just gave you a plain old UNIX filesystem to do whatever you want with.
It's too bad nobody does that ...
Doesn't iCloud Drive fit the bill?
Access to it is slightly obscure, i.e. ~/Library/Mobile Documents/com~apple~CloudDocs/ but wouldn't that work?
It's free for me to use. That's because I'm already paying Apple $10/mo to backup the family's iPhones. We're only using a little over 200 GB out of the 2000 GB we have. (I'm sure that Apple is counting on most people not using their full amount).
I've only put a few files out there, so maybe there are a lot of potential pitfalls. But it doesn't get much simpler than using cp or mv.
In reality it's most emphatically not a "plain old UNIX filesystem". Apple is doing some magic and storing blobs out in Amazon S3 or in their own datacenters. But to me it has the appearance of a Unix (Posix?) filesystem.
I realize that rsync.net couldn't survive with a business model that limits users to 2000 GB, which is Apple's maximum. But I thought I'd mention it, since it just might be the perfect "free" solution for a lot of people.
Same for archiving. At my former workplace we however figured that out only after the horse left the barn.
Disclaimer: Happy Backblaze Mac client and B2 customer, no other affiliation.
EDIT: @yev: I took the signal out after the sibling reply :) Appreciate the responses as always. Please stay awesome.
*Edit -> Well, those plus the load balancing servers =D
> So how did they manage to get rid of those hidden costs? Or is the new S3 compatible API more expensive?
The new S3 compatible APIs are the same cost as the original native B2 APIs.
I'm the author of that original blog post, and we were able to get rid of SOME of the internal costs, but in the end we ate the cost of the load balancer. Internally I voted that we externalize that cost, but it's in the spirit of our "no-sales-friction" pricing model to get rid of decision points for customers and just let them get their stuff done.
The internal math is that we believe the additional storage that (hopefully) will result from supporting the S3 API will help make up the cost of the upload balancers. It eats into our margin, but not by enough to justify a "friction pain point" that might reduce sales as customers struggle to decide which API to use. When you do the math over thousands of customers, we make the vast majority of money simply from renting the storage. Many many many customers upload things to Backblaze, and let them sit for long enough, that the cost of the load balancers becomes pretty small as a percentage.
One of our goals of the pricing of Backblaze Personal Backup (flat fee of $6/month regardless of how much data you backup) and not charging much for transactions is we aren't trying to nickle and dime customers. We just want to make our margin and provide a solid service that makes customers happy.
In any case, thanks for B2 (happy customer) and good luck. Sounds like an exciting time.
Now that Amazon has deprecated the single URL version and replaced it with region-specific URLs (e.g. s3.dualstack.us-east-1.amazonaws.com) and tooling has been mostly updated, this huge reason for not supporting the S3 API is gone.
Heh, kinda like how FTP worked. That's funny to see again.
Amusingly the price to store 1.2TB of data is the same as the cost of their backup plan, so if your disk is smaller than that, you could save a few bucks running your own backups. Until you have to restore (from what I can tell restores are free on their backup plans but would cost money on the S3 plan).
You could, but if i read correctly (s3fs-fuse limitations): "random writes or appends to files require rewriting the entire file".
So changing 1 bit of a 10GB file, means re-uploading 10GB.
> random writes or appends to files require rewriting the entire object, optimized with multi-part upload copy
Now changing one bit means re-uploading 5 MB, the minimum S3 part size.
That being said, I wish B2 performance was better. Throughput is dramatically slower than S3.
- Thankful Personal Backup Customer
I find it bizarre how in India you can get 100GB of LTE for a few dollars but cdn bandwidth can cost content providers more than that - which is absurd.
Already 4 networks have exited the market and 3rd & 4th largest networks(Vodafone and Idea) have combined due to cash crunch. Airtel(earlier largest) has been raising outside money in hopes that it can survive the low prices. So there are only 4 networks remaining. Only recently they started increasing prices.
That billionaire is also going into Fibre(purchased his bankrupt bother company's infrastructure), maybe we'll see that competition extend to DC and interconnects.
When I contacted their customer supported , I was pointed to the following URL where they explain in detail how they handle these errors https://www.backblaze.com/blog/b2-503-500-server-error/
Essentially they are not considered as errors and expect the client to retry loading the file. This approach won't work in our use case.
Even if you're doing multi region s3 replication you'll run into this for external clients semi-occasionally.
but when we get a 503 response it should be considered an an error and acknowledged by the provider that it's an error.
in my use case, we were using a CDN which was configured to pull files from B2. When B2 responds with a 503/500 I have no control on the retry mechanism.
The error rate was around 5-10%
I've been using B2 for backup storage for some personal projects. It doesn't necessarily do anything "better" than S3 from what I've seen, but never having to log into AWS's dashboard is a reward enough on its own.
They do have a command-line client that's a quick PIP install, so you can do something like:
b2 upload-file bucket-name /path/to/file remote-filename
Which is, of course, nice for backups.Same since it seems to be on the new storage service launch checklist right after "buy hard drives".
Currently I have to save the tar.bz archive to disk first before uploading to balance. Took several hours to do so (huge spinning disks array, not as fast as ssd), while uploading to B2 is blazing fast. Saving the archive to ramdrive essentially solved this, but as the data grows I don't have enough memory to spare anymore for a ram drive that can fit the whole archive.
b2 upload_file bucket <(tar -cj huge-directory) archive.tar.bz2
The argument the command sees will be something like "/dev/fd/42", and the shell will provide the output of tar through that file.Just had to be sure to omit the B2 external storage folder from the backups on my Nextcloud server.
Now only if Virtualmin (YC ‘08) supported virtual server backups to S3-compatible B2 cloud storage… There’s an open ticket for this at https://www.virtualmin.com/node/65024
It's like brand names that become so common that people use no more the material name, but the brand.
Edit: Confused Hetzner and OVH, sorry
It is... a little slower than I'd like but with Cloudfront in front it has been manageable. I love tips from Backblaze on how to increase performance there beyond caching to CF.
BIG gotcha btw: You have to choose between US and EU when you create your B2 account! You can't have buckets in the other location, so that means you'll need two accounts if you want to do that.
The killer feature of Google Cloud Storage in my eyes is its ability to be strongly consistent, if you set the right HTTP headers. This is not possible for Amazon S3, which is always eventually consistent and makes it unusable for many use cases where you need to be able to guarantee that customers will always see the newest version of a file.
Yes - B2 is strongly consistent. When you upload an object using either the B2 Native or S3 API - the object is persisted to the final resting place before the upload completes. Therefore, you can list/download the file immediately after your upload completes.
Glacier pricing in us-east is .0004 vs .0005 for B2. There is always pricing obfuscation with cloud, but AFAICT, there is no need to move off Glacier for a backup use-case.
Amazon wants to charge you $90/TB* to get data to the outside world, compared to B2's $10, but you can mitigate it in various ways. At the low end that's using a lightsail instance as a VPN, depending whether you think the TOS allows that. At the high end it's paying flexify.io $40 to move your data to B2, then paying B2 $10.
There might be other ways to improve S3 egress costs. It's a very hard thing to search for. I only learned about flexify from this post.
So if you have to restore less than half of your data each year, Glacier Deep storage will save you money. It's worth considering, unlike normal Glacier which is almost entirely downside.
* There's also a $2.50/TB fee to get things out of Glacier, but that's dwarfed by the other costs.
I use this as an offsite backup- that as long as disaster does not strike, I will never use, and even if does, I can be patient about restoring.
This is for home use- which maybe I was wrongly assuming that most Synology users are. In a business context you might have a more routine need to recover backups and those costs become more tangible.
Synology Hyper Backup absolutely works. Details are here: https://help.backblaze.com/hc/en-us/articles/360047171594
(If you saw my old answer, ignore it. I was misinformed.)
I literally have NEVER seen SFTP being used for blob storage in any python project - is this a real thing somewhere?
It seems to me you would only get a very narrow subset of the functionality.
having access to primary storage and cheap backup storage using the same S3 API will make us reconsider that and will probably make it worth the effort to dump our rsync-based solution for B2.
When they first launched B2, I inquired about ability to enter into a BAA (Business Associates Agreement) for HIPAA compliance and was told that it wasn't "on the roadmap". It sounds like B2 has come a long way on the compliance side. Would be great if they were open to this.
It was just super hard to make the code perform well. Like you have to manage client sessions on your side and chose optimizations on your side. Like you have to spread things manually. Which is hard to do. Whereas S3 is maximising your bandwidth with no custom code required. It's not really S3 compatibility that was needed but B2 API wasn't good.
There is a price/performance trade-off: B2 has higher request latency than S3, no matter where you are (my experience), but they also are 5x cheaper on storage costs, 10x cheaper on download bandwidth, and have no price gimmicks like minimum object sizes or minimum object lifetimes like many other services (S3 IA for example).
To make up for B2's request latency it is more important to issue requests from multiple threads, especially for short-running requests like removing files.
Another key difference is that B2 always uses SSL whereas S3 can be accessed without SSL with little security impact because each S3 request is individually signed with a secret key. Setting up an SSL connection is more overhead, so another key to performance is to reuse connections.
Both of these suggestions apply to S3 as well, just more to B2 because of the latency difference.
I am not a lawyer. So this is a genuine / dumb question.
> Can Amazon actually patent their API (per the google vs oracle case) - basically like prevent other vendors to provide S3 APIs so that Amazon can lock in users.
Most likely yes. Backblaze plans going forward are to fully, uncompromisingly maintain our original native B2 APIs for a few reasons including this concern.
It's probably up to Amazon whether they want to boot all 3rd parties off their S3 API. Backblaze has a viable fallback if that occurs. I hope for customer's sake Amazon doesn't declare war in that fashion.
If Amazon decides on this path, internally at Backblaze we have discussed immediately doing the opposite - declaring for all of time anybody can copy our B2 APIs. Remember, our APIs are technically superior to the S3 APIs. They are lower cost to implement, and are shockingly easier to use for developers. They don't make all the mistakes S3 made. We had the luxury of learning from all their mistakes over the years. :-)
So google cloud is actually expected to also have this potential legal time bomb?
Amazon can sue you and retroactively force you to pay them right? So all they need to do is to wait for the alternatives to become popular
It is just awful to see, how everyone tries to reinvent the wheel and not to be compatible with anyone else.
I see this isn't available for old buckets, is there a straightforward way to duplicate a bucket to make it compatible or do you have to use something like rclone?
If their load balancer is smart enough it can call the dispatcher, and make use of something like https://zaiste.net/nginx_x_accel_header/ to figure out where to forward the request. Unfortunately this still requires uploads be proxied through the dispatcher.
You could get crazy and involve a CDN (akamai or cloudflare or fastly) that could do some smart logic, especially if you can emit your dispatcher as a lookup table that's updated frequently. I don't know what bandwidth costs would be for that though. Probably high.
It's an interesting problem space and I'd love to talk to these folks about it.
> The DNS resolver library of your client is allowed to cache the IP address for a given hostname for up to TTL
Not only that, but one mistake a lot of developers made early on was asking for a location to upload for every upload. That was NEVER the intention. In fact that annoys our servers also.
Developers are supposed to request a location to upload ONCE, and then upload to that location for hours, or even DAYS. Unless you have a bug in that software, it really shouldn't come anywhere close to being a high runner in DNS. We're talking 9 or 10 requests per day, at most, if you are unlucky. Feel free to reach out to our support if you aren't seeing that!
These things are not particularly difficult, but they require additional mind-space to accommodate. Most developers will just do the simplest thing that works, performance be damned. If the simplest path is slow, they'll just remember "B2 is slow", no matter how unfair that is.
That is what we did for the S3 protocol. It adds cost via a load balancer.
The whole original storage design was based on the fact that in our original product line (Backblaze Personal Backup) we owned both ends of the protocol - our servers on the back end, and our client on the customer laptop. We were able to eliminate all load balancers from our datacenter by being a little tiny bit more intelligence in the client application (maybe 50 lines of code). The client asks the central server where there is some free space. The server tells it. Then the client "hangs up" and calls the storage vault directly, no load balancer required! Then the client uploads as long as that storage vault does not fill up or crash. If the storage vault crashes, or is taken offline, or fills up all the spare space it has, the client is responsible to go back and ask the central server for a NEW location. This fault tolerance step in the client ENTIRELY eliminates load balancers! Normally you need an array of servers and a load balancer to accept uploads, because what if one of the array of servers crashed, had a bad power supply, or needed to update the OS? The load balancer "fixes that" for you by load balancing to another server. Pushing the intelligence down into the client saved us money. Nobody ever noticed or cared because our programmers could write the extra 50 lines of code, to save the $1 million worth of F5 load balancers (or whatever solution Amazon S3 has).
We based our original B2 api protocols on this cost savings and higher reliability, but it does push the 50 lines of code logic down to the client. It caused a lot of developers this extreme, extreme angst. They just couldn't imagine a world where their code had to handle upload failures and retries. They would ask us "how many retries should we try before we just fail to backup"? Should I try 2 retries, or 3 before the backup entirely fails and the customer loses data? Our client guys had a whole different approach, since it was a computer we just went ahead and tried FOREVER. Never endingly, until the end of time, in an automated fashion. A couple times a year one client gets unlucky and it requires several round trips before getting a vault to upload to, but who cares? It's a computer, it can retry forever. It never gets tired, never gives up.
But S3 never figured this out, and they require the one upload point have "high availability". It saves any app developers about 50 lines of code and a lot of angst, but then we (Backblaze) has to purchase a big expensive load balancer, or build our own. We mostly built our own.
Developers working with cloud storage APIs generally need to get used to the idea that not everything is going to work all of the time. Retries and proper status code/error handling are critical to making your application work properly in real-world conditions, and as "events" occur. Every major cloud storage provider has circumstances under which developers must retry to create reliable applications; Backblaze is no different. For GCS, we document truncated exponential backoff as the preferred strategy [1].
Google has its Global Service Load Balancer (GSLB) [2], which handles...let's just say an enormous amount of traffic. GSLB is just part of the ecosystem at Google.
It's hard to design a storage system that's "all things to all people"! There are a series of tradeoffs that need to be made. Backblaze optimizes for keeping storage costs as low as possible for large objects. There are other dimensions that customers are willing to pay for.
[1] https://cloud.google.com/storage/docs/exponential-backoff [2] https://landing.google.com/sre/sre-book/chapters/production-...
"Developers are supposed to" is a scary way to start a sentence. ;)
Edit: also think DigitalOcean Spaces and B2 might be better off merging together, or Spaces being a whitelabel B2 in disguise (both are part of BWA).
> I really want to see lightning fast response times and TTFB (Time To First Byte Served)
If a file is "cold" (nobody has requested it in the last 24 hours) then it needs to be reconstructed from the Backblaze Vaults and there is a little delay. After that, it should serve pretty fast for the following requests (off of a caching layer with SSDs).
In the end, Backblaze B2 is a good solution for some customers, and not ideal for others. If your application requires blinding speed, like sub 1 millisecond serve times, Backblaze B2 may not be perfect for you. But how often is that the case? Certainly not when fetching a web page, or storing a backup for a year, right? In those cases a small delay is FINE. This is an example web page served by Backblaze B2 here, how does it load for you? https://f001.backblazeb2.com/file/ski-epic-c/full/2015_scotl... Fast? Slow? How is it?
For comparison, my regular hosting provider serving the same web page here: https://www.ski-epic.com/2015_scotland_will_macdonald_birthd...
Personally I can't tell any difference. I still look silly in a kilt in both versions. :-)
> Second pain point is the number of retries needed for uploading a large batch of small files.
It really shouldn't take any retries, or geez, at VERY MOST something like less than 1% - why is that an issue? Software should handle the tiny failure rate. I'm honestly curious, we want to know why people aren't choosing our solution!!
I asked the engineers that work on that code, and they pulled a random sample from the logs (we time all of this) and said for files less than 1 MByte, it averaged around 250 milliseconds to reconstruct the file from the Backblaze Vault and get it onto the cache servers where it is then served up. In 95% of requests completed within 900 milliseconds, but there were a few up over 1 second (1.2 seconds was the highest they found). Those are live production numbers so it includes all the load on those Vaults.
A couple other notes just to add color. Any one Backblaze account is bound for life to what we call a "cluster", for example there is one cluster in Europe so all files are stored in Europe for any account in Europe. There is a load balanced array of "cache servers" in front of all the vaults specific to that cluster (the caching servers are physically located close to the vaults for latency reasons), and our biggest cluster has something like 20 of these SSD based caching servers. Ok, so the cache layer is not "shared", meaning each cache server only pulls directly from the Backblaze Vault. So if you were serving a file, and 20 separate customers got amazingly unlucky, the file would get the 250 millisecond lag every time for those first 20 fetches. The cool parts of this architecture is that then you have 20 populated caches that are completely unrelated to each other so you have 20x the bandwidth available to serve it up (and a rack of really fast 20 servers to serve it). Plus they are all totally independent so they can crash or be brought offline to upgrade the software without any downtime.
We can add these cache machines as we need them, they are these 1U units and we have "warm spares" for a variety of things. When we have had spikes in load in the past we toss some hardware at it pretty fast.
We upload a few hundred GiB to B2 daily and have this issue as well. Really annoying…
If you want to sync to B2 specifically with a lightweight tool, check out https://rclone.org/
You can configure any cloud storage backend (B2, S3, GCS ...) and combine it with other utility storage backends, like "crypt" [1], "cache" [2] and "chunker" [3], I highly recommend it to anyone searching for a backup solution.
The only feature I miss from Rclone is automatic directory monitoring and mirroring, which I solved using Syncthing (but forces me to host an additional server for it).
[1] https://rclone.org/crypt/ [2] https://rclone.org/cache/ [3] https://rclone.org/chunker/
Have been restoring ponctual data easily. Did not do any major restore yet (hopefully).
I found them really good and cheap. I can only recommend.
https://fedoramagazine.org/automate-backups-with-restic-and-...
edit: one gotcha is that using "user" systemd units, they will only run when the user is logged in. So, for a personal device like a laptop, this is fine, but for a server, it might not do what you're thinking. For the server use case, you probably want to enable linger for that utility user, so that units will run even with no active logged in sessions:
# loginctl enable-linger username
Here are my systemd files for reference.
Service:
[Unit]
Wants=network-online.target
After=network-online.target
[Service]
Type=simple
ExecStart=/bin/bash -c "/usr/bin/restic unlock && /usr/bin/restic backup --verbose --one-file-system --exclude-caches --exclude=$XDG_CACHE_HOME %h && /usr/bin/restic forget --keep-last 5 --prune && /usr/bin/restic cache --cleanup --max-age 30"
[Install]
WantedBy=default.target
Environment (fill these in with your keys): [Service]
Environment="B2_ACCOUNT_ID="
Environment="B2_ACCOUNT_KEY="
Environment="RESTIC_REPOSITORY="
Environment="RESTIC_PASSWORD="
Timer: [Timer]
OnCalendar=*-*-* 00:30:00
RandomizedDelaySec=20min
[Install]
WantedBy=timers.targetI'm considering ditching the VM and using ZFS on linux, and just doing straight rclone to both the USB disk and Backblaze, for the following reasons:
TL;DR - my thread model is essentially me accidentally deleting something, and bitrot, and I think this setup is overkill for that.
* All operations mentioned take forever. Cutting the encryption I think might speed things up a lot (not sure about this, the key question is if the encryption is causing a lot of extra io operations).
* I'm afraid of losing my restic encryption key. I have multiple copies of my keepass file but if syncthing decides to delete them, both of my backups become useless.
* Backblaze has earned my trust, and in the unlikely event someone hacks my data, there's not much valuable there. I would probably switch to encrypting a single directory with the few things I care about being exposed. Even if I keep encrypting the whole thing, I would use something more boring than restic.
* I'm not sure I've ever needed to restore an old version of a file with restic (there may have been one time early on). I think deduplication is overkill for my needs as well. It's cool tech, I just don't think I need the complexity.
They've been doing incredible work in the open (storage server design, hardware reliability data, etc) and I'm really happy they've grown to where they are today.
Not sure how unique that URL is, looking at the structure it could depend what data centre your bucket gets created in.
[1]: https://community.cloudflare.com/t/cloudflare-how-not-to-vio...
> There is no mention of the durability guarantees that s3 has.
I wrote this blog post doing some of our math around this: https://www.backblaze.com/blog/cloud-storage-durability/
But here is the thing: if you value your data, like if you will really go out of business if you lose it, then you should store three copies with AT LEAST two separate vendors. No matter how reliable any one vendor is, "stuff can happen" like your credit card is declined and the vendor deletes all of it.
I would recommend you use two separate vendors like Amazon S3 and Backblaze B2, and use two separate credit cards that expire on different cycles. I believe the credentials for login should be different on those two accounts, and the same one employee shouldn't have the credentials to both. Because one disgruntled employee should NOT have the ability to put you out of business. If you want some other thoughts, here is a blog post Backblaze wrote called the "3-2-1 Backup Strategy": https://www.backblaze.com/blog/the-3-2-1-backup-strategy/