Object Storage: AWS vs Google Cloud Storage vs Azure Storage vs DigitalOcean
chooseacloud.com
chooseacloud.com
B2 is cheaper than every option here for both storage ($0.005/GB) and egress ($0.01/GB).[1] Their transaction pricing is also cheaper.[2] Despite being cheaper, it’s still hot storage, so you can immediately download buckets, in whole or in part. I’ve personally used it to backup (and restore) terabytes of data for over a year. I doubt it has an SLA like GCP or AWS, but DigitalOcean doesn’t either, yet it’s listed here. I find the B2 API documentation to be very readable as well.[3]
I’ve used AWS S3, Glacier and GCP Nearline, Coldline. I can’t think of a specific thing that has disappointed me about B2, and the reliability has been excellent. The nature of my work is that I have very large datasets, and B2 becomes extremely competitive when you’re backing up tens of terabytes or more.
_________________________________
1. https://www.backblaze.com/b2/cloud-storage-pricing.html
Also, if you upload an object with the same name (e.g. myphoto.png) it creates a new version, and there's no way to stop this. I don't want or need a new version!
I have a feeling this restriction is in place because of the underlying "vault" implementation.
I wouldn't use it for non backup object storage, as it's in a single data center. S3 and Google Cloud are completely different, I'm not sure on the Google specifics but S3 has data replicated across three AZs.
(edit: I’m talking about geographic redundancy. B2 does use redundancy inside the datacenter to protect from drive failure, etc)
When you say 'key', do you mean the opaque object ID granted to objects in B2 which object stores like S3 doesn't have at all? I don't understand the rant here. S3 only operate by 'name' (key).
I have had no problem integrating B2 side-by-side with S3 and GCP in the products I have written. Their high-level models are largely compatible.
It would be quite simple for backblaze to implement, though, so I guess it might come as a future option.
Just to clarify you statement about atomicity: All writes to S3 and B2 are 'atomic' (and B2 can also verify the hash and reject on failure for an extra layer of security). The difference you mention is just that you can superficially "disable" versioning for S3, so that each key only stores one object version at a given time.
The only difference to the upload when S3 versioning is disabled is what happens during the metadata update: With versioning, a version is appended. Without, the version is replaced.
For B2, simulating disabled versioning is two operations: Upload a new object, and delete the old one. As long as the object is only referenced by name, this will also be atomic.
It's only slightly different from the name+version approach, but I personally like the concept of unique identifiers better. It feels nicer.
Disclaimer: I am in risk management.
Google Cloud Storage's multi-regional support (asynchronously replicated to two or more geographic locations) is even easier than AWS: You simply specify, on bucket creation, whether the bucket should be regional or multi-regional.
I stand by my statement about the benefit of multi vendor object replication.
The world is not only Google, Amazon, Azure and DigitalOcean!
transfer costs and latency makes it unusable for anything interesting.
I'd also add - unless you are going to include a median/mean latency number many options will look absurdly good even though we know they're a horrible idea, e.g. https://en.wikipedia.org/wiki/GmailFS. I also think "unknown" availability looks a lot like a durability issue..
I wrote recently a series of articles comparing most of those providers and explaining how to use them with JavaScript:
- Amazon S3: https://medium.com/@javidgon/amazon-s3-pros-cons-and-how-to-...
- Google Cloud Storage: https://medium.com/@javidgon/google-cloud-storage-pros-cons-...
- Microsoft Azure Blob Storage: https://medium.com/@javidgon/microsoft-azure-blob-storage-pr...
- Backblaze B2: https://itnext.io/backblaze-b2-pros-cons-and-how-to-use-it-w...
- DigitalOcean Spaces: https://medium.com/dailyjs/digital-ocean-spaces-pros-cons-an...
- Wasabi Hot Storage: https://medium.com/@javidgon/wasabi-pros-cons-and-how-to-use...
trying comparing performance: latency, bandwidth up/down and scalability in the face of many concurrent requests. And then do that wrt. compute resources in various locations.
Basically, there are loads of issues with rejected requests because of rate limiting (returns a lot of 503 "slow down" responses). Basically, I don't recall ever receiving this from S3. You can check the forums to see more in-depth discussion.
The good part: This is a solvable problem, and I hope they relax these limits very soon.
Another great anecdote: their API is 99% compatible with S3. In fact, the official recommendation is to use the AWS SDKs on the server, which I am doing!
That being said, last month I wrote over 60 million files to S3 and the number of failed writes were tiny (solved by simply retrying the write)
I've not used DOs spaces yet, I'd love to know if they guarantee read after write consistency. I know S3 has some issues with that depending on use case. Maybe it's time for another look at spaces.
Amazon S3 is much better in the sense that there IS dynamic scaling if it notices spikes.
To the defense of DO, they are newer and their business model is "cheap,cheap,cheap", so they can't compete at the same level.
This sounds complex but is actually fairly easy to do. For example, if you are not dealing with randomly generated ID's, try using Base64 encoded keynames rather than the keynames themselves. (A more limited character set is helpful, though.)
A better name than user_f38c9123 is f38c9123_user, which will allow a statistically even shard size across the full 16 character range of the first character. As more performance is needed, S3 will automatically shard the second character into 16x16 (256) possible shards, etc.
Also, using a more limited character set such as hexadecimal (that is, [[a-f][0-8]]*) (or just numberic digits) for the first characters of a filename will shard more evenly than a full alphanumeric [A-z][0-9][etc].
Jeff Barr had a blog post on this a while ago.. here it is:
https://aws.amazon.com/blogs/aws/amazon-s3-performance-tips-...
"By the way: two or three prefix characters in your hash are really all you need: here’s why. If we target conservative targets of 100 operations per second and 20 million stored objects per partition, a four character hex hash partition set in a bucket or sub-bucket namespace could theoretically grow to support millions of operations per second and over a trillion unique keys before we’d need a fifth character in the hash."
We actually do exactly this in Userify (blatant plug, SSH key management, sudo, etc https://userify.com) by just switching the ID type to the end of the string: company_[shortuuid] becomes [shortuuid]_company. It makes full bucket scans for a single keyname a bit easier and faster if you use different buckets for each type of data, but you actually will get better sharding overall by mixing all of your data types together in a single bucket. The trade-off is worth it for the general case.
The example is 200GB storage with 2000GB data transfer (out) every month. That's a LOT of data going out every month, so I am guessing the scenario is if you are hosting a photo library and lots of people are downloading every month.
If however you are just using the service as an online storage to hold <100GB of data as backup (i.e. mainly only transfer in), then S3 turns out way cheaper than DO.
Not knocking either service - I actually use both, for different use cases.
Support alternatives even if they are a bit more expensive.
The main thing stopping me from using Digital Ocean is AWS RDS which is amazing. If they could bring out a similar solution with backups and MySQL and PostgreSQL, that would be amazing. :) I could then run my apps in docker.
Yes, AWS S3 is cheaper if you have no egress traffic, and the monthly storage cost falls below $5, which is a minimum price for DO.
This scenario is very much a corner case. If you have egress traffic, or more than $5 worth of storage, then DO runs ahead quite fast.
I'm slightly disappointed that Backblaze B2 was not on the list, though.
It's difficult, sometimes, to guess at numbers that may or may not play to a specific service strength. A lot of times there's something hidden in the details that you're missing as well. In our case, we didn't notice that Cloudinary didn't have an overage rate so if you go over any of the capped limits on a particular plan you have to move up to the next plan level. It was a very unpleasant surprise. We didn't notice it to compare because it simply...didn't appear anywhere.
But what's the use-case for uploading 100 gb a month, and then... Not deleting it (keep paying for storage) and not accessing it?
If not, you need to get this fixed.
Bandwidth costs have dropped by a huge factor over the past few years; but none of this has been passed on.
I really hope backblaze and/or DO manage to cause the big three some hurt on this and get them to reduce prices significantly; 7c/GB is really high these days.
I'm not expecting it to be at the cost of a 10gig cogent connection at LINX, but some drop makes sense. DO/Backblaze pricing seems to be more on the money.
Keep in mind these provides also do charge on top for a lot of other networking services (VPCs, NATs etc) which often can really add up, so I'd expect some of the SDN capex to be absorbed by that.
The larger the provider is the more they are able to negotiate a lower rate from transit providers. The larger a provider is the more peering they get which reduces the total amount of bandwidth they pay for in general.
Then when you consider the cost of the equipment vs the cost of the throughput you will see that as a network becomes larger, network equipment isn't the main cost driving factor, nor are network engineers.
Simply put Amazon is over charging customers by 10x on bandwidth fees.
If we are able to sell bandwidth to customers at $0.01 cents per GB profitably, which accounts for paying for transit, network equipment, and network engineers, turning a profit for reinvestment, then AWS should be offering you a price that is 10x less than $0.01 per GB because their bandwidth cost should be significantly better than ours.
Instead they are charging you a 10x higher price.
You can’t beat the offerings of AWS, but there are definitely some compliance scenarios that are easier to fulfill on Azure.
We rarely get customer requests for Google Cloud. Seems like it’s mostly Azure and AWS (at least at the enterprise level).
For automatic cross-region replication for storage, are you talking about Azure's GRS class? That's the same as GCP's multiregional class and AWS allows you to setup entire bucket replication to anywhere else in a few clicks.
Regional GCS is the storage class equivalent to standard S3 and is $0.02/GB.
That being said, the clouds are great if your compute is co-located in the same place because the transfer fees are waived. Otherwise DO or B2 are probably better options for less usage or more neutral network locations and egress.
Likewise what is the replication story? Can you get event notifications when objects are uploaded/deleted? Is there versioning? Static website hosting? Lifecycle management?
So they say, but the SLA doesn't make any promises about durability. https://aws.amazon.com/s3/sla/
Details of durability here https://docs.aws.amazon.com/AmazonS3/latest/dev/DataDurabili...
As a side note, I think lifecycle management is very useful for backups. A server push its backups to the object storage, but cannot overwrite or delete previous backups. This is useful if the server is hacked...
In theory it should be more reliable, as it's decentralised and your data gets split among multiple servers around the world. The question remains what if the Sia network itself stops being profitable and people all exit at the same time. Although the same could be said for Amazon?
Sia will actually soon be adding a backend to Minio too [3].
The only thing that has stopped me using Sia is you have to have the blockchain running on the machine.
2. https://siastats.info/storage_pricing
3. https://blog.sia.tech/introducing-s3-style-file-sharing-for-...
$0.015 per GB per month
$0.05 per GB downloaded
I see it being used as a base layer for a Glacier type of product, which is still useful because it might give everyone a cheap-ish way to store media, but I wish it could be something more. As far as I know, there's no way to coordinate the sharing geographically, and furthermore the bandwidth is pretty bad, which I think might be due to bandwidth not being counted in the pricing models? There would be huge money in creating a blockchain like Sia that created a more granular marketplace for distributing shards and that factored geography/bandwidth into the pricing. The killer app for this technology is a public market for servers (specifically media hosting). Imagine if Netflix could continually shift their distribution platforms around the world as offices emptied and people turned off their PCs
The problem I have with this article (very brief table), is that the author is comparing two "enterprise" solutions to a consumer solution. With "enterprise" solutions, you get guaranteed uptime and speeds. With Spaces, and the rest I've mentioned, you don't get any of that. Only that your data will still be there as long as you pay out.
As far as I know, the advantage comes from exposing a course-grained API over the network, which is generally more efficient and reliable than doing many small operations. It would be hard to implement Minio's distributed mode using the filesystem API.
Disclaimer: I have no affiliation with Minio. In fact, they're very nearly a competitor with the project I work on. However, Minio is full of ex-Gluster people and I consider many of them friends.
[1] http://pithos.io/ [2] https://www.exoscale.com/object-storage/
Spaces provides a RESTful XML API for programatically managing the data you store through the use of standard HTTP requests. The API is interoperable with Amazon's AWS S3 API allowing you to interact with the service while using the tools you already know.
It surprised me that this isn’t standard on other services.
https://docs.microsoft.com/en-us/rest/api/storageservices/un...
Storage is something that won’t make much money in a few years I believe. I think the Egress overcharging maybe finally seeing decent competition.
> If your use case creates an unreasonable burden on our infrastructure, we reserve the right to limit your egress traffic and/or ask you to switch to our Legacy pricing plan.
> Wasabi’s hot cloud storage service is not designed to be used to serve up (for example) web pages at a rate where the downloaded data far exceeds the stored data or any other use case where a small amount of data is served up a large amount of times
Also since you're the cofounder of nodechef, you should add a disclaimer when mentioning your own product.
But I have to admit, the $5 fees some of these new offerings have is probably to filter out some type of customers.