Amazon S3 and Glacier Price Reductions
aws.amazon.com
aws.amazon.com
The bandwidth costs are so far out of line with what the network transfer actually costs, it just feels like price fixing between the major cloud players that nobody is drastically reducing those prices, only storage prices.
Charging 5 cents per gigabyte (at their maximum published discount level) is the equivalent to paying $16,000 per month for a 1 gigabit line. This does not count any operation costs either, which could add thousands in cost as well, depending on how you are using S3.
There are several providers that offer a unmetered 1gbps line PLUS a dedicated server for ~600-750/mo. Providers like OVH offer the bandwidth for as little as 100/month. ( https://www.ovh.com/us/dedicated-servers/bandwidth-upgrade.x... ) I am just not sure how amazon can justify a 160x price increase over OVH or a 30x increase over dedicated server + transfer.
For the time being, the best bet is to use S3 for your storage and then have a heavily caching non amazon CDN on top of it (like cloudflare) to save the ridiculous bandwidth costs.
A consulting customer came to me a year ago, with a growth from 200TB/year in data production to over 6PB/year and their budget couldn't sustain that jump (or anywhere close to it)
Having come from the mass-facilities and data center space with MagicJack, I knew the wholesale cost of bandwidth, power and drives were continuously falling.
There are certain clients and use cases that need access to their data all of the time and the very bones they are built on is based on collaboration (Genomics).
For example, this client is now storing 6PB of data with us, 3 copies in separate data centers. We are half the price of S3, and we include all the bandwidth for free, but limited to a 10GigE per PB stored. This has worked out extremely well - we were about 20% (!!!!) the price of Amazon after you factor in bandwidth.
There are lots of challenges we faced, like over zealous neighbors in the environment, storing lots of small objects and high usage of ancillary features like metadata but for customers of any size. By putting the "tax" on bandwidth, a lot of these business cases are solved. I see why Amazon does that.
AWS is truly great, but as you get into very high scale (specifically in storage - 2PB+), it becomes extremely cost prohibitive.
However, S3 has the same egress pricing as EC2. Do you think it's really a "business case tax" they're applying across all services?
In the current world, they can keep prices for some products below costs but make their money with bandwidth and the other services people are forced to use to avoid egress traffic.
Which AWS products are loss leaders?
S3 storage pricing is not exactly cheap. Neither is EC2 instance pricing.
"Otherwise everyone would use S3 together with Google compute engine and Azure databases (let's assume they'd be cheapest). In this scenario all providers would lose out."
No, S3 would do well, GCE would do well, Azure would do well. Providers only lose out to the extent their products no longer compete on merit alone.
I think the three providers are smart enough to know why they charge that much for bandwidth. And this is the only reason I could think of why all 3 of them charge that much. And I'm pretty sure that some products run at a loss, they do for nearly every company. But AWS won't tell us which ones.
It's reasonable to think that S3 is loss making or about breakeven on its own but recoups costs due to bandwidth charges.
Feel free to send me a note, my contact info is in my profile (I helped build preemptible VMs and I'm sort of fascinated you're doing this).
Sorry for the delay - yes it's ploid.io - nothing up there yet. We've been in stealth mode while we've been building the system for our first client - HudsonAlpha institute for biotechnology Feel free to ping me at brandon at ploid.io - happy to share any insight we've gained!
Feel free to ping me at brandon at ploid.io - happy to share any insight we've gained!
Yes, it makes me wonder. For many applications, the ridiculous bandwidth costs will be substantially higher than compute costs, so it makes no sense that nobody in this supposedly competitive market wants to compete in this area.
> There are several providers that offer a unmetered 1gbps line PLUS a dedicated server for ~600-750/mo. Providers like OVH offer the bandwidth for as little as 100/month. ( https://www.ovh.com/us/dedicated-servers/bandwidth-upgrade.x.... ) I am just not sure how amazon can justify a 160x price increase over OVH or a 30x increase over dedicated server + transfer.
OVH and other large volume providers are probably relying on the fact that many customers won't use 100% of their purchased capacity, but even taking that into consideration, AWS/GCE/Azure bandwidth pricing is insane.
Part of building a moat? On one side we have a big stack of custom APIs. On the other -- once you have a few hundred terabytes of data in our systems -- it will be very costly to emigrate.
I knew as soon as I saw the headline that the would be reducing storage costs (who cares????) and not reducing the bandwidth costs which have been stuck for years, which are price-fixed with the other providers, and which account for the vast majority of our S3 bill.
S3 already have request pricing in place: https://aws.amazon.com/s3/pricing/#Request_Pricing
In addition to that, the pricing for "Data Transfer OUT From Amazon S3 To Internet" is exactly the same as for "Data Transfer OUT From Amazon EC2 To Internet", so this is not specific to S3 but to EC2 and AWS as a whole it seems.
AWS appear to have really expensive egress costs (or really profitable egress margins) compared to OVH and Hetzner. If so, then something is stuck, either the costs are not being addressed or the margins are not being passed on.
Also, it means using DNS with them as well.
Here is a blog post that shows the same problem we were having: http://goldfirestudios.com/blog/135/CDN-Benchmarks%3A-CloudF...
There is also a GitHub comparison for all object storage providers (this price reduction isn't included yet as I see) [2].
[1] https://www.ovh.com/us/cloud/storage/object-storage.xml [2] http://gaul.org/object-store-comparison/
I wish the Route53 would go down... Hard to beat, even if I were running job-replacing amounts of income.
At that point, I think the reliability and multicast nature of Cloudfront or some other CDN is worth the $5 premium over DO/Vultr/OVH....
B) Not everything is about selling ad space, or even generating revenue.
I had wanted Amazon to wrap it in something where they managed that complexity for a long time. Looks like they finally did.
Now the only thing Amazon needs to do is expand free tiers on all of their services, or at least very low cost ones. I prototype a lot of things from home for work -- kinda 20% time style projects where I couldn't really budget resources for it. The free tier is great for that. All services ought to have it -- especially RDS. I ought to be able to have a slice of a database (even kilobytes/tens of accesses/not-guaranteed security/shared server) paying nothing or pennies.
Glacier has supported for something like 18 months now, the ability to put a policy on your vault that capped your maximum retrieval cost. Whenever your request would cause you to exceed that limit, it would get throttle response that the SDK handles happily. I've used it when I needed to retrieve a whole bunch of data and wanted to do it faster than the free tier supported. I set it at $5 and just left the retrieval running.
$11.50/month doesn't sound too hard to budget for.
it seems the idea of the Amazon free tier is to give people a taste of AWS so they can decide to go in more or not. It's not really designed to be a free prototype for existing large customers product. Like the other poster said, you can host a tiny VM for $6 or so a month, not a big expense.
If you are asking for $4000 a month production cluster, yes that is harder to just get..
Glacier was, from my point of view, literally unusable. Now, it's usable. I still may not use it but at least I can.
Of course, then you're using Dynamo.
And even worse, there is no way to prerender an SPA site for search engines without standing up an nginx proxy on ec2, which completely eliminates almost all of the benefits from Cloudfront. This is because right now S3 can only redirect based on a key prefix or error code, not based on a user agent like Googlebot or whatever.
This means that even if you can technically drop a <meta name="fragment" content="!"> tag in your front end and then have S3 redirect on the key prefix '?_escaped_fragment_=', that will be a 301 redirect. This means that Google will ignore any <link rel="canonical" href="..."> tag on the prerendered page and will instead index https://api.yoursite.com or wherever your prerendered content is being hosted rather than your actual site.
Not only is it a bunch of extra work to stand up an nginx proxy as a workaround, but it's also a whole extra set of security concerns, scaling concerns, etc. Not a good situation.
edit: For more info on the prerendering issues, c.f.:
https://www.fwdeveryone.com/t/QOU4DQDbS8e4tddfI22J0w/s3-prer...
Of course this won't yet show up in Google until we get that nginx proxy stood up or this feature gets implemented. :-)
Also, doesn't google penalize sites that serves different content to googlebot and other UAs?
The content you're serving to Google needs to accurately reflect the content users are seeing.
You can setup "Origin Custom Headers" in CloudFront ;)
What exactly are the benefits over simply setting up nginx if not simplicity? Yeah, it's great to just serve the asset from s3 but the complexity of what you just described negates it almost entirely.
We weren't fully aware of these limitations when we decided to host our site on S3. Had we been, we may have just used nginx.
There are obviously a ton of benefits to S3 and Cloudfront, it's just than in practice you can't really get them if you need Google to index your site. And while Google claims they can now execute javascript and include async content, in practice this isn't true for any real Angular or React site.
And even if every search engine were to magically execute js correctly, you'd still need to prerender your site in order for Facebook and Twitter to populate the preview cards for your site with the proper headline, summary, and image.
It does work if you want to host a static site, and it's nice that they offer a bunch of extra niceties to help make those work... but expecting things like user-agent-specific redirects is a bit much for what's essentially a filesystem.
Yes it gives some scalability, but so do many cloud providers. Digitalocean and Vultr both have SSD storage that you can attach to VMs. The speed I've seen is fast enough to easily saturate the bandwidth. Of course you'd need to scale up when you need bandwidth of more than 100mbit, but this is still cheaper than paying for AWS bandwidth and servers.
Cloudfront or API gateway help you integrating with non static resources.
Bandwidth price is identical with ec2.
With lambda out there, for me ec2 is just legacy.
but since ?_escaped_fragment_= is a suffix, not a prefix, i don't think redirection rules help.
https://www.fwdeveryone.com/t/mdDouBesQwCv0za_o7GjCA/serving...
[1] https://prerender.io/js-seo/angularjs-seo-get-your-site-inde...
In that case, can't you do a cloudfront custom behaviour on /?_escaped_fragment_=*? I havent't tried this.
Install the AWS CLI (https://aws.amazon.com/cli/) and choose whatever method you like for making an encrypted local backup. Then sync that backup partition to S3 every day.
Here's an example command you can adapt for cron to call via a shell script:
/home/tom/.local/bin/aws s3 sync /mnt/backups/daily s3://your-s3-bucket-name --storage-class STANDARD_IA --acl private --sse
The -sse means server side encryption which is redundant since I've encrypted the data prior to uploading, but why not?To backup nearly a terabyte of photos costs about $4/month in storage costs. Uploading costs a bit extra due to the pricing for requests.
I looked into Glacier a couple of years ago, and, from memory, restore costs were insane.
I'd highly recommend not repeating my mistake - use a real backup service for your actual data. Though rolling your own can be fun and interesting, it's probably a bad idea to bet your data on it.
As others have said, if you're just trying to back up your Mac, take a look at Arq.
Disclosure: I work on Google Cloud (and get $30/month of credit, so my first 3 TB are free).
Note: Glacier is not a backup service, it's an archive service. It's for long term backups, vs the relatively short term of standard backups (Full, delta, delta.. cycles). With Glacier if you delete an archive before 90 days, you'll still be charged for the full 90 days of storage.
The data trove is fairly unique, and valuable in being the only of its kind, but we don't need anywhere near instant access to most of it.
Hypothetical: you create a logging service for users to send all their log data to you. You promise 365 days of archives, but 30 days of data accessible at any time. You create a lifecycle rule on your S3 bucket to automatically archive data to Glacier 30 days after creation. On the 31st day, your user decides they want to look at an old log. They click the big Download button. You display a message saying they'll get an email from you when that data is ready to download.
http://docs.aws.amazon.com/cli/latest/reference/glacier/init...
From there you can submit a "DescribeJob" request, with that Job ID as the parameter, and the Glacier service responds with the state of the job.
Once the job is marked as complete, you submit a "GetJobOutput" request with that Job ID. That response is the archive body. (similar to how you'd do a GET request from S3).
You've got 24 hours to start the download of the archive before you'll have to repeat the entire InitiateJob->GetJobOutput cycle again.
I think through the API you do not leave the connection open, you check with whatever frequency you want and when it's ready, the response will include the temporary location on S3 for the file.
[1] http://highscalability.com/blog/2014/2/3/how-google-backs-up...
[2] http://www.everspan.com/home
[3] https://www.microsoft.com/en-us/research/publication/pelican...
[4] https://code.facebook.com/posts/1433093613662262/-under-the-...
But this wouldn't explain why the read rate factors into cost. Maybe they scatter data across tapes as well, and higher bandwidth requires loading more tapes concurrently?
Again, total conjecture. Please let me know which parts are wrong and which are right.
However using Glacier as a simple store from the command-line seems horribly convoluted:
https://docs.aws.amazon.com/cli/latest/userguide/cli-using-g...
Does anyone know of any good tooling around Glacier for the command line?
Again, huge disclosure: I work on Google Cloud.
I bet that when Backblaze increase their scale by adding data centers they will decrease prices, not increase them.
According to Ford's laws of service: Price, volume (scale) and quality are never opposed.
1. Decrease price and you can increase volume.
2. Increase volume and you can increase quality.
3. Increase quality and you can increase volume.
4. Increase volume and you can decrease price.
As far as I am aware GCS does erasure coding across sites?
Backblaze could do multiple tiers of erasure coding and they would still be able to reduce prices given more scale, ceteris paribus.
It's not a question of number of replicas, data centers or technical implementation, but a question of pricing policy.
Does one want to use volume and scale to drive prices down (and cheaper prices to increase volume) or does one want to use volume and scale to bloat margins? Backblaze are arguably doing the former.
Does one want to lock customers into an ecosystem by enforcing excessive bandwidth prices or does one want to pass on bandwidth cost-savings to customers? Backblaze are arguably doing the latter.
Backblaze would continue to be cheaper because their pricing policy serves customers across all dimensions.
More scale is definitely less dollars not more (even if it means a fraction of a few more erasure coded shards across sites).
Disclosure: I do not work for Backblaze.
Either way, good news on the storage price reductions :)
See my other comment, it got link to article about S3 costs optimizations which got more detailed recommendation.
I am guessing it comes down to the algorithm used to compare and upload/download files. I believe the two solutions above use a separate 'index' file on S3 to track file compares.
Users authenticate with an API gateway endpoint, we do a PUT to store a descriptor file, send a presigned PUT URL back so they can upload their file, we then process the file and do a COPY+DELETE to move it out of the "not yet processed" stage and finally do another PUT to upload the resulting processed file.
Despite a lot of data, the storage bill is barely scratching $40, but we're at almost $700/mo on API calls.
Edit: Sent!
We just use the AWS SDK on our Ruby back end. The user file is first uploaded to the (EC2) app server, then we use the SDK call to transfer it to the S3 bucket. Our storage and request costs are about equal at this stage.
Using Lambda/Node, I guess that the SDK is not an option and you have to use the pre-signed URL method? Or else use Python and the SDK library?
By the way, I wrote article, how to reduce S3 costs: https://www.sumologic.com/aws/s3/s3-cost-optimization/
Best source I could find was: https://storagemojo.com/2014/04/25/amazons-glacier-secret-bd...
Shouldn't that be $5.025? Or did I misunderstand?
Am I understanding this right? $0.023/GB/month for Glacier, so * 12 months/year = $0.276/GB/year, which means:
10GB = $2.70/year
100GB = $27.00/year
1TB = $270.00/year
...
And this is only the storage cost. This doesn't take into account the cost should you actually decise to retrieve the data.So considering a 1TB hard drive [0] costs $50.00, how is this cost effective? I can buy 5x1TB hard drives for the price of 1TB on Glacier.
I understand there is overhead to managing it yourself. So, is this just not targeted to technically proficient folks?
[0] https://www.amazon.com/Blue-Cache-Desktop-Drive-WD10EZEX/dp/...
It's a running service (rather than cold storage) which means you need servers, power network ports, etc. Again, times 3.
Finally, include the software development to build the server stack and the staff for 24x7 ops and security.
It's possible to beat their pricing but most people who actually do manage it by cutting in an area which they don't need.
The reality though is that no business would be comfortable with the plan being "our data is replicated offsite at Steve's house". And, other than maybe a pair of NAS boxes (one at each end) the cheapness of the solution assumes you have a great network connection between the two, a machine to plug it into, and only need a single hard drive. That is, how would you do offsite, active backup of say 50 TB?
Disclosure: I work on Google Cloud (and do offsite backup to GCS Nearline, that I'll move to Coldline when I get a minute to play with our new per-object lifecycle rules).
> tl;dr www.feralhosting.com is down, database lost, slots are up and will remain up. We're moving to the honour system for paying bills. ETA 25th November.
Not exactly a good first impression, to say the least.
Hardware is usually not a business' main cost but it does matter for home users, small businesses or startups that didn't get funded yet, some of whom might consider Tarsnap or some other online storage solution which uses Glacier at best and S3 at worst. Now you could suddenly be 7× cheaper off if you do upkeep yourself (read: buy a raspberry pi) and if you throw away drives after one year.
Google cloud nearline costs $0.12 per gigabyte-year with prices that will continue to fall. For a typical 500g hard drive that saw perhaps 700g of unique data, that's $84/year to have an outside-the-house replicated backup using something like Arq.
If you want to compare them, you have to buy space on a different continent, and store your backup there.
Your pricing assumes that the drive is never powered.
LTO (-4) tape had gotten capacious and cheap enough that I went back to tape (I'd outgrown DAT); if I didn't have a big sunk cost in a well working tape system and pool of tapes, which are very easy to put in e.g. a safe deposit box (they're a bit fragile, but nothing like a hard drive), I'd already be using one of S3, Glacier, or Backblaze, maybe even GCS since suddenly and irretrievably losing access to my backup data because a bot decided I was evil would not likely coincide with a total data loss at home (Google simply cannot be trusted if you're small fry like myself, as HN has been discussing as of late).
As Glacier has gotten sane enough to use without twisting your mind into a pretzel, with the new price reduction for slow retrieval I can seriously think about adding it to the mix and switching to it when my LTO-4 tape drive dies someday (e.g. ~3TiB for ~$12/month per my quick calculation just now), instead of buying another tape drive.
Retrieval in all Google Cloud Storage is instant and for Coldline is $.05/GB (and Nearline $.01/GB). If you value that instant access, it seems the closest you'd get with the updates to Glacier is via the Expedited retrieval ($.03/GB and $.01/"request" which is per "Archive" in Glacier). Then you have to decide how much throughput you want to guarantee at $100/month for each 150MB/s. (It's naturally unclear since it was just announced what kind of best-effort throughput we're talking about without the provisioned capacity).
If you're never going to touch the bytes, and each Archive is big enough to make the 40 KB of metadata negligible then the new $.004/GB/month is a nice win over Coldline's $.007. Somewhere in between and one of the bulk/batch retrieval methods might be a better fit for you.
But again, it's still a bit of a challenge to go in and out of Glacier while Coldline (and Nearline and Standard) storage in GCS is a single, uniform API. That's worth a lot to me, and our customers. But if Glacier were a good fit for a problem you have, and you're talking about enough money to make the pain worth it, you should seriously consider it.
Disclosure: I work on Google Cloud, so naturally I'd want you to use GCS ;).
https://cloud.google.com/pricing/tco/storage-nearline
(work at Google)
I just noticed that the new pricing for Standard (2.3¢) is now less than the pricing for Reduced Redundancy (2.4¢)! So there appears to be no reason to use Reduced Redundancy anymore.
- The price reduction on S3 is good! Kudos AWS.
- The price change on glacier is a fucking disaster. They replaced the _single_ expensive glacier fee with the choice among 3 user selectable fee models (Standard, Expedited, Bulk). It's an absolute nightmare added on top of the current nightmare (e.g. try to understand the disks specifications & pricing. It takes months of learning).
I cannot follow the changes, too complicated. I cannot train my devs to understand glacier either, too much of a mess.
AWS if you read this: Please make your offers and your pricing simpler, NEVER more complicated.
(Even a single pricing option would be significantly better than that, even if its more expensive.)
> AWS if you read this: Please make your offers and your pricing simpler, NEVER more complicated.
I am reading this. We made them simpler. Decide how quickly you need your data, express your need in the request, and that's that. If you read the post you will see that these options encompass three distinct customer use cases.
I don't think you understand just how insanely, incredibly, bizarrely complicated the old glacier pricing was.
I don't think you realize how insanely complex the entire S3 pricing model once you get out of the "standard price".
Maybe I just have too much empathy for my poor devs and ops who try to understand how much what they're doing is gonna cost them. It's only one full page, both sided, of text after all.
There are so many combinations of hard drives that will result in different performance for different situations all with different costs. Then you start talking about cold storage as well and you've moved into other media formats.
Just because there is a page worth of a pricing model doesn't mean AWS or any cloud provider is doing anything incorrectly. You're paying for on demand X and engineers who are going to utilize that should understand it as well as they would understand how to build an appropriate storage solution of their own. On demand just means now they don't have to take the time to design, implement and operate it themselves.
I'm saying that anyone who thinks this is more complicated than it was does not understand just how crazy glacier pricing was before. Three static glacier pricing tiers is a lot better than the previously system which is so complex that earlier versions of the AWS pricing page just gave up and called it "free".
(Briefly: The old Glacier model's pricing wasn't based around data transfer, but on your automatically retroactively provisioned data transfer capacity based on your peak data transfer rate, billed on a sliding scale as if you'd maintained that rate for the entire month. There's probably a less intuitive way to bill people for downloading data, but if so I've never seen it. It was a system seemingly designed to prevent users from knowing what any given retrieval would cost.)
I can assure you that if AWS had announced a simple flat-rate structure that was more expensive for everyone there would be plenty of unhappy existing customers. It's a tough call to balance simplicity and efficiency, by making this complex for you, they allow you to opt-in to the "screw it, give me my bytes back as fast as you can" model. You could just pretend that's the only option ;).
Again, I'm not criticizing your unhappiness with the complexity of Glacier. But I think it's only fair to recognize that the folks at AWS have just released a major improvement that provides real value for some customers.
Disclosure: I work at Google Cloud.
byte[] buffer = new byte[128*1024];
while(socket.read(buffer) != -1) {
// do something with buffer here
Thread.sleep(1000);
}(If you end up never downloading the file, you still get charged)