Introducing Google Cloud Storage Nearline
googlecloudplatform.blogspot.com
googlecloudplatform.blogspot.com
In short, if I'm understanding correctly:
$0.01 per GB per month storage
$0.01 per GB retrieval
Normal egress fees on top, so additional ~$0.10 per GB if you want to retrieve outside of Google Cloud.
Early deletion fee. Effectively just a minimum storage charge of 1 month.
It seems this is cheaper than Glacier and quite a bit simpler. The speed restrictions are interesting though.You will be surprised by a retrieval fee of 1000 / 4 * 720 * 0.01 = $1,800.
If you want to get your 1000 GB for free, you'd have to request no more than 69 MB per hour. It will take you about 20 months to get your data!
With Glacier, limit retrievals to less than 5% per month / 30 days / 24 hours = 0.007% of total data stored per hour to avoid these ugly surprises.
Source: http://aws.amazon.com/glacier/faqs/#How_much_data_can_I_retr...
In a disaster situation, a company is going to want hard drives sent to them next day. As others have mentioned, this isn't a money thing, it's a time issue. But it isn't available with Glacier (probably not with Google either...)
https://aws.amazon.com/glacier/faqs/#How_much_data_can_I_ret...
So Google looks cheaper/simpler, but the amazon model is actually quite good for disaster recovery where you may be in a situation where you "need it now at any cost"
The biggest gotcha is that you can only access 0.17% of the data you've stored in any given day without extra charges. So if you've stored 1000 GB, you can only access 1.7 GB per day for free.
The cost for going over your daily "retrieval allowance" can be large, because cost is driven by "peak retrieval rate". They find the highest hourly retrieval rate over the entire month, and charge the whole month's retrieval at the peak rate.
This can get expensive fast. Again, with 1000 GB stored, if you retrieve 200 GB over the course of 4 hours, your hourly retrieval rate is 200 GB / 4 hours = 50 GB per hour. Your free retrieval allowance is 1.7 GB per day / 24 hours = 0.07 GB / hour. So your excess retrieval is 49.93 GB per hour.
Amazingly, that 49.93 GB per hour for 4 hours is charged for all 720 hours in the month, so that's 49.93 * 720 * 0.01 = $359.49.
That's an astonishing $1.79 per GB just to retrieve one-fifth of the data from Glacier storage.
So Glacier only makes sense for true cold storage, where you are very unlikely to touch more than 0.00007 of it in any given hour.
Google's Nearline is much simpler and faster for the same storage cost -- and far cheaper if you need to access more than a tiny fraction of the data.
Source: http://aws.amazon.com/glacier/faqs/#How_much_data_can_I_retr...
I guess he would say he was... sunglasses FROSTBITTEN YEAH!
AWS has now added the ability to set a spending limit to avoid runaway retrievals. Nonetheless, from my point of view the ultra-complicated pricing scheme is nothing short of a disaster for Glacier as a product. I think it will continue to seriously impact its uptake, and Google is smart to exploit the opportunity to offer a simple, straightforward pricing scheme.
edit: Also as others have pointed out, Nearline is limited in retrieval time as well, so the cost difference isn't nearly as large.
Incentivizing users to bring the data online once a day to trickle it out seems bad for all involved.
If Amazon store your data in a tape archive (I don't know if they do, but they at least seem to have similar constraints), they can only access a small portion of the stored data at a time, so they need to control how often people request data.
They could just rate limit everyone, but this way allows people to pay for priority in an emergency while still discouraging everyday read requests.
The pricing makes more sense if you're a large user with data spanning several tapes than if you're in the single terabyte range, but the low limit still discourages you from making requests causally, which helps them keep their SLA.
If they predict that you'll trickle out your small file, they can just read out everything on the first access and cache it online, so there's no extra trips to the archive for them.
Glacier uses low-speed (5400 RPM) consumer drives, which they then clock down lower to save energy. Any given 'rack' only has enough power to power a few drives on that rack, the rest is powered down.
To prevent multiple customers from all trying to pull their data out they needed to introduce a rate limiting system, which they did with this exorbitant pricing.
As long as we are comparing effective pricing, which is the only number that matters with Glacier and "Google Nearline" (and, to some degree, S3) it should be noted that rsync.net PB-scale is 3.0 cents with no additional charges.
That is, the effective price, no matter what your use-case or traffic/usage, is 3.0 cents per GB.[1]
Of course, you do have to buy a petabyte of it ...
> "a two year contract is required." [1]
> "There are no contracts, overages, fees, or license charges at rsync.net." [2]
[1] http://www.rsync.net/products/petabyte.htmlSo should I believe in Google's good-will? I would be fine trying out some services, which are in Google Beta. But my valuable data? They should have a SLA right from the start to gain the user's trust.
I think SLAs are literally worthless since I don't think they encourage even slightly more effort in minimizing whatever issues the customer is concerned about. No one wants servers to go down, no one wants to shut down a service. That a few Google bucks might be on the line would have zero impact.
But if the 3 seconds becomes 6 seconds... Or the price goes up... Or they announce they're end-of-life'ing the product...
Just move your data somewhere else.
Sure, it'd be inconvenient (and maybe expensive) to move. So, you balance all of that out in your mind, and maybe this is the right service for you, and maybe it's not.
If I lose my AWS Glacier stored data (or Amazon bump the prices intolerably), I'll upload it to a competitor _from my local copies_...
Admittedly, I've only had to deal with storage topping out in the tens of terabytes range, so I've never needed to go beyond a dozen or two consumer-grade drives to keep a pair of rsynced copies locally - but I think that same kind of techniques scale all the way out to building your own Backblaze style storage pod if needed.
You're thinking about data that can't be lost, or your customers are screwed. Not all data is like that.
Log files come to mind. They're _nice_ to archive for a long time, but in many businesses, they're certainly not _critical_ to archive for a long time.
Intermediate files, too. You retain the original files in secured storage. But because the intermediate files are large and expensive to re-create, you keep them here in AWS.
You just want cheap and fast.
And if you aren't comfortable even investing development efforts against a beta, don't. Different customers have different risk tolerances, and the fact that an early access product doesn't meet yours doesn't mean it shouldn't be available for those whose risk tolerances it is suitable for.
Surely the whole point of not having an SLA (and indeed the whole point of calling it a "beta") is so that you don't trust it with your valuable data. I'm sure it will have an SLA soon enough.
Once it has an SLA, I can use my tool to work with valuable data.
Don't place all your eggs in the same basket.
Far better to use multiple redundant solutions, and although Nearline is only beta* it offers us an easy and cheap way to increase storage diversity.
*How long did we all rely on Gmail while it was beta?
Any experience with them? how is their stability and speed?
Note that there are separate zones for Google Compute Engine: https://cloud.google.com/compute/docs/zones#available
(edit: add /month)
* Dropbox doesn't guarantee 11 9s like this service does - when I'm backing critical data up, I want to make sure it's _there_. * Dropbox likely wouldn't take kindly to me storing 10PB, whereas that is what this service is designed for. * I've got SLAs and guaranteed speeds with this, Dropbox isn't designed for me to suddenly download 10PB very quickly.
Those claims are from glacier: http://aws.amazon.com/glacier/faqs/ which is priced the same per GB of data.
In this case, the durability refers to the loss of data/objects stored per year - if you're sending multiple PBs of data off to Glacier, you want to be able to retrieve them many years later. Even 5 9s would mean that 1 object out of 100,000 is lost every year, which is quite poor.
http://en.wikipedia.org/wiki/High_availability#Percentage_ca...
In this case, it's not as much service uptime as it is data retention. If you're storing 5+ PB of data, even 5 9s of data loss per year can have a measurable impact.
source: SE at Google Cloud
Also retrieval speeds increase with your data set size: Note: You should expect 4 MB/s of throughput per TB of data stored as Nearline Storage. This throughput scales linearly with increased storage consumption. For example, storing 3 TB of data would guarantee 12 MB/s of throughput, while storing 100 TB of data would provide users with 400 MB/s of throughput.
With storage becoming essentially free the last excuse for not using self-hosted, secure backup tools like Arq[1] disappears.
My thinking is that most people likely max out well below the 0.5T that the same $5 buys you on Google now, making that variant actually cheaper for them than Backblaze.
For those people it means paying less per month and getting privacy and better control (data retention!) in return. Should be a no-brainer.
I'm not sure if you're being serious or sarcastic. ;)
Anyway, if you don't care who has access to your backups then obviously my argument has no relevance to you.
I care very much, however, whether hackers will have access to my files, as they can cause havoc with things like tax returns that the NSA won't (I mean, the government already has my tax returns....).
So I was mainly curious if there was some significant flaw in Backblaze's encryption that should worry me from the perspective of a non-nation state adversary.
You can also set multiple backup locations, by sharing space on your other computers.
They have a proprietary Java client, but it runs cross platform. It's completely painless to add your own encryption key to all backups.
Beware flat rate service providers - it's not a consumer-friendly model:
http://blog.kozubik.com/john_kozubik/2009/11/flat-rate-stora...
Your interests and the interests of your provider should be aligned, not opposed.
Hardly. I pay for support. I pay for liability. If something goes haywire with my backups, I have a phone number I can call. A number that's not software developer/ sys admin who has no idea why I'm calling about my mom's missing holiday photos.
Google "{service} lost all of my data" and see what kind of support these people got. I get that it gives you warm fuzzies but when these companies loose all of your data, you have almost no recourse.
A more likely scenario is that the data is not lost, but my mother (or whoever is calling them) simply has things misconfigured. These companies _do_ provide useful support. Support that they're not going to get if they backup to a developer oriented service directly.
If they lose any data, they're liable for that as well. All the way up to taking them to court on it.
I can hardly recommend it high enough.
The only minor niggles that I've run into were:
- If you have a very large directory (hundreds of thousands of files) then opening the backup-catalog can give you the beachball for minutes. However this really only happens for ginormous directories and is easily resolved by splitting the job into a few smaller ones.
- If you're low on diskspace then the temp files that Arq creates during a backup can drive you over the edge and into a disk-full situation.
Other than that it has been rock-solid for me, the author really knows what he's doing.
Does anyone know an Arq alternative for linux?
It's really great. I'm using it instead of Glacier for all my personal backups.
No thinking about installing software (e.g., Arq, and finding something equivalent for Windows), keeping that software up to date, checking that it's up and running properly, etc. (BackBlaze tells me when it hasn't been able to reach a given machine after some period.)
So, no, I think the SaaS backup industry has nothing to fear from cheap online storage, at least for ordinary folk. (Hackers are a vanishingly small segment of that market.)
We have a "HN readers" discount which is fairly substantial.
rsync.net is 20x more expensive than Google Nearline ($0.20/GB).
Why would anyone choose to use it for Arq when Arq supports both?
rsync.net just saying: "Hey, we're not Google or Amazon" is already a big selling point for some.
Further, Amazon S3 is a better comparison, as far as pricing goes - we're fully live, online, random access storage - not nearline or weirdline like google/glacier.
As they generally do tit for tat on price wars, it will be interesting to see what Amazon responds with.
Glacier is the other big offering from a comparable cloud vendor in this space, so, yeah. The performance claim is directed largely at Glacier, as is the consistent access one.
4MB/s is only around 4% of the bandwidth of a modern 1TB hard disk.
> You should expect 4 MB/s of throughput per TB of data stored as Nearline Storage.
But still, you are correct in that it will take about 1TB / 4MB/s = 2.9d to retrieve all your data.
If you need the ability to do a restore faster than this, then you need to pay more for storage, is all. For many of us, waiting 3d to recover from a catastrophic failure isn't a big deal.
In 3s on one disk you can do roughly 5k seeks and read up to 300MB. They only need to do 1 seek and read 4MB.
Maybe I should start my own service to harness all this infrastructure with something like swift!
The unusual thing about this market is that by constantly improving the price point of their product but keeping their profit margins the same, AWS has forced every competitor in it to actually compete at the limits of their capability - so any scheme like this that got any serious traction would presumably be self-undermining. The economists will have considered all of the unused provided capacity in the model before they priced it. Sorry to be miserable ;)
The large cloud providers sell on brand recognition, convenience and trust, not cost. They may compete with each other on cost to try to cannibalise each others markets, but they're not even close to pushing the envelope on cost efficiency in terms of the prices they offer.
If you have your own servers I can see running zfs and replicating snapshots every few minutes to a remote machine but if you aren't, what else is competitive?
I do agree that a lot of AWS services aren't cost effective though, especially at any sort of non trivial scale.
If you are comparing against a self managed solution, are you factoring all costs into that equation (fully burdened labour costs, disposal, cooling, power, etc.).
In general, you can beat AWS with 3x replication across multiple data centres even with renting managed servers, especially as your bandwidth use grows as AWS bandwidth prices are absolutely ridiculous (as in, anything from a factor of 5 to 20 above what you'll pay if you shop around and depending on your other requirements). Lease to own in a colo drops the price even further.
Anything from 1/3 to 1/2 of AWS costs is reasonable with relatively moderate bandwidth usage, with the cost differential generally increasing substantially the more you access the data.
Most people also don't have uniform storage needs. If you use your own setup, people tend to be able to cut substantially more in cost by reducing redundancy for data where it's not necessary etc.
See also the Backblaze article that mentioned Reed-Solomon - if your app is suitable for doing similar you can totally blow the S3 costs further out of the water that way.
Of course there's good reasons not to roll your own solution until you reach a certain scale, but it's certainly possible to compete with Google and AWS on price.
Well, it would be price competitive if you considered only the hardware costs and not, e.g., the labor cost involved.
I have plenty of servers sitting in racks that haven't been touched in 5+ years; if you were to operate on a "buy and throw out in a year" practice, my typical labour costs would average a couple of hours per new server for setup. Or if you rent managed servers, you don't ever need to touch (or see) the hardware.
Conversely, you can get bandwidth at a tiny fraction of the cost of AWS bandwidth, so the moment you actually transfer data to/from your setup, AWS gets progressively more expensive.
Tax status Business
This service can only be used for business or commercial reasons. You are responsible for assessing and reporting VAT.
and then makes me enter a VAT number... well, I will stick with AWS and Azure, since neither of those "REQUIRE" a vat number or business status.
PS: i know a business name costs about EUR20 to setup, but the VAT requirement is the pain in the left one...
We tried to manually using it, like APIs, Scripts, etc… it’s impossible, it was very had
So we tried cloudberry to help us upload, it will upload but also It won’t work for huge data, with glacier you won’t be able to search, list, find any file you want to download easily, also its not practical to manage all the backup and millions of files manually, also we got millions of photos we needed a way to find them easily like thumbs So Glacier had so many restrictions like 3-5 hours restore time, 5% restore quota hard to use even with utilizes like cloudberry, cant list, search, it’s not useable as its! But the price was attractive
We considered Seagate Evault , but it was expensive and so many hidden fees and complicated for our case Then we tired another solution called Zoolz ( http://www.zoolz.com ), Zoolz does not use your own AWS Account, but utilize their own account and from what I heard from them they got 5 Petabyte, so they got massive restore quota, around 5% of the 5 Petabyte, also its simple to use, they offer Zero restore cost, it’s like Mozy or Crashplan but for business and on glacier , like they internally create thumbs for your photos and store it on S3, so you have instant preview, also you will be able to search and browse your data easily and instantly, they got Servers and Polices, and got a reasonable price, all we wanted , we got the 50TB for $12,000 / year, its more expensive than using Glacier by itself, using Glacier will cost around $6000, but its not practical for a company to use it as a standalone storage
The only disadvantage is that when/if we needed a file we have to wait 4 hours to get it, which is fair, it was faster than when we used to use Tapes :), awe tried to restore 1 TB with 1.2 million files, it took us around 10 hours to complete, which was okay
I don't know which they use but HTTP traditionally uses a `Content Length` or an empty chunk to represent the end of data.
This particular story is about a Google service that is an extremely late entrant into the market (Amazon Glacier, etc) and offers almost nothing new. I'm not saying it's a bad product or not useful to folks, but it's not, by any means, groundbreaking. I'd much rather see the top spot of HN occupied by some startup's new idea or a researcher's new findings.
I also don't mean to imply that "big corporate" == bad. Certain products -- self driving cars, Space X automated landings, etc -- are absolutely worthy of our attention and discussion. I just hope that people would think twice before upvoting a story merely because it's from Google.
http://techcrunch.com/2015/03/11/google-launches-cloud-stora...
This information should challenge your argument that Nearline is not groundbreaking, at least in the "cloud archival storage" space.