Google Cloud Storage Nearline graduates to general availability
googlecloudplatform.blogspot.com
googlecloudplatform.blogspot.com
And its only 1 month free at that level of commitment; anything more than 1 month requires a migration commitment to more than the 1PB.
EDIT: At the time this comment was written, the thread title referred to the 100PB free for up to 6 months offer, not the fact that Nearline was now in GA. While the content of the comment is still true, the title that is being clarified by the comment is no longer present. (Similar things seem to be true of a number of the other top level comments.)
That is a pretty steep hit when you can use a lower quality object store:
https://www.runabove.com/storage/object-storage.xml
For $.01/GB and $.01/GB to pull.
Yeah, it may not be as highly available as Nearline but for the 99% of the time it is available...you don't have a 3 second delay and you aren't paying a 13x premium to pull data out.
For instance, lets say you have a 50GB backup that you automatically verify via a testing process. You don't want to burden the disks of your primary datastore longer than you have to [the source of the backup] so your workflow is:
Database Server -> Create Backup Locally -> Push Backup to Object Store Bucket [$.50/month to store the backup, assume you store 7 days] -> Spin up VM which pulls from Object Store Bucket and verifies the backup [$.06/day for an 4GB Linode for an hour to do this which is plenty of RAM and time ] -> Pull the Backup & Verify [$.50/Day]
30 Days Backup Cost:
$20.30/month to maintain a week of verified backups with RunAbove = (7 * 50 * .01) + (30 * .06) + (30 * .01 * 50)
$200.30/month to maintain a week of verified backups with Nearline = (7 * 50 * .01) + (30 * .06) + (30 * .13 * 50)
I don't know why anyone would willingly pay an order of magnitude more for cold storage.
(7 * 50 * .01) + (30 * .05) + (30 * .01 * 50) = $20.00/month
https://cloud.google.com/compute/pricing#localssdpricing
$0.113
So its:
[apples to apples; local ssd]
$186.95 = (7 * 50 * .01) + (30 * .05) + (30 * .1213 * 50)
[persistent provisioned ssd]
$54 = (7 * 50 * .01) + (30 * .05) + ( ( ( (.17 * 96) / 720) + .01) * 50 * 30)
[1] https://cloud.google.com/storage/pricing#storage-pricing
The bandwidth cost is simply too high. They really need a "low quality" v. "standard quality" option. Plenty of people sell bandwidth at $.01-$.04 a GB with 99%-99.99% availability
100 PB = 102400 TB = 104857600 GB = 107374182400 MB.
You have 365/2 * 24 * 3600 seconds in 6 months, that's 15768000 seconds.
So you need to upload ~7 GB/sec, that's 60 Gb/sec constant upload.
We don't have a cold storage product, so perhaps not relevant for nearline ... but it is relevant for their online options. You do have the ability to migrate to another provider.
Two random notes:
- we're announcing zfs send/recv[1] as a transport in the next few days and there will be s3-competitive pricing for this product at TB+ levels.
- "HN readers discount" is still available all these years later. Just email.
[1] over SSH
While I think ZFS is great, it's just not a solution for really large storage.
Also, if you are still using s3cmd, that only tells me how inexperienced you are.
I'm not sure you understand how s3cmd is being used in the above context ... the idea is this:
ssh user@rsync.net s3cmd get s3://rsynctest/mscdex.exe
Also, we've been doing this longer than Amazon has.Can we have a brief (one or two sentence) summary of what people should be using instead?
Do they have fixed costs or is this one of the last ways they have to to make money on their platforms?
Google's Network is a vast, secure, and performant Global SDN. This allows us to do some cool things. Therefore, Google has a fundamentally unique value proposition:
- Traffic between zones/regions at Google never leaves Google Network. Therefore, no need to setup VPC or VPN between zones/regions. Makes deployments much much simpler.
- Google Compute Engine is a VPC out of the box, even cross-region. You can carve out your own sub-VPCs easily using firewall rules.
- Traffic from Google to end-user is not being dumped off of Google network as soon as possible. Google carries your traffic as close to the end-user as possible.
Edit: take a look at this as well.. http://googlecloudplatform.blogspot.com/2015/06/A-Look-Insid...
Cloud Storage considers retrieving an object through its HTTP url is considered a class B XML request type which is priced at $0.01/10,000 ops. This is 1.3x-2.5x more expensive than cloudfront and s3 respectively. This is also not well documented and caused me some trouble. I didn't expect HTTP GET requests to be counted as XML API requests.
That + it being more expensive than cloudfront and s3 came as a rude shock :(
In general I find google cloud documentation and the service much better and more pleasant to work with than AWS but in this case it is not. Both S3 and Cloudfront have much clearer pricing (and positioning). In network pricing, both S3 & Cloudfront are significantly cheaper
“Networking is a red alert situation for us right now,” explained Hamilton. “The cost of networking is escalating relative to the cost of all other equipment. It is Anti-Moore. All of our gear is going down in cost, and we are dropping prices, and networking is going the wrong way. That is a super-big problem, and I like to look out a few years, and I am seeing that the size of the networking problem is getting worse constantly. At the same time that networking is going Anti-Moore, the ratio of networking to compute is going up.”
http://www.enterprisetech.com/2014/11/14/rare-peek-massive-s...
So the bigger news here is Google Nearline Storage graduating to general availability.
Its 100PB free for 1 month, if you commit to transfer in at least 1PB in the first three months, and commit to keep at least 1PB for 12 months (Its "up to 6 months" free of 100PB, but more than 1 month free requires greater than 1PB commitment.) At $0.01/GB/mo, 1PB of data for 11 months (12 months minus 1 month free) is a $110k commitment.
Its not "100PB for every customer", at least not free, its a variable size discount that covers up to 100PB for 1-6 months of a 12 month commitment period for customers able to make up-front commitments of $110k or more.
2. It's only if you upload at least 1 PB of data in the first three month, which removes a lot of companies and only leaves the serious heavy data potential customers
Amazon has Import/Export which lets you ship drives, is this the best option?
50 TB / 0.000000625 TBps = 80000000 seconds = 925.92 days
AWS Import can probably import at 100MBps or ~6 days.
As is the fact that you have to migrate the data from non-Google servers -- either your own or some other cloud provider -- to qualify, since you're going to pay for the sending the data, junk or otherwise, you are dumping into it.
Cold storage is absolutely a necessity when you reach the scale of imgur, unless you want to run the service incredibly inefficiently.
imgur keeps photos forever [0]. To do so without some kind of cold storage infrastructure in place for all the millions of photos that will probably never be seen again, or perhaps at a rate of once per year, would be infrastructure cost suicide.
Facebook faces a similar challenge and has gone in to detail about how it developed a cold storage strategy to deal with it[1].
I would say that for a service like imgur there is probably a level of parity between the amount of importance they need to place on content delivery and on cold storage.
[0] https://help.imgur.com/hc/en-us/articles/201476457-How-long-...
[1] https://code.facebook.com/posts/1433093613662262/-under-the-...
To me, "SLA" is an enterprise oriented term, which a lot of folks (us) don't care about because we're scaling horizontally and architect with failure in mind.
99% for a typical non-scaling enterprise app is crap.
Because at least part of the target market for this product cares about quantifying the guarantee, whatever it is.
> To me, "SLA" is an enterprise oriented term,
I think that's excessively narrow in terms of who cares about it, but Nearline is clearly in large part an enterprise-targeted offering, so, even if it was a purely enterprise-oriented term, its appropriate.
> which a lot of folks (us) don't care about because we're scaling horizontally and architect with failure in mind.
Knowing expected failure characteristics can be an important input to intelligently architecting with failure in mind.
> 99% for a typical non-scaling enterprise app is crap.
Nearline isn't an app, its one of a set of closely related storage offerings that, by design, would usually be used in coordination with each other and possibly other storage systems by an app. For its role in that stack, 99% doesn't seem immediately unreasonable, to me.