New storage classes for Google Cloud Storage
cloudplatform.googleblog.com
cloudplatform.googleblog.com
For me this is in no way acceptable and it seems to be a vicious attempt to sneak in some extra profit without having the customer noticing this upfront. Shure, these harsh words seam like a big exaggeration, but I literally never ever hear or read something about traffic costs in Googles fancy blog posts, and I would make the assumption that only a fraction of the HN community is aware of this fact. 120 bucks for a sloppy TB of traffic is just way too high.
I'm paying about $0.01/GB right now, and I've seen market rate at half that. And that's not even directly using IP transit providers. You can get a gigabit unmetered for $450-1000/mo which you can shove a theoretical 324TB through every month. The difference in the numbers is so staggering I sometimes wonder if I'm even doing the math right.
Perhaps their bandwidth is better somehow (prove it), but 12-18x better it is probably not, and having truffle shavings added to your IP transit really adds up when you're hauling a lot of traffic. If you're doing something with heavy BW usage and low margins, be careful with stuff like this. It quickly becomes much more expensive than doing it yourself.
I'd love to be wrong here. I'm sitting next to 60 pounds of storage servers I'm setting up for a data center, they're taking up my entire living room. I would love to get out of the data persistence business forever. But at these BW rates, it's never going to happen.
any experience using backblaze b2 as hot storage, and glacier or google coldline as backup
There are trade offs. Aside from more upfront costs and a fixed monthly, Ceph is ridiculously complicated. Interface wise, they need some much better abstractions.
B2 was a strong candidate, their BW rates are a little high still but approaching reasonable. The one issue is that they seem to have inconsistent latency. Not a problem for most use cases, but I need nearly all requests to come back in <100ms consistently, as I'm using this for web hosting. Their use case seems more focused on "hot standby backup" than on high availability ATM. I'm strongly considering them as the backup provider for my storage cluster.
FWIW, GCS is not winning any best-of-show awards for latency either. S3 at the time I ran tests was doing a better job.
It really shouldn't be this complex. I would love to just be able to boot an executable with a simple config file and be done with it. SeaweedFS shines a light on how this could be improved: https://github.com/chrislusf/seaweedfs
SeaweedFS is fundamentally trying to do a different thing than Ceph is, and it benefits certain use cases. Reading the Facebook Haystack paper will give you an idea of the differences.
That said, I'm extremely impressed with how simple the interface is. It's very easy to get started.
I work at Backblaze, and B2 is our second product (Online Backup was our first product). We are less experienced at serving up data that has fixed latency requirements, but we're learning and actively improving this area all the time. I'm curious when you last tried it and I would be interested in hearing about your experiences (in a personal message if you like).
In a nutshell, the first time we read a file we build it from the vaults (vaults are our lowest layer which is reliable but slowest) and that should be fairly consistent with some caveats (see below) and then for at least the next 24 hours it should be extremely fast and consistent being served out of our SSD cache layer.
Caveats: If you upload a ton of files it is getting loaded into the most recent vault we have deployed. So for three or four days you and everybody else are uploading tons and tons of files to this exact vault causing great loads. After ten days the vault will be "full" and we'll deploy another vault and suddenly serving up your files will become a lot more consistent and easier and faster because the vault is almost entirely idle. What we just started doing is deploying twice as many vaults at once to lower their load in half while they "fill up" making the response in the first 10 days faster and more consistent.
I'd also like to point everyone at Zach Bjornson's wonderful blogpost from last year [1], which was both thorough and independent!
Disclosure: I work on Google Cloud.
[1] http://blog.zachbjornson.com/2015/12/29/cloud-storage-perfor...
I'm sharing it with a few other people and it's great.
- It's IP transit and a DC from HE, which means it's possibly not BGP peering through other transit providers? If this is true, if HE's network fails, so does your server's internet connection. Typically you want to be multi-homed to ensure redundancy, unless you're doing something like an Anycast CDN and you don't care if it craps out for a few hours.
- HE is the cheapest transit for reasons. It may not be a big deal for what you're doing, but be aware of the differences https://news.ycombinator.com/item?id=5348624
- They only provide a 15A circuit (likely only 80% usable) for a 42U rack, which is pretty ridiculous. You'll blow through that pretty quickly if you're using dual Xeon servers. It's cheaper under their pricing to get a second 42U rack than it is to get the proper amount of power needed for a full cabinet.
- If you're doing Anycast, be aware that HE doesn't provide any BGP communities except blackholing, which could make it hard to tune your network.
(If any of this sounds annoying to deal with and think about, you know exactly why I'd prefer that the cloud providers had better BW costs so I could not have to do this anymore. Again, if the business model works with GCS, congratulations, use it.)
The gigabit transit that comes with the cabinet is HE-only, yep. You can however buy transit and an interconnect from any other carrier in the datacenter (there are quite a few in FMT2 at least).
> They only provide a 15A circuit for a 42U rack, which is pretty ridiculous.
Yeah, it's pretty shitty. I ended up going with Xeon D for the power efficiency. Works pretty well.
Disclosure: I work for Google Cloud.
[1] https://cloud.google.com/cdn/pricing#cache_egress_pricing_ta...
I really think you could differentiate a from your competition hugely by just using that $0.02/GB point for all your customers, regardless of the amount of egress they're using. It's a big chicken-and-egg problem to have those bandwidth costs initially as a startup, but then be able to handle them as a bigger organization. It made us move to our own infrastructure, and once we're here, I already have my infrastructure in a datacenter at that point, so why bother with the cloud?
Even better (for me anyways) would be to provide an option for the "this isn't VOIP or gaming" bandwidth. Not everybody needs that. Actually, probably most people don't. For what I would use GCS for I only need transit to other datacenters. That bandwidth is much cheaper than major ISP peering bandwidth because there aren't monopolies like Comcast extorting everybody for peering.
The call out of 12ct for throughput and 2ct for storage actually reflects a typical blended data environment. In short, data access frequency definitely follows the pareto principle. The vast majority of data is write only, with very little of it being accessed frequently. For those generalized storage populations the pricing actually aligns.
Related is the issue of "stalled" throughput of dense storage like HDD. 4 years ago it was 4TB spindles at 5-7,200 RPM. We're up to 10TB per spindle these days. But rotational latency is still 5-7K rpm, limiting throughput to say 20-200MB/s depending on how much you like latency and queueing. As I recall HDDs are going to be up to ~20TB per spindle over the next few years. And still 5-7K rpm. Storage keeps getting cheaper, throughput is not.
And lastly its really important that youre not paying for traffic. Youre paying for your providers network. That's their datacenters, their backbone, their leased fiber investments, their hundreds of millions in capital. It's not something as simple as "oh I can get cogent for $0.2/mbs." It's an old post, but for idea see https://news.ycombinator.com/item?id=7479030.
"Egress to a different Google Cloud Platform service within the same region => No charge"
Also, it supports multi-region whereas S3 does not.
Love the direction Google Cloud is taking <3
It's the primary reason I've switched over to Datastore and App Engine.
edit: durability, not reliability.
Then ask yourself, "is the chance that someone made a mistake in designing the system which causes catastrophic data loss higher or lower than one in a billion?"
Yes, the number is total fantasy. The only actual meaning of the durability SLA is that if the durability SLA is violated, your cloud provider will give you some service credits. If the data is worth more than some service credits, this is small consolation.
In practice, if you absolutely must compare durability, what you are trying to do is compare the frequency of black swan events, but actually estimating the frequency of those events requires access to proprietary data which you don't have.
Both providers are undoubtedly storing the data in multiple locations on multiple types of storage systems with multiple layers of error coding.
After thinking about this issue earlier, the most likely data loss scenario in my opinion was "I get injured, the company hires someone incompetent who doesn't pay the bill for a year while I'm in the hospital."
With regards to your Asteroid Strike (And other disasters such as Riot, Insurrection, War, Hurricane, Earthquake, etc...) - these are all disclaimed in the agreement you sign with Amazon under a clause known as Force Majeure - which essential means, "Acts of God". The durability clause comes into effect under the normal course of business, not exceptional events. For those types of scenarios, you'll want to have a business continuity plan in place, not a durability formula on your storage service.
My point is that there is no point in taking that kind of durability guarantee into account, because there are much bigger threats to data storage. It's kind of like saying that driving by car is very safe if you don't get into an accident--while technically true, it's not a very useful fact.
But lets get back to that 400B object customer who's losing 40. How much does it cost to verify (resilver) all those bits, even annually? Is it even tenable?
No, that's absolutely not how to think about it. In short, it's a dangerous simplification and it's not at all representative of how failures work in a system like S3 or Google Cloud Storage. It does make good marketing copy, but if you are an engineer or a program lead you have a certain level of responsibility for understanding why cloud storage does not, in practice, give you eleven 9s of durability per object every year. (At a first approximation, you would at least expect the data loss to follow a Poisson distribution, but…)
Drilling down into the techical guts, S3 and Google Storage do not store separate objects in the lower storage layers, it's too inefficient. So below the S3 / GCS object API, you have stripes of data spread across multiple data centers with error encoding, along with a redundant copy on tape or optical media. Randal Munroe estimated Google data storage at 15 EB (https://what-if.xkcd.com/63/), so taking that number, let's suppose a stripe size of 100 MB (just picking a number out of the air that seems reasonable) and you get 1.5e11 stripes. Taking Amazon's 11 nines, that gives a loss of 1 stripe every 8 months.
So, all we've done so far is look at how a system would be implemented, and we've already completely destroyed the notion that you would expect 40 of 400B objects to disappear due to bit rot in any particular year. Supposing the objects are 10KB in average size, you might expect most years to lose no objects at all, and if you lose any you might lose 10,000 at the same time—and the entire extent of your recourse is to get a service credit from your cloud provider.
The gotcha is that the system simply isn't that reliable. First of all, engineers at Amazon and Google are constantly pushing new configuration and software updates to their stack. Some of these software updates can result in catastrophic data loss, and some of these errors will not get caught by canaries. "Catastrophic" might mean metadata corruption, it might mean the loss of many stripes all at the same time, but from the cloud provider's perspective they're still meeting SLA for most of their customers so most of their customers are happy. On top of that, you also have to take into account the possibility that a design flaw in the storage media would cause massive data loss across multiple data centers simultaneously, or other nightmare scenarios like that. Given that I've personally experienced data loss due to design flaws in storage media and I only have ever owned twenty hard drives or so in my life, you can imagine that a fleet with millions of hard drives presents some unique reliability and durability problems.
You can pretend that these configuration and programming errors are "unusual events" but the fact is that stripe loss for any reason is already an unusual event, and you might as well include the most probable cause of data loss in your model if you are going to model it at all.
So, what is the SLA? It's part of a contract. It defines when the contract is performed and when it is broken. It's also a piece of marketing and sales leverage. That's all. It's not a realistic or particularly useful description of how a system actually works—so the responsible engineers and program managers at companies which use cloud services are always asking themselves, "What happens if Amazon or Google violates their SLA? Will I lose my job?"
(A footnote: You don't need to verify the bits yourself, cloud providers will send you messages when your data is lost. If you want more durability then you go multi-cloud or buy a tape library.)
(Disclosure: I work at a company that provides cloud storage.)
So, perhaps another way of looking at this, is the 11 9s of reliability means that Amazon is providing a guarantee that they will lose at least that many objects, but, for all the reasons that you highlighted, there underlying redundancy mechanism means that there are all sorts of reasons why a catastrophic loss could result in many orders of magnitude lost of data, not only during exceptional events, but just under the normal course of business.
I would note that your anecdote about losing "data on a hard drive" isn't as relative at cloud storage levels, because one of the things that has been drilled into my mind by a colleague who works on a cloud system at scale, is they not only assume they will lose a single device, but they also scale so that they can lose (in order), a complete Rack, a complete PDU, and a complete Data Center - and still continue to provide availability to storage. That is, in the normal course of business, they plan on losing data centers, and continuing to provide full availability (albeit, at reduced durability in that event until they restore that data center). Google takes this up a notch, and provides availability in the event that an entire region is lost. Cautious companies can roll their own Business Continuity Plans on top of this as well.
What would be interesting, but unlikely, is for Amazon/Microsoft/Google to share with their customer what their actual loss of data was in the prior periods (Say, per year), and then provide a rolling graph of actual loss. Also useful (and almost guaranteed not to be available) - is what percentage of customers lost more than their SLA each month.
The passing comment of distribution of errors is important. However the "1 stripe is 100Mb, and objects are 10KB, ergo you'll lose 0 or 10,000 objects" bit is bizarro. I suspect youre letting your personal experience lead to assumptions that may not be true in other implementations.
[paraphrased] All storage classes are designed for 99.999999999% durability
As the previous replies to your comment mention, at that order of magnitude there are many other causes of data loss to worry about.
Disclaimer: work for an Alphabet company but not doing anything related to this.
http://blog.zachbjornson.com/2015/12/29/cloud-storage-perfor...
The issue with Glacier is the convoluted retrieval pricing. I understand they want to dissuade people from using it as a primary storage, but the potential for a surprise bill is hard to swallow.
Interesting how they still offer unlimited storage through their consumer Amazon Drive service.
Meanwhile other competitors (e.g. Nearline) promise to have your file available within seconds...
Disclosure: I work on Google Cloud and am a happy GCS customer.
Paying for storage, read and writes per GB and per request, rounded to 128kB blocks. WTF.
gsutil -m rsync -P -d -r -x "node_modules|tmp|..."-e ~/ gs://you-will-spend-too-much-time-thinking-of-a-clever-bucket-name
(although I should expand it a bit to better handle ignoring stuff by parsing .gitignore)
FQDNs to the rescue, I very much doubt anyone but me is going to call a bucket backups.ninjagiraffes.co.uk (and if someone does so now, well, good trolling)
if you own a domain, only people who pass domain ownership verification can make buckets with that domain suffix.
Combined with https://cloud.google.com/storage/docs/hosting-static-website it means you can easily host static websites by just populating appropriately named buckets with your domain suffix.
I think this is a very cool feature.
Of course Amazon Cloud Drive looks nice for people who wants to backup their media with little effort. I'm using a bit different approach, I'm backing up my data to a home server, but I need a reliable mirror in case of emergency.
[0] https://cloud.google.com/storage/archival/ - At the bottom
In fact, it's implied they believe Coldline is the biggest new thing they're doing, for it's the first thing in the announcement they discuss in detail. It's certainly got me interested.
"We may use, access, and retain Your Files in order to provide the Services to you, enforce the terms of the Agreement, and improve our services, and you give us all permissions we need to do so. These permissions include, for example, the rights to copy Your Files, modify Your Files to enable access in different formats, use information about Your Files to organize them on your behalf."
that's very scary if they can terminate my service and delete all my binary backups because they did not like binary files.
https://www.amazon.com/gp/help/customer/display.html/?nodeId...
For comparison, CrashPlan has a similar clause (see section 11)[0].
Ditto Backblaze (see "Our Rights")[1].
I wouldn't be surprised that all "unlimited" consumer products have similar clauses. If this is a concern for you, use a metered product.
[0] https://support.code42.com/Terms_And_Conditions/End_User_Lic...
For our Personal Backup product (fixed $5/month for as much data as you can keep on your laptop) your data is encrypted on your laptop BEFORE UPLOADING to Backblaze, and we have no idea what is inside your files and I assure you we don't want to know. (Look into setting your own private encryption key if you are worried at all about this.) In our 10 year history we have never once removed a customer's file because we objected to the file contents in that customer's Personal Backup because we simply don't know what is in the files.
For our other product line (B2 Cloud Storage for half a penny per GByte per month) the problem is you can have a "Private" bucket at which point we absolutely DO NOT care what you store. But if you have a "Public" bucket that is a public website. If you have public bucket serving up illegal content such as a phishing website or sharing bootleg songs or movies that you do not have legal rights to share, then we may shut you down (turn the bucket "Private" so nobody but you can get the data).
TL;DR - use our Personal Backup product or a "Private" bucket and store WHATEVER YOU WANT. But if you break the law serving up a website from Backblaze we have an obligation to shut you down.
It does local encryption before sending your data up to the cloud.
Notably Arq restores all macOS permissions/meta-data, and there is an open-source test kit to show whether such a program does so.
It also works on Windows.
I have no relationship with the company, other than as a happy user.
From what I read, you can still offload data from your local machine by choosing to "archive" a directory.
But thats an extra step and I haven't read how it handles conflicting directory that are created in the future after archiving the original.
My goal is just two off-site backups (AWS and Google).
ssh user@rsync.net gsutil cp mscdex.exe gs://my-bucket
ssh user@rsync.net gsutil rsync -d data gs://mybucket/data
Although it should be noted that, circa 2016, the cool kids are all backing up to rsync.net with borg[1][2] (the "holy grail of backup software") which limits the use cases of gsutil with an rsync.net account.HN Readers discount. Just email and ask.[3]
[1] https://borgbackup.readthedocs.io/en/stable/
[2] https://www.stavros.io/posts/holy-grail-backups/
[3] info@rsync.net
I really like the look of the service but I can't justify paying over 3x for it.
Well, our storage platform is online and random access so it would be inappropriate to compare it to either nearline or glacier.
The appropriate comparison is to S3 - or in this case, the multi-regional GCS option that this discussion points to.
In that case, our attic/borg pricing is 3 cents as compared to (roughly 3 cents) for s3 and 2 cents for GCS - and that assumes you use no bandwidth. Since we charge nothing for usage/bandwidth, the comparable prices from amazon/google would be slightly higher.
It sounds like a steal to me and every day plenty of people agree enough to commence using our services.
Also, unrelated, you know you can just drive up to an rsync.net location and get your data - even if the Internet is crippled.[1]
That would be true but you stated "circa 2016, the cool kids are all backing up to rsync.net with borg", so comparing your service pricing as a backup product to Nearline is appropriate I believe.
Don't get me wrong, your service definitely has a lot of good applications and the free bandwidth is quite great but for those of us that just want somewhere to park a whole pile of data we aren't going to touch much, Rsync.net is quite expensive.
Also a downside of rsync.net vs. AWS/GCS etc. is that you can't direct users' HTTP requests to rsync.net. That dampens the usefulness of online storage and unlimited bandwidth, since it essentially limits it to servers under our own control. If I could run HTTP off it directly, my feelings would be very very different.
But, I would like to quibble with the "online" comment about Nearline or Coldline: there's no access penalty (now). We kind of quietly announced it in June or so, but Kirill the PM linked to it again. I had understood the "default" rsync.net choice to be in a single location, so comparing to our multi-regional (as opposed to Regional, Nearline or Coldline) seems incongruous.
Disclosure: I work on Google Cloud.
(0.7 * / 100) * 3000 = $21 / month * 12 = $252 / yr
For that price, I'm better off just paying $60 / yr for CrashPlan, which is unlimited storage. I'd say that applies to anyone who has a terabyte or more of stuff to back up.
I should probably "test" doing a recovery from CrashPlan at some point, but I fear that even getting a third of it (1TB) would probably trip some sort of overcharge from my ISP.
Plus the price isn't too bad.
Echoing another comment, I'm not sure if this is a big a danger as a programmer's or operator's potential to wipe out all your data no matter how many datacenters it's located in, but....
I have data on my laptop, an external disk that I keep at home, and BackBlaze. The likelihood of all three failing in the same week so that I can't restore from any single copy is low.
I work at Backblaze, so I'm biased. :-) But the use case is important, meaning I wouldn't group together all "cloud storage solution" as one use case.
Backblaze does not have compute at all (like EC2) so Backblaze is a terrible choice for a company that will spend a lot of cycles analyzing the data stored in the cloud or doing compute on the data stored in the cloud. That is a much larger issue than the one data center for that use case.
On the other hand, it is my profound belief that for long term durability of data you should have AT LEAST three copies of the data stored by profoundly different vendors. Hopefully different file systems and stored by software written by different programmers so the same bug that affects one copy won't affect the other copies. Hopefully stored at separate physical locations. Put a different way: the only way to get more reliable than data stored at Amazon Glacier is to have one copy in Amazon Glacier, another copy on Google Coldline, and another copy in Backblaze. In that use case, the single data center of Backblaze is obviously a non-issue.
I've found that you can alleviate some of the headache by assigning the disks unique drive names. The reason for the deletion is that you might plug in your external HDD that gets mounted as drive F. You let it finish backing up, then disconnect it. If you later plug in a tiny flash drive that also gets mounted as F:, CrashPlan will see the contents of that drive and assume that you deleted everything from the previous drive and accordingly mark is as deleted in the backup.
If you mount the external HDD as drive Z:, however, when you disconnect it, CrashPlan will simply show the device as "Missing", but won't delete it unless something else gets mounted as Z while CrashPlan is running.
- Multi-regional is what we used to call Standard (highlighting that it's always been awesome and replicated)
- Regional / DRA collapse into one
- New offering: Coldline
- Massive price cut on operations
- Per-object storage class, plus lifecycle (e.g., automatically go from Nearline to Coldline after 60d)
Disclosure: I work on Google Cloud (and use Nearline at home!).
[Edit: Formatting. I always forget to put two newlines]
Yes, I know unrelated, but I like GCS much more than S3, and I can't use it since by "network" effect caused by GPUs eveyrthing have to run on AWS.
Disclosure: I work for Google Cloud (contact info in profile).
[1] https://cloud.google.com/compute/docs/disks/gcs-buckets
My only complaint is Google Cloud persistent SSD disk pricing at $0.17 per GB / month. AWS EBS SSD is noticeably lower at $0.12 per GB / month.
There are two outages there, one that only affected a few projects and one that affected only service in the central US (we have regions over much of the world).
Anecdotally, from watching and working with other services internally I don't think most of our outages affect all regions. We actually spend a significant amount of engineering effort ensuring that we're as decoupled as possible.
Disclaimer: I'm an Engineering Manager on Google Cloud Storage.
Other than that, the general approach is to minimize global control planes and dependencies in our software stack. In the case of GCS, we do have a single namespace which means we need to look up the locations of data early in the request. Once we know the locations of data we can route the request to the right datacenter to serve it. That global location table is highly replicated and cached, of course.
When outages happen, most are caused by changes to the stack, so we also are careful to roll out code or configuration slowly and carefully, slowly increasing the blast radius after it's been proven safe. For example rolling out new binaries first to a few canary instances in one zone, then to a few instances in many regions, then to a full region, then to the world, all spread over a few days.
Disclaimer: I'm an Engineering Manager on Google Cloud Storage.
Not sure what happened there, but think a deploy should be rolled out to 1 availability zone at a time for hosted things (like load balancer)
Much like with GCS (and any service at Google), the most common source of outages is a rollout. While we strive to offer zonal and regional services where appropriate, some like GCS and PubSub do have real value as "global" APIs. Trust me though, hurstdog spends a lot of his time struggling with that balance ;).
Disclosure: I work on Google Cloud.
0.2ct per Gb.
Or do we upload to the European region and it gets replicated to the US automagically?
And the whole nickel and diming charges that force users to needlessly seperate their computing needs into storage, compute, memory, bandwidth, reads, writes and what not are actually forcing considerable complexity on users.
This cannot be brushed aside as a good model for cloud when valuable user time is being wasted grappling with needless intricacies that have no reason not be flat.
End users cannot buy bandwidth at low bulk rates and for things like backups may be faced with considerable overcharges on their local connections.
Besides capricious pricing changes, their support is notoriously bad, even for paid apps for business.
They don't seem to understand business customers needs nearly as well as Microsoft/Amazon, from what I've seen.
Really not who I want running my business's server infrastructure, even if it's completely free.
In fact the whole thing could do with being rewritten in a simple bullet-pointed skimmable form. That way I wouldn't have to come to HN comments to find out what the hell it's all about. Is it too early in the day for a drink?
1. It was 4 in the afternoon here when I posted that - a touch too late for breakfast even if it was a bit early for a drink
2. I quite like words when they are strung together well ;-)
I still would like to hear someone defend that as a piece of marketing copy. I've re-read it and stand by my earlier diatribe.