Google Cloud Platform Tokyo region now open for business
blog.google
blog.google
This is not just a theoretical issue. In the past week we've been doing a bit more than 5 request/sec to Google Cloud Storage and according to NewRelic the average response time was 8 seconds! I.e. the service has been down and not been responding at all for large periods of time. I've been in contact with their support team and they've refused to reimburse us anything.
If not, drop me a line at jani at google dot com with a reference to your support case and I'll be happy to take a second look. (Yes, I work in Google Cloud Support.)
That said, we have very recently (as in, late October [1]) introduced a new pricing model for GCS with the explicit goal of reducing latency, and the SLA may be due for an update accordingly. I'll look into this.
[1] https://cloudplatform.googleblog.com/2016/10/introducing-Col...
Also, the HTTP 500 thing is specific to GCS only, other services like GCE [2] define downtime more broadly as "loss of external connectivity or persistent disk access".
Don't you agree that it's odd to only include HTTP 500 errors in the error rate? Let's say someone hacks your DNS servers and points storage.googleapis.com to 127.0.0.1. Then the entire service would be down completely but according to your SLA you'd have 100% up time.
If a mobile app can't connect to GCS it could be that GCS is down - but more likely the user just has a weak signal.
Unfortunately these things can occur in the darndest of places, as a bug in Google CLoud, an incident at GCS, or maybe even in your monitoring stack. I would encourage you to hold off judgement until root cause is identified.
One assurance I can make is that Google SRE monitors these things very carefully 24/7, and such levels of latency in the service would be treated as an incident. So it's likely something else is going on.
(work at Google Cloud, but not on GCS or support)
So Google claims that it does not have control over its own DNS servers and is therefore not to blame if the DNS is pointing to the wrong IP. Not very reassuring.
For many SaaS businesses the service credits are quite useless, because you provide so much value on top of the cloud services you purchase. You pay $1 for cloud and charge $50 from your customer for your app. If cloud is down, you get $0.10 as credits and need to credit $5 for your own customer (in good case).
(I'm not blaming the cloud providers for this. If they would offer better terms, they would need to anyways transfer the risk to their customers and significantly raise the prices or take the risk of going bankrupt in case of major problems).
No SLA I've seen guarantees a time for full body because that time fluctuates too much with both the size of the object and the current state of the internet. The new-ish refresh of the GCS lineup of services says you get sub-second access, but that has to be time to first byte, and I have a hunch that NewRelic shows you time to last byte.
If my assumptions are accurate, I would say the data you get from NewRelic does not warrant reimbursement from Google, though I might side with you if all of your objects are tiny.
Unfortunately the problem is still occurring after I got this message (although less frequently).
(In my defense, I believe the main product pages were out of date when I posted.)