Here's the specific example https://cloud.google.com/billing/docs/how-to/notify#cap_disa...
Here's the specific example https://cloud.google.com/billing/docs/how-to/notify#cap_disa...
In our case, we racked up a $10000 bill on BigQuery in ~6 hours, when a job was failing and auto-retrying.
We had set up every alert correctly and our reaction time was about 5 minutes (about $100 of usage, no big deal). So how did we get a $5000 bill? Google's alert was 6 hours late (according to them, this was root-caused to us, because we were submitting jobs continuously). They pointed to their TOS and said they don't guarantee on-time delivery of the alert.
I had to write up a blog post with fancy graphs and prepare it for social media before they finally agreed to eat the bill.
EDIT:
Link (see blue info box a a bit below the anchor on which the page is opened):
https://cloud.google.com/billing/docs/how-to/notify#cap_disa...
But if 1,000,000% lower doesn't work ($7 vs $70k) then...?
you misunderstand the intent of this - you basically set this. even if it fails (because messages are delayed), google will refund.
This has happened to us before - they do a refund - since you had set the limits correctly. In general, they are not super assholes. I actually dont know a case, where they have refused to refund.
AWS is better here - since GCP doesnt have a support dashboard. So the "chasing them" experience is much worse.
This looks like it has the same problems as the post, because it also relies on those budget alerts that can happen a long while after you've exceeded them.
"Resources [...] might be irretrievably deleted."
Also it's not automatic, you have to manually write code to do it, and test it, and make sure not to break it.
A reasonable implementation of this feature would be built into the console, guarantee a maximum spend, not require writing your own fallible code, and provide an option to preserve storage (at normal cost) so that all your data isn't deleted when your compute/API stuff is shut down.
- hard limits caused downtime more often than they prevent these blog posts
- hard limits were inconsistently enforced, even within GAE
- platform wide quota notifications were implemented (reached "GA"), leaving the question of "how a developer wants to handle this" to the developer, not the platform
- maintenance burden
The "I bankrupted my startup by running tests in an infinite loop" blog posts happen ~once a year, while the number of customers (including internal teams!) who inadvertently went down because of this quota was staggering. I feel like I used to see one a week, at least. Most often someone on the team was like "oh I'm going to turn this down to zero because we don't want to spend any money during development", never told anyone, and then they go live and they forgot to turn the knob back up (or didn't properly estimate traffic/costs and set it too low).
I can tell you it hurts revenue a lot more when a large customer goes down for 15 minutes due to quota issues and their usage drops to zero (both in terms of revenue and customer credibility) vs when tiny developer accidentally blows through 10k in a month and we refund it (since, obviously, the providers cost is a lot less than that).
When I also think of the fact that Google tied it to requiring a credit card for almost every single transaction even if it is free gives the impression that it is for financial purposes (aka a way to get more out of developers or those who might be free-loading on the free tier of App Engine)