What went wrong with UniSuper and Google Cloud?
danielcompton.net
danielcompton.net
https://github.com/hashicorp/terraform-provider-google/blob/...
While we're at it, it also looks like the provider couldn't provision stretched clusters at all until mid-April. I don't know what I think this means for the theory presented in the article. Maybe Uni was new to TF (or even actively onboarding) and paid the beginner's tax? TF is great at turning beginner mistakes into "you deleted your infra." It's an uncomfortable amount of speculation, but it's plausible.
Relevant discussion is on https://github.com/GoogleCloudPlatform/magic-modules/pull/10... and relevant code changes are on https://github.com/hashicorp/terraform-provider-google/pull/...
Fwiw -- and I know there are plenty of googlers here -- the OMG isn't locked down and it was legitimately a unique bug. I appreciate the forensic analysis of public statements that the author of this post strung together, but it doesn't really advance the conversation.
Used to work on GCE, but as the article mentioned, there are a lot of safeguards built in over there to prevent you from accidentally deleting things and account wipeout has a grace period to prevent a lapse in billing from deleting anything.
It could be valuable in the sense that scenario 2 has me shrugging my shoulders, but scenario 1 goes on the ever growing list of "reasons to avoid Google _anything_".
Superannuation funds are often industry-aligned, and some are managed by trade unions. UniSuper is the industry super fund for the university sector.
EDIT: As others below have pointed out, yes, I meant a 401(k).
Will this well-publicized event materially alter cloud spend (e.g., cross-cloud replication or backups)?
This seems like an amazing time for AWS and Azure to come out with statements how they prevent accidents like this and why single staff members aren't capable of nuking a large company's cloud account.
That's exactly the opposite of what he said:
> [...] putting out a competing statement blaming or contradicting your customer is a bad look with that customer and with all future customers.
In Google’s case, Google already has a reputation for poor-to-miserable support and for arbitrarily removing its users’ access to their data. (Heck, another incident of this sort was on HN today.) GCP gets considerably revenue from very large users, and those users and their decision makers will do just fine moving from GCP (which is not #1!) to AWS. [0]
Imagine an aircraft manufacturer or regulator in a similar situation. A plane crashed, and it behaved as documented or intended given pilot inputs. But it still crashed, and the factors causing it should be identified and appropriate improvements, if any, should be implemented.
[0] Some of them will grumble about how AWS’s user experience is dramatically worse than GCP’s in a lot of very obvious ways, but the overall comparison tilts strongly toward AWS here. Sure, AWS makes it miserable to configure Organizations and Accounts in line with best practices, but Google might arbitrarily delete a project/account/whatever! Google should get out ahead of this.
Absolutely. Not only should they get ahead, the should be providing detailed technical details as to what went wrong and how they are ensuring that it won’t happen again. The lack of details here just screams unseriousness.
I don't buy it. Google could explain what the customer did to shoot itself in the foot and how GCE will modify the UX to hide the footgun. Instead, it just says that the customer had a unique configuration. That sounds like a GCE bug.
As I've said many times before, Google doesn't use GCE for critical applications, so you shouldn't either. Amazon and Microsoft use AWS and Azure respectively, but you should carefully consider which cloud services they offer are also used internally.
At the end of the day terraform can have a bug. You really want to control blast radius with permissions. Makes me wonder if the GCP VMWare integration is a boundary that doesn't expose granular permissions.
If it was operator error with terraform that should set off alarm bells through the industry. Who else is one fat finger away from total annihilation.
"Hey terraform just output a wall of text, it wants to know whether or not to proceed."
"That's what it does mate. Let it do its thing, she'll be right."
Even if they do fix this particular issue, I see this as a warning to be very very reluctant to adopt these integrations. A product level integration (eg internally hosted github) seems ok. But infra level integrations can have ugly failure modes.
(I guess an investigative journalist could start interviewing staff promising anonymity)
https://www.cnbc.com/amp/2024/05/01/google-cuts-hundreds-of-...
From the article:
>The Core unit is responsible for building the technical foundation behind the company's flagship products and for protecting users' online safety, according to Google's website. Core teams include key technical units from information technology, its Python developer team, technical infrastructure, security foundation, app platforms, core developers, and various engineering roles.
Basically, Sundar Pichai is taking the McDonnell Douglas approach to engineering and just deciding to coast on Google’s previous engineering.
The danger is that while you may have short-term stock returns, you destroy the engineering culture and it is only a matter of time before the doors blow off mid flight like a 737.
You mean Boeing?