Cloud TPU v5p and AI Hypercomputer
cloud.google.com
cloud.google.com
They take weeks to respond to anything, they change their minds constantly, you can never trust anything anyone says, their internal communication is a complete disaster, and someone recently told me they outsource a lot of their GCP personnel.
We went back to AWS and had a whole fleet of GPUs up and running within the week.
This is my 3rd extremely bad experience on Google Cloud. My last unicorn startup had several GCP-caused P0 production issues. They would update something internally with no announcement to customers and our production workloads would completely break out of the blue. It would usually take them days to weeks to fix it even with us spending tens of millions with them and calling support constantly. Everyone at our company was baffled at how bad the experience was compared to every other cloud provider.
I would not put anything serious there and I would never partner with GCP again.
How frequently are you engaging your account reps? You should be able to get the ear of a PM within 48 hours in most cases.
That makes sense because GCP doesn't have all that much capacity for scale-out GPU compute (they offer TPUs for that).
I posted this in another comment, but the default provisioning for V4 and V5 TPUs is ZERO. They don't tell you this anywhere. So when we'd try to allocate V5 TPUs on our GCP account, it would just fail with a generic error and a huge error number that led to nothing in a search.
So I reach out to our GCP rep we had been working with. After about 10 days and 3 follow-up emails/calls that went unanswered, she replies, "you have to fill out this form." I click on it and it's a Google Form. The same type of Google Form you and I can make.
I submit it. To this day I have heard nothing back. I reach out to various executives at GCP we had been talking with. They said, "you have to fill out a form." I tell them we did, and they say, "oh, it's usually pretty fast." I heard this so many times from so many people. It seems that not a single person who actually works in sales or account management knows how this process works.
When I finally got a response, they told me that "my account has no billing account associated with it." I showed them yes we do, they replied that they were trying to provision for the wrong account. It took another couple weeks to get a follow-up response.
Luckily they eventually connected us with one of their partnered consultants who was finally able to help us, but by then we just decided to go back to GPUs on other platforms because it was such a miserable experience and in that time, all of our providers came through with the volume we needed.
At my previous company, we even had an announced partnership with Google and did lots of co-marketing. That doesn't mean much when a GCP engineer changes something in GCP and breaks your production.
I wish it worked out. I liked working in GCP outside of these problems. I really like GKE.
Also those are examples, or details, even. Not references.
I think if you search across the internet you find a lot more GCP horror stories than AWS ones. Considering the market ownership of the two platforms, I'd expect the opposite to be true.
Giventhe worldwide shortage of AI compute, and extreme demand, How is Amazon able to onboard new clients and have fleets of GPUs ready to go?
Does Azure/AWS/Google have a huge collection of AI compute that is just waiting?
Are these machines just the ones running the custom Amazon or Google accelerator chips?
We also need GPUs for two things: training and inference. Our training needs a bunch of A100/H100 or large TPU nodes. But our final models are all pretty small (we're in ASR, TTS, translation, etc) relative to most commercial LLMs.
So we run a large amount of inference on GPUs like T4s, L4s, etc that are much easier to obtain.
Our intent was to try to find a unified, scalable platform that we could train/infer on that didn't have the cold boot problem AWS GPU instances have when our inference needs go beyond our dedicated hardware. With AWS the P99 cold boot times were over 5 minutes which means we'd have to queue inference and we'd end up paying for 5 extra minutes of GPU time. This isn't that big of a deal, but when you're spinning thousands of GPUs up and down it adds up to a ton of money every month.
Since our models require a lot less kernels than a lot of modern LLMs, TPUs seemed like a really good fit. But after dealing with GCP the last few months we abandoned our plans to port our model to JAX and decided to stick with GPUs for now.
Fuck yooooouuuuuuuu. Fuck you, fuck you, Fuck You. Drop whatever you are doing because it’s not important. What is important is OUR time. It’s costing us time and money to support our shit, and we’re tired of it, so we’re not going to support it anymore. So drop your fucking plans and go start digging through our shitty documentation, begging for scraps on forums, and oh by the way, our new shit is COMPLETELY different from the old shit, because well, we fucked that design up pretty bad, heh, but hey, that’s YOUR problem, not our problem.
We remain committed as always to ensuring everything you write will be unusable within 1 year.
Please go fuck yourself,
Google Cloud Platform
Source: https://steve-yegge.medium.com/dear-google-cloud-your-deprec...
Whilst free GCP support is utter rubbish you will get several Googglers including direct line of comms to product management if you’re spending that much.
I don’t believe your claims
At my previous company we had dedicated account managers of course. You are completely correct that we had line of communication to engineers working on the problem. That just doesn't mean very much when they don't know how to fix the problem quickly.
At that company we were multi-cloud (AWS, GCP, Azure, Rackspace, Hetzner) and had even more workloads on AWS. Never had a problem as bad as the problems we had on GCP.
How do you get access? Not by talking to anyone at Google, but by filling out a Google Form. Yes, a Google Form.
I spent many hours talking to everyone from support to executives at GCP. This is not an exaggeration. They all told me to fill out "the form." They said it "usually goes pretty fast." They don't even know the people or the department who approves it.
It took us months to get provisioned. By that time my company completely lost interest as we cannot move that slow in the most competitive space in technological history.
Also one last thing, we never even got notified once we were provisioned! There's no response to your form, you just have to keep checking.
Just absolutely wild amateur stuff.
Before long, the form is the 'API', but both sides are using automation and the Google Forms backend is starting to rate limit responses and someone is forced to develop a proper API, try to track down all the form users, and try to persuade them all to migrate to the new API.
Busywork.
* Like @leetharris, credits were pulled/not distributed, what was promised had to be cajoled out with most of it being sent to some weird SaaS product that we'll never use
* The GCP rep literally ghosted us halfway through the month where it when we had some expiring credits and were in the middle of training
* Not that the credits mattered, our quota requests for lifting GPU or TPU was rejected twice. It was impossible to get any GPUs that were within our credits, even writing a script to try to look for machines for weeks didn't work.
* Right after the credits expired, suddenly our last quota request, which was hanging around for weeks was approved. I assume they have an internal system setup to do that, but like we literally couldn't pay for GCP if we wanted to.
* Also, GCP rates are like 2-4X the market rate. Like you can get an H100-80 from Runpod (and actually get one) for what GCP charges for an A100-40.
Basically, the lesson learned was that no one should ever depend on GCP unless your time is worthless and you're not serious about getting any work done. They can go suck eggs.
Kelsey Hightower told me at a GopherCon (many years ago) that Google doesn't run any internal workloads on third-party GPUs mainly because it costs significantly more (b/c cooling iirc), though they are happy to help you run your workloads on such GPUs.
> AI is all about how much compute dollars it can generate for the cloud providers.
If the providers wanted to extract more money they would not create custom hardware which reduces overall costs and prices to users.
I would argue that this is actually more about ensuring NVidia doesn't have a monopoly on hardware and alleviates us from having to pay for Nvidia profits through our cloud providers.
Azure is going down the same path here: https://www.theverge.com/2023/11/15/23960345/microsoft-cpu-g...
Extracting money is about margins, not revenues. If they reduce your costs (and their revenue) by 20% with a TPU, but they can produce TPUs for 50% less than buying gear from Nvidia, it's still a profitable move.
The "extracting" word typically comes with abusive connotations when used in the context of money, which doesn't feel like the right word for the win-win outcomes imho
He also shares what he has learned for free, rather than putting paywalls in front of his content, which is quite rare these days.
Higher efficiency results in greater utilization.
I have no doubt that Nvidia has extensive optimizations to get SOTA performance, but I am curious what is attainable off the shelf. If you could design a 5nm chip, would it be possible to hit 15% of a NVidia chip? Significantly more?
Of course, there is more to a GPU than just the matrix multiplication, but I am wondering how much effort it would take to get something off the ground for the well financed organization. Presumably China is actively finding such efforts.
I'm not usually one to point out redundancies like this but this one seems egregious.
Large Large Language Model Models
Small Large Language Model Models
Papa Bigfoot was the biggest Bigfoot of them all. Mommy Bigfoot wasn't as big as her husband, but she was still a bigger Bigfoot than her daughter Little Bigfoot, who was the smallest Bigfoot of the family.
One day Little Bigfoot slipped in a stream and hurt her foot.
The little Little Bigfoot foot hurt so much and she cried a lot
[1] https://cloud.google.com/blog/products/compute/the-worlds-la...
https://www.theinformation.com/articles/to-reduce-ai-costs-g...
Edit: As the commenter below points out, "block floating point" is the common name, not "batch floating point."
Datatypes are really tricky. Hardware designers tend to be conservative in my experience, and don't want to waste die space on things that might not be useful.
Edit - The original: https://www.abebooks.com/first-edition/Rounding-Errors-Algeb...
Google is so fucking lame these days