Trying to get a bunch of GPUs on Google Cloud
twitter.com
twitter.com
First try I'm allowed to spec up a beefy server with AVX-512 and click "create", but it never actually starts. Turns out the default CPU limits are tiny - I need to apply to bump two separate limits and wait for that to go through.
Second try I'm actually able to spin up a couple of servers. I set the compute job running and go for lunch. An hour later both are terminated and I get a message about "violation of our Acceptable Use Policy".
There is a link that says crypto mining is forbidden on free instances. I send a reply informing them that a) these are not free instances, and b) I am not mining crypto.
I get a reply back saying more or less "Yes it says free instances but why would you expect anything we write in the documentation to be correct. Also you might think you weren't mining crypto but actually your VM was probably compromised because no one would buy compute instances and actually run compute jobs on them. We're not giving you back access to your instances, maybe delete them and create new ones lol."
I actually find it an easier product to get started with than AWS, but man they really really don't want your money. When you are actively trying to prevent your customers from scaling up resources you are failing at being a cloud.
People need to do their due diligence.
Basically it improves the predictability of resource usages by limiting outlier spikes.
The needs of those wanting to compute should absolutely outweigh the needs of those wanting to not compute.
Then they wouldn't configure overcharge limit, would they?
Allow people to pre-pay (deposit) an amount of their choice and set an account flag "Do NOT exceed account balance".
I mined Ethereum in azure during the height of the crypto boom, and looking at my statistics of vms i was able to get i am pretty sure some other people did to.
On 5k$ spend on azure i got round about 7k$ in Etherum when sold the next day, this includes german Sales tax and some loses i made starting out when not accounting for the german sales tax early on.
I think the whole timeline where this worked was about one and a half weeks(Germany, if you played around with exchange rates and sales tax potentially a bit longer), but it worked pretty well.
Azure was also perfectly fine with me doing it.
Maybe a flag you could set in your account that says "I know what I'm doing and I promise not to ask for refunds for my mistakes." Obviously if the money is lost due to a Google screwup then yeah, I'm going to want that back, but if I forgot to turn off a few thousand max GPU instances over the weekend then that's on me. Maybe combine that flag with a prepayment system? And if you blow out your account then it halts your instances until you top it off, hopefully with a bunch of text messages and emails to the admin.
I suspect it's mostly Google being Google and horribly at customer service, even when the customers are waving money at them.
Please do better Google.
I've had three software engineer jobs, learned this lesson on the first one and watched others get tripped up by it in the other two. Nobody ever believes me when they say "We will just spin up a instance for every file in the directory" or similar and I respond with "No you won't, they aren't going to give you a thousand instances". This is a lesson that can only be learned first hand.
AWS regularly gives me a few hundred instances at a moments notice.
Sure, they’re given back just as quickly after the jobs are done, but it’s certainly not impossible.
It's basically an auction market. You set your max price to bid on instances, and will get them as long nobody else bids higher.
More details here:
https://aws.amazon.com/blogs/compute/new-amazon-ec2-spot-pri...
You set your max price, up to the on-demand price, and then you'll get your VM... or not. If you do get the VM, you still don't pay the full price in most circumstances unless demand is very high.
Or the "with enough GPUs I can quickly launder a large pile of cash I scammed off the elderly/businesses/etc. into crypto currency before this account gets frozen" problem.
Since they only accept credit, this is exposing Digital Ocean to the risk of handing out cash-equivalents for potentially bad credit. This "forces" them to be aggressive against abuse. Unfortunately, heuristics aren't perfect, so they also occasionally execute a hapless business by accident.
John C made the same comment that I did. Other than outright crime (child porn, etc...), there ought not be any limits on instance usage in clouds, provided that customers pay cash.
Heck, let them crypto mine if they wish, as long as they pay cash! Who cares? It's just computation.
"Just take cash" isn't an option.
Those are much harder to claw back, particularly if you have a signed agreement from the customer saying "We want to pre-pay for $50k in compute resources. We understand that once the pre-paid amount is used, it cannot be refunded."
I'm sure GCP and any other legitimate cloud provider could also take wire transfers for sizable payments, without issue.
GPU prices didn't come down because crypto crashed, they came down because ETH switched to proof of stake making mining on GPUs obsolete. Anyone could predict when this would happen because it was pre announced.
So we built it and the moment we ran the crash test (we tried to go up to 1000 simultanious instances), it turned out eu-west-2 had just 48 instances available. We followed up and AWS confirmed that we just capped their region capacity for those instances. They didn't raise any quotas and suggested we use other regions to bypass the cap.
I can imagine they don’t keep a few thousand really expensive instances around (in every region/availability zone).
Google-sized megacorps should not be able to impede innovation and economic activity like this. We're all losing because of it.
Besides, their incompetence allows companies like Lambda to fill the niche, and probably do a better job of it than Google do.
You're just restating the current status quo as if that some how refutes the notion that businesses can be regulated, and that businesses of this size should be regulated in a manner that promotes competition.
Bad producers can only impede as much activity as stupid consumers give them.
It makes much more sense from every perspective to buy them and put them in your bedroom.
Cloud GPUs are slow, VERY expensive and not even available.
You shouldn't need the superstar power of being John Carmack to get access to computing resources.
The promise of the cloud is for the most part, sssnake oil.
One of the big deals about data center GPUs is that they have a lot of RAM, more than any consumer card.
If you're not running something that needs NCCL et al, yeah, go run something under your desk.
Anecdotally, having a single high powered modern GPU in my office generates enough heat that it is in part why I keep a temperature monitor on my desk to ensure I'm not cooking myself. Having more than one gets far more interesting.
The problem is 1) they won't just let you pay them and 2) even if you could, they still won't just let you use some GPUs.
Similarly everyone was convinced Gmail was an April Fool's joke when they gave you 1 GB of storage when everyone else just gave you 3 MB.
Can you really say that across YouTube, Search, Gmail, Drive, Cloud, Open Source Contributions (Kubernetes, Tensorflow, etc) and so on they don’t add any real value anywhere at all?
https://support.google.com/googleplay/android-developer/answ...