We use AWS GPU machines for our CI but for serious ML training workloads we use GCP L4 GPU instances. Even in GCP we couldn’t provision or A or H100 (our quota itself is just 1 GPU of these instances) but we’ve never had issues provisioning L4 GPU’s and I think that’s enough for smaller not LLM Scale Models. For LLM scale startups, it’s tough provisioning GPU’s even if you have money.
(We’re based in Bay Area btw)
https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-capa...
Really feels like if you need accelerated compute GCP is the better option these days. At least there you can rewrite in Jax if it comes down to it and opt for TPU's.
The trick is your strategy and what's in place to survive and turn into those waves when they come in the end. Seems like you did this as you had more success elsewhere.
Quotas in this case refers to needing to request access to GPUs. Those reviews takes minutes or days depending on your relationship, and are often final.
Its possible that none of the cloud providers offering GPUs will ever allow you quotas high enough to scale to profitable margins, and there's nothing you can do about it except try and host in house (which would be mad). But the GPUs being on cloud makes that very uneconomical and they are in high demand.
This was all as true in 2019 or 2022 (more so with TPUs in 2020 I should say) as it is now. It's not a ChatGPT thing.
2013 is approximately when university and national computing units became relatively useless compared to GPU cloud compute. This stopped a wave of would-be university spin offs from having access to sufficient compute to compete.
Where?