If you want on-prem, wait a few months. The supply of 5000 series (probably announced at CES in a few days) should push more 4000 on the market and, maybe, for a bit, over-supply and push the price down.
Nvidia stopped manufacturing the 4000 a few months ago because they don't have endless factories. Those resources were reallocated to 5000 series and thus pushed the price for the 4000 up to the ridiculous place it is now (about $2,000 on ebay)
I think the current appetite for crypto and ai is big enough to consume all 4000 and 5000 series cards to a point of scarcity (even 3090s are still fetching about $1000) but there should be a window where things aren't crazy expensive coming up.
There's no evidence supply will continually outstrip demand unless something unusual happens.
It's probably somewhere between 12months-never depending on how the market shakes out. Maybe 2 years is a good idea ... really, if power is cheap/free and the machine is on and idle then it's free money - that's the way to look at it.
https://cloud.vast.ai/host/setup
There's a lot of competition in the "airbnb gpu" so if you don't like us, the number is around 12 or so globally. We're probably either #2 or #3. Companies don't really disclose these things so it's hard to know.
Some people probably list on more than one platform. There may be some host management software somewhere that helps with that. I haven't actually checked.
I'd be happy to talk more about these privately. Some are better than others and I've got no interest posting less than charitable things about our competitors publicly, regardless of how accurate I think it is. My email is in my profile.
We aim at $1200/y for 3090, so around a year given descent electricity prices.
Highly recommend setting a lower power limit (usually 250W for 3090).
The industry is full of effectively "imitation companies" right now. For instance, runpod, quickpod, simplepod and clore are the ones cloning us at vast right now.
We see them in our discord, they try to snipe away customers, get in our comment threads on reddit and twitter with self-promotes, clone our features ... this is the ferocious wild west days of this industry. I've even gotten personal emails from a few who I guess scanned their database looking for registration addresses from other companies in the space.
There's even companies like primeintellect which are trying to become the market of markets - but they have their own program - it's clearly a play to snipe other customers by funneling them through some interface where they'll eventually push out the other companies and promote their own instances.
Then there's interesting insider hype players with their own infra like sfcompute who are trying to pretend like they invented interruptible instances and somehow get a bunch of people treating them like they're innovators. The resellable contracts they talk about are a pretty common feature and especially from the host's programmatic command line controller, it's just usually tucked deep in the documentation. They're doing effectively a re-prioritization play.
I guess my angle is "highest integrity possible". It's certainly a gamble - scammy companies sometimes capture a market then become unscammy - I'll hold my tongue but there's plenty of examples.
It's interesting times.
I guess what I’m missing is, what’s scammy about them?
even in the web3 space, AI gpu compute markets are oversaturated
but why is an end user supposed to case about the user acquisition strategy?
if they’re cheaper, more profitable for the gpu owner, or solving a need better, that’s all that matters
You can multi sell a machine, use qemu to lie about the hardware, have hidden fees... there's a bunch of hustle
> AI gpu compute markets are oversaturated
This is not the case. We see a moving average of over 90% utilization of our network. There's a lot of players, but the demand is outstripping supply
> why is an end user supposed to case about the user acquisition strategy?
Well hn is founder/insider talk but for a more direct answer, more legit institutions get higher retention and easier customers.
We're a two sided marketplace so we need to create a platform where people see integrity.
There's also the hypocrisy of complaining about competitors jumping in on "their threads" in a comment on a competitor thread.
Yes, this comment of yours is highly unethical.
4090s are too small for training and you'll have to write your own suboptimal batching.
Unless you value the learning, it'd be better to rent GPUs in the cloud for training.
This might pull you down a path towards distilling and quantizing models, for instance.
I’m glad to know
But pricing is okay-ish, have a look at Geohot's Tinybox for turnkey solutions.
In Deep Learning it depends on your sharding strategy.