I'm working on the new one!
of course we can run any model if quantize it enough. but I think the OP was talking about the unquantized version.
You can do it via `-ot ".ffn_.*_exps.=CPU"`
Not affiliated, just a (mostly) happy user, although don't trust the bandwidth numbers, lots of variance (not surprising though, it is a user-to-user marketplace).
I don't even know how these Vast servers make money because there is no way you can ever pay off your hardware from the pennies you're getting.
An alternative is to use serverless GPU or LLM providers which abstract some of this for you, albeit at a higher cost and slow starts when you first use your model for some time.
Not for the GPU poor, to be sure.
For the hobby version you would presumably buy a used server and a used GPU. DDR4 ECC Ram can be had for a little over $1/GB, so you could probably build the whole thing for around $2k
Mobo was some kind of mining rig from AliExpress for less than $100. GPU is an inexpensive NVIDIA TESLA card that I 3D printed a shroud for (added fans). Power supply a cheap 2000 Watt Dell server PS off eBay....
[1] https://bsky.app/profile/engineersneedart.com/post/3lmg4kiz4...
There's already a 685B parameter DeepSeek V3 for free there.
For example you could use it to summarize a public article.
Truthfully, it's just not worth it. You either run these things so slowly that you're wasting your time or you have to buy 4- or 5-figures of hardware that's going to sit, mostly unused.
This guy ran a 4-bit quantized version with 768GB RAM: https://news.ycombinator.com/item?id=42897205
There's a couple of guides for setting it up "manually" on ec2 instances so you're not paying the Bedrock per-token-prices, here's [1] that states four g6e.48xlarge instances (192 vCPUs, 1536GB RAM, 8x L40S Tensor Core GPUs that come with 48 GB of memory per GPU)
Quick google tells me that g6e.48xlarge is something like 22k USD per month?
[0] https://aws.amazon.com/bedrock/deepseek/
[1] https://community.aws/content/2w2T9a1HOICvNCVKVRyVXUxuKff/de...
Software: client of choice to https://openrouter.ai/deepseek/deepseek-r1-0528
Sorry I'm being cheeky here, but realistically unless you want to shell out 10k for the equivalent of a Mac Studio with 512GB of RAM, you are best using other services or a small distilled model based on this one.
If speed is truly not an issue, you can run Deepseek on pretty much any PC with a large enough swap file, at a speed of about one token every 10 minutes assuming a plain old HDD.
Something more reasonable would be a used server CPU with as many memory channels as possible and DDR4 ram for less than $2000.
But before spending big, it might be a good idea to rent a server to get a feel for it.
With an average of 3.6 tokens/sec, answers usually take 150-200 seconds.