Complete hardware and software setup for running Deepseek-R1 locally
twitter.com
twitter.com
And since X sucks today just about as much as it sucked yesterday: https://nitter.poast.org/carrigmat/status/188424436990727810...
We need Nitter and BitTorrent to have a baby, p2p so we can share the load.
This would involve eliminating parts of the page that are unlikely to be stable (i.e. ads, pagination based on screen size, session specific details like usernames), and use that for the CID. That way the only time that users end up storing separate copies of the page is when it has actually been changed substantially. You'd probably also need some code to fluff the normalized content back up into something that makes sense for the device you're reading it on.
It's probably for the same reason I can't view many other sites like Reddit, Imgur etc, because our VPN exit IP is from a cloud hosting provider.
https://www.dailydot.com/debug/poast-hack-leaked-emails-dms/
I just finished a test of running this Deepseek-R1 768 GB model locally on a cluster of computers with 800 GB/s memory bandwidth (faster than the machine in the twitter post) and I can now extrapolate to a cluster with 6000 GB/s aggregate memory bandwidth and I'm sure we can reach higher speeds than Groq and Cerebras [1] on these large models.
We might even be cheap enough in OPEX to retrain these models.
Would anyone with cofounder or commercial skills be willing to set up this hosting service with me, it will take less than $30K investment but could be profitable in weeks?
[1] https://hc2024.hotchips.org/assets/program/conference/day2/7...
As someone not familiar with investment sourcing or SME financing. Could you break down the maths/accounting? How do you go from sinking 40k in a business to losing 6.5k if you turn the lights off at the end?
I would invest more than the initial $30K on optimization after the servers have found paying customers and thus have proven commercial viability. I would invest in software development, finetuning, retraining and above all reverse engineering GPU and neural engine instruction sets and adapting these open source models to the more than 2 quadrillion operations per second that these 48 servers can do.
You would be wise to do the software development I mentioned, do more sales and support than was covered under my initial $3000 labour fee. But that you can pay for with the revenues, it would not be the initial investment to see if it is viable as a business.
Because if that exists, I want to buy them all.
Given the propensity for these big tech companies to hoover up/steal any information they can gather, running these models locally, with local fine tuning looks quite attractive.
at the end of the day you still have to sell this product to the sorts of companies that are far and away all microsoft 365/google workspace clients and we're gonna have to figure that out one day or another
However, if both training and operating costs of a DeepSeek-like model are as small as they are, the companies best able to offer this service are... Microsoft, Amazon and Google. And second best are... teams inside the would-be customer enterprises themselves. $6M to train and $6K to run is effectively free for such companies; there is no moat here. The services that enterprise customers would happily buy instead of building are... operations, and assuming legal liability if the model turns out not to be safe from copyright infringement lawsuits. But those are exactly the services those companies are already buying from Microsoft, Amazon and Google.
Maybe they could skin the robotic bureucrats in vintage scifi appearance as well to have the whole consistent experience when you go to the building permits bot, there could be small talk about the latest Beatles record etc.
(I'm reiterating my prediction wrt. AI and moats - the only mid-term moat there can be is in human labor. Hardware vendors benefit from selling better hardware to more people for less; software and research are cheap to scale, datasets eventually leak or get reproduced. Human labor is the one thing that doesn't scale, and except for an economic crisis, only ever gets more expensive with time. Whatever edge one can get by applying human labor that cannot be substituted by AI - like RLHF and its evolutions - is the one that will last all the way to AGI; past that, moats won't matter anymore.)
One of the many reasons I'm firmly on the side of making the training of large neural models exempt of copyright considerations for everyone.
edit: apparently in the EU the situation is complicated by new AI specific legislation in the works: https://www.morganlewis.com/pubs/2024/02/eu-ai-act-how-far-w...
think inside the box
I myself just have proof of a single customer having their own private Cerebras rack. There are rumors about several more customers with on-prem Cerebras.
So it's true that you can't encrypt compute tasks of this type end-to-end, so you can't know if unauthorized parties mine your data. However, Microsoft is very unlikely to mine your data (for "you" being e.g. any of the many multinational corporations that already run all their office work through Azure-hosted Outlook, Office, SharePoint, etc.), or to let others mine it, because if it ever came out, your customers' lawyers would be after you, your lawyers would be after Microsoft, and the whole thing would explode into a multiple-billion-dollars shitshow and might even get a government or two involved.
That's the working assumption that makes Microsoft well-positioned to eat any fledgling self-hosted DeepSeek market in the business space. They already have things set up at a level that is trusted by governments as well as corporations in critical industrial sectors, with huge financial and legal exposure.
(Presumably Google and Amazon are in a similar position here, though I've only seen this personally with Microsoft/Azure, so that's what I can comment on.)
For contract breach civil crime like this, there is zero chance it ends with jail time.
On top of that, "everything is securities fraud" - and since that does carry potential jail time, corporations generally try to avoid pissing off parties that would be able to frame a contract breach (and its consequences) in terms of investment fraud.
EDIT:
For starters, almost all data a multinational corporation generates and processes is subject to export control regulations, which are broad, full of special cases, vary over time, space and politics, and most importantly, violations of them come with huge fines and criminal penalties[0] for both businesses and individuals involved. The only reason Microsoft can get a corporation like this to migrate to O365 and run their back-office in Azure cloud is by solid, tested contractual guarantees that the data will be processed in ways that will keep the customer compliant with applicable regulations. Now, I'm not a lawyer, but it's not particularly hard to draw a line from "Microsoft snooping on enterprise customers" to securities fraud.
I mean, even in context of hosting a DeepSeek derivative, we're talking about a cloud service offering enterprise customers secure training on company data. "Company data" may involve, e.g. detailed documentation or specs for software for designing advanced optical systems, which may sound benign until you make the connection[1]: "advanced optics" includes applications in advanced laser systems, which basically means weapons (e.g. ranging, missile targeting, anti-missile countermeasures). Obviously, regulators around the world (and the US in particular) would be very unhappy to see such information crossing through the wrong borders. For both the affected customers and the cloud service, this is high stakes game; a random startup isn't in a position to enter it.
--
[0] - E.g. in US, up to $1M per violation and up to 20 years in prison, possibly at the same time; see https://www.bis.doc.gov/index.php/enforcement/oee/penalties.
[1] - This was a real intro example used in export control training I went through some years ago.
A small random startup is unlikely to play in the securities sandbox until they have enough resources to hire enough lawyers to keep themselves out of prison and the fines "reasonable"(i.e. not enough to incentivize actually doing something about the fine being imposed other than to at least temporarily stop doing the action).
When was the last time securities fraud ended in jail time by any S&P 500 company? My quick web search returned no instances ever(but I could be wrong).
My point here is that OP's startup won't be able to compete with incumbents for enterprise money, and since the incumbents already provide this kind of service cheaply and reliably for customers of any size, all while handling applicable security concerns, OP's startup won't be able to compete with them for smaller customers either.
Given the current US government is headed by a person that just looks to take what he wants - your assurances aren't comforting.
Sure, but that's not some unexpected gotcha - it's just a plain fact of geopolitical reality, managed by international treaties and accounted for in laws and contracts around the world. A multinational enterprise isn't like a person subscribing to a free plan of a random SaaS because the "sign up" button was the right shade of green - there are armies of lawyers on both sides, tasked with navigating applicable regulations (including GDPR and export control laws) and finding out a way to make things work.
When they can't, the deal simply doesn't happen.
What actually happens is you have people seeing no evil, hearing no evil and speaking no evil - by going lalalala - hoping that because everybody else is doing it they won't get fired.
This happens because alternatives seem too hard.
There is no evidence of this happening in the last 20 years. None.
And if there was it would be the complete unravelling of the entire cloud concept.
So you're talking about solving a problem no one has.
Plenty of evidence of companies and governments using spying for commercial/national ( sometimes the same ) advantage.
So let's say you are a big company, and suddenly the US government decides you are a competitor in a nationally strategic industry - is your data safe if held by a US company?
GCP offers hundreds of models in its Vertex AI, including all "open source" (actually open weights) models, and the ability to fine tune for your specific needs. This blog post is from 2023 [1].
(disclaimer: I work at Google, but not on the Cloud team)
[1] https://cloud.google.com/blog/products/ai-machine-learning/s...
The reality is that many datasets include stuff that has been absolutely stolen or haphazardly stored online, stuff like emails, texts, conversations, our credit histories, really any information about people that exists in a large quantity has already made its way into these LLMs one way or another.
An AI trained solely/additionally on specific datasets, with the intent of predicting human behavior exactly, should yield successful and highly accurate predictions from rather general demographic data combined with an individuals location history/current their location, plus really any third substantial thing - could be browsing history, or purchases, activity on social media, that should be all that's necessary to rather accurately put everyone into a box.
The most frustrating thing is that there is a box for the people that refuse to be put into a box.
In my house I currently have almost 900 GB/S memory bandwidth in aggregate but only 132 GB total DRAM.
Especially because these models <think> for a long bit before actually answering, so they generate for a longer period. (to be clear - I find the <think> section useful, but it also means waiting for more tokens)
Personally - I end up moving down to lower quality models (quants/less params) until I hit about 15 tokens/second. For chat, that seems to be the magical spot where I stop caring and it's "fast enough" to keep me engaged.
For inline code helpers (ex - copilot) you really need to be up near 30 tokens/second to make it feel fast enough to be helpful.
For DeekSeek 1 token ~= 3 English characters.
See: https://api-docs.deepseek.com/quick_start/token_usage/
It comes out to around 1-3 words/second. This is not so slow that it's maddening (ex 2 token/second is frustratingly slow, like walk away and make coffee while it's answering slow), but it's still slow enough to make it hard to functionally use and not break flow state. You get bored and distracted reading at that pace.
I'll ask it to do a task, then it will do it in a few steps and notify me when it's done. I use that time to take care of something else.
[0] direct link with login https://x.com/carrigmat/status/1884244369907278106
[1] alt link without login but with ads https://threadreaderapp.com/thread/1884244369907278106.html
Edit: someone posted xcancel link above - no ads https://xcancel.com/carrigmat/status/1884244369907278106
It's going to take some time, but the farce is gone. We'll have parity to Chat GPT on consumer hardware soon enough. 6k is still too much. I suspect the community will be able to get this down to 2K.
I'm tempted to cancel my Chat GPT subscription!
Not every single person needs to have it. But if someone in your circle has the needed hardware...
https://www.newegg.com/p/N82E16819113866
https://www.newegg.com/supermicro-h13ssl-nt-amd-epyc-9004-se...
As for memory, these two kits should work (both are needed for the full 12 DIMMs):
https://www.newegg.com/owc-256gb/p/1X5-005D-001G0
https://www.newegg.com/owc-512gb/p/1X5-005D-001G4
Since it would be a 2DPC configuration, the memory would be limited to 4400MT/sec unless you overclock it. That would give 422.4GB/sec, which should be enough to run the full model at 11 tokens per second according to a simple napkin math calculation. In practice, it might not run that fast. If the memory is overclocked, getting to 16 tokens per second might be possible (according to napkin math).
The subtotal for the linked parts alone is $5,139.98. It should stay below $6000 even after adding the other things needed, although perhaps it would be more after tax.
Note that I have not actually built this to know how it works in practice. My description here is purely hypothetical.
> Also, an important tip: Go into the BIOS and set the number of NUMA groups to 0. This will ensure that every layer of the model is interleaved across all RAM chips, doubling our throughput. Don't forget!
This does not actually make sense. It is well known that there is a penalty for accessing memory attached to a different CPU. You don’t get more bandwidth from disabling the NUMA node information and his token generation performance reflects that. If there was a doubling effect from using two CPU sockets, he should be getting twice the performance, but he is not.
Additionally, llama.cpp’s NUMA support is suboptimal, so he is likely taking a performance hit:
https://github.com/ggerganov/llama.cpp/issues/11333
When llama.cpp fixes its NUMA support, using two sockets should be no worse than using one socket, but it will not become better unless some new way of doing the calculations is devised that benefits from NUMA. This might be possible (particularly if you can get GEMV to run faster using NUMA), but it is not how things are implemented right now.
Also how much would stuffing a GPU or 3 (3090/4090) improve speeds, even with heavy CPU layer offloading, or would the penalty be too big? I know in some cases you're swapping data into the GPU, but in others you're just doing parts on the CPU. I'm curious what the comparison for speed would be.
Chips and Cheese suggests things are even worse than this as the per CCD bandwidth is limited to around 120GB/sec, which probably ruins the idea of using the 9015, as that only has 2 CCDs:
https://old.chipsandcheese.com/2024/10/11/amds-turin-5th-gen...
https://www.techpowerup.com/cpu-specs/epyc-9015.c3903
Anyway, leveraging both sockets’ memory bandwidth would require splitting the layers into partitions for each NUMA node and doing that partition’s part of each GEMV calculation on the local CPU cores. PBLAS might be useful in implementing something like that.
As for a speed up from using 3090/4090 cards, that is a bit involved to estimate. The model has 61 layers. The way llama.cpp works is that it will offload layers and the computation will move from device to device depending on where the layers are in memory. You would need to calculate roughly how long it takes for each device to do a layer. Then multiple by the number of layers processed by that device and sum across the devices. Finally, normalize to get the number of tokens per second and you will have your answer. DeepSeek R1 has 61 layers (although I think llama.cpp will say 62 due to the embedding layer if it counts for DeepSeek like it does for llama 3). It has 37GB of activated weights, so you can do 37GB / 61 / memory bandwidth to get the time per layer. You probably want to multiply by 1.25 as a fudge factor to account for the fact that these things never run at the full speed that these calculations predict. Then you can plug in these numbers into the earlier calculation I described to get your answer.
Since tensor cores are so fast you still come out ahead when you send the weights over PCIE to the GPU and return the completed products back to main memory.
Complete hardware + software setup for running Deepseek-R1 locally. The actual model, no distillations, and Q8 quantization for full quality. Total cost, $6,000. All download and part links below:
Motherboard: Gigabyte MZ73-LM0 or MZ73-LM1. We want 2 EPYC sockets to get a massive 24 channels of DDR5 RAM to max out that memory size and bandwidth. https://t.co/GCYsoYaKvZ
CPU: 2x any AMD EPYC 9004 or 9005 CPU. LLM generation is bottlenecked by memory bandwidth, so you don't need a top-end one. Get the 9115 or even the 9015 if you really want to cut costs https://t.co/TkbfSFBioq
RAM: This is the big one. We are going to need 768GB (to fit the model) across 24 RAM channels (to get the bandwidth to run it fast enough). That means 24 x 32GB DDR5-RDIMM modules. Example kits: https://t.co/pJDnjxnfjg https://t.co/ULXQen6TEc
Case: You can fit this in a standard tower case, but make sure it has screw mounts for a full server motherboard, which most consumer cases won't. The Enthoo Pro 2 Server will take this motherboard: https://t.co/m1KoTor49h
PSU: The power use of this system is surprisingly low! (<400W) However, you will need lots of CPU power cables for 2 EPYC CPUs. The Corsair HX1000i has enough, but you might be able to find a cheaper option: https://t.co/y6ug3LKd2k
Heatsink: This is a tricky bit. AMD EPYC is socket SP5, and most heatsinks for SP5 assume you have a 2U/4U server blade, which we don't for this build. You probably have to go to Ebay/Aliexpress for this. I can vouch for this one: https://t.co/51cUykOuWG
And if you find the fans that come with that heatsink noisy, replacing with 1 or 2 of these per heatsink instead will be efficient and whisper-quiet: https://t.co/CaEwtoxRZj
And finally, the SSD: Any 1TB or larger SSD that can fit R1 is fine. I recommend NVMe, just because you'll have to copy 700GB into RAM when you start the model, lol. No link here, if you got this far I assume you can find one yourself!
And that's your system! Put it all together and throw Linux on it. Also, an important tip: Go into the BIOS and set the number of NUMA groups to 0. This will ensure that every layer of the model is interleaved across all RAM chips, doubling our throughput. Don't forget!
Now, software. Follow the instructions here to install llama.cpp https://t.co/jIkQksXZzu
Next, the model. Time to download 700 gigabytes of weights from @huggingface! Grab every file in the Q8_0 folder here: https://t.co/9ni1Miw73O
Believe it or not, you're almost done. There are more elegant ways to set it up, but for a quick demo, just do this. llama-cli -m ./DeepSeek-R1.Q8_0-00001-of-00015.gguf --temp 0.6 -no-cnv -c 16384 -p "<|User|>How many Rs are there in strawberry?<|Assistant|>"
If all goes well, you should witness a short load period followed by the stream of consciousness as a state-of-the-art local LLM begins to ponder your question:
And once it passes that test, just use llama-server to host the model and pass requests in from your other software. You now have frontier-level intelligence hosted entirely on your local machine, all open-source and free to use!
And if you got this far: Yes, there's no GPU in this build! If you want to host on GPU for faster generation speed, you can! You'll just lose a lot of quality from quantization, or if you want Q8 you'll need >700GB of GPU memory, which will probably cost $100k+
The memory bandwidth might be an issue, and it would be a pretty small percentage of the model, but I'd guess the speedup would be apparent.
Maybe not worth the few thousand for the card + more power/cooling/space, of course.
Because in 2x CPU system, the model may have to be passed via NUMA, which has 10% - 30% of memory bandwidth bandwidth
I'm not so convinced that the Nvidia panic is justified.
I do agree that it's very funny though.
You need to think really hard to get to an answer, because that’s more fine grained than the way you usually think about words and letters.
I think this is a deeper question about the bounds of human desire... which seem virtually limitless. We seem to have an unlimited appetite for answering questions on the complexity of existence. Pair that with arms race issues, and you have an obvious need for massive compute regardless of how efficient the algorithms are.
How about the amount of compute to take a buggy hextree implementation I've been poking at for a few years and completely rewrite it into a fully functioning implementation (with test cases) without me having to write a single line of code? Well...other than me having to break out the printf debugger to help with tracking down some deep bugs that is.
I've been trying to get the original implementation to work correctly for so long I don't even remember what I wanted to use it for in the first place, I just mess with it a bit here and there when I have nothing better to do.
And that's just me, a half-assed self-taught junior woodchuck coder, playing around to see what all the hype is about.
I can only imagine that I'm giving them valuable training data as some of the bugs were very deep and took a whole lot of 'thinking' to track down the root cause. It does take a bit of prodding to get it to look in the right place but so far it has found and fixed them all. The last bug was an overflow in the tree iterator's stack it uses to track state across iterations that I was concerned would time out as it was thinking for a long, long time.
I'm not really one to defend the robots but this one is actually uesful.
--edit--
Oh... I guess that's a meme now. Nothing to see here, move along...
And M4 Ultra Mac Pros are probably only weeks away too.
That said, a similar upgrade has been done on the raspberry pi 4, so it is theoretically possible:
https://hackaday.com/2023/03/05/upgrade-ram-on-your-pi-4-the...
(You don’t have to take it from me: if CPU were good enough, AMD’s valuation would be 100x its current value.)
https://github.com/ggerganov/llama.cpp/issues/11333
The TLDR is that llama.cpp’s NUMA support is suboptimal, which is hurting performance versus what it should be on this machine. A single socket version likely would perform better until it is fixed. After it is fixed, a dual socket machine would likely run at the same speed as a single socket machine.
If someone implemented a GEMV that scales with NUMA nodes (i.e. PBLAS, but for the data types used in inference), it might be possible to get higher performance from a dual socket machine than we get from a single socket machine.
v3 yes w/ 37B activated params, yes, but terrible on 405B as it's a dense model.
These guys say they got really good results with 2.51-bit quantization of the original R1. The original has 671B params weighing in at 720GB - that's what they're running on this $6000 setup. According to these guys[0] they get really good results at 2.51bit quantization which would be this model[1] which is still 671B params, but weighs in at 212GB.
[0] https://unsloth.ai/blog/deepseekr1-dynamic [1] https://huggingface.co/unsloth/DeepSeek-R1-GGUF/tree/main/De...
Not completely clear what changing numa nodes per socket from 1 to 0 does, possibly gives linux less information about when to migrate threads across x64 cores? (didn't upset llvm compile time so I'll leave it on nsp0)
If I wanted to run a production-grade service using the full Deepseek model, with good tokens/sec and the ability to serve concurrent requests, what sort of hardware are we looking at?
This is going to drop by half when Nvidia starts shipping DIGITS. I think we’re all going to want one. It’ll probably have a much bigger impact than Apple VisionPro, that costs the same.
I can already think of using it as a much more intelligent local Siri/Alexa to control devices. It’s something that can actually keep the kids engaged with useful trivia/knowlege (better than watching mindless trash on YT) or it can just humour me whenever I want - all without needing to worry about privacy.
Apple is going in the direction of total integration. Louis Rossmann and iFixit will hem and haw, but soon there will be a MacBook whose motherboard, besides cooling, PSU, and ports, consists simply of a single component that houses CPU, GPU, RAM, storage, I/O port controllers, radios (wifi, Bluetooth, etc.), firmware for all of the above, and a security module plus keys, all directly on the CPU bus, and it will be glorious. It will absolutely lap any PC laptop, and in single-core performance will smoke even high-end AMD Epyc beastbox builds because of the aggressive elimination of inter-component latency.
But there won't be a need to fix it. If it breaks you just recycle it and buy new, being sure to sync your data back from Apple Cloud -- but it probably won't break. Kinda like how unibody cars are both safer and more reliable than the much more fixable cars of the 60s, even if they crumple like tinfoil and must be totaled upon experiencing any sort of impact.
Source: first figure from https://www.anandtech.com/show/17024/apple-m1-max-performanc...
Hopefully this year's M4 Ultra systems will at least allow a much higher top spec.
Now this is basically the moment of Apache/ngnix being free and open source.
You then have share hosting phenomenon out of it.
But that's fixable.
Since it's memory bound, it might be possible to reach 15 tok/sec with this build.
> ...if you want Q8 you'll need >700GB of GPU memory, which will probably cost $100k+
Now it makes sense.
Still undecided how I feel about having the ability to use all that quality in the full size model if one could only retrieve it at 6-8 tokens per second.
Screenshots or mirrors (without the login requirement) are okay.
(I use this firefox extension to automatically rewrite twitter links into it: https://addons.mozilla.org/en-US/firefox/addon/privacy-redir... )
Nitter Chrome redirection extension: https://chromewebstore.google.com/detail/nitter-redirect/moh...
Once the brain dead greedy MBAs get involved, is just how much you can steal. It should all be sold short, as we watch the world burn.
$6k is much less than the millions that would be required to run anything by OpenAI. And it's a first pass. It could get much lower by the end of the year