How to Run DeepSeek R1 671B Locally on a $2000 EPYC Server
digitalspaceport.com
digitalspaceport.com
This [1] X thread runs the 671B model in the original Q8 at 6-8 TPS for $6K using a dual socket Epyc server motherboard using 768GB of RAM. I think this could be made cheaper by getting slower RAM but since this is RAM bandwidth limited that would likely reduce TPS. I’d be curious if this would just be a linear slowdown proportional to the RAM MHz or whether CAS latency plays into it as well.
[1] https://x.com/carrigmat/status/1884244369907278106?s=46&t=5D...
Given the model has only so few active parameters per token (~40B), it is likely that just being able to hold it in memory absolve the largest bottleneck. I guess with a single consumer PCIe4.0x16 graphics card you could get at most 1tps just because of the PCIe transfer speed? Maybe CPU processing can be faster simply because DDR transfer is faster than transfer to the graphics card.
I've had to disable the overload safeties in LM Studio and tweak with some loader parameters to get the model to run mostly from disk (NVMe SSD), but once it did, it also used very little CPU!
I tried offloading to GPU, but my RTX 4070 Ti (12GB VRAM) can take at most 4 layers, and it turned out to make no difference in tps.
My RAM is DDR4, maybe switching to DDR5 would improve things? Testing that would require replacing everything but the GPU, though, as my motherboard is too old :/.
Some math:
DDR5 6000 is 3000mhz x 2 (double data rate) x 64 bits / 8 for bytes = 48000 /1000 = 48GB/s
DDR3 1866 is 933mhz x 2 x 64 / 8 / 1000 = 14.93GB/s. If you have 4 channels that is 4 x 14.93 = 59.72GB/s
Which suggests it wouldn't be quite the right fit here -- the precomputed constants in the model aren't changing, nor do they need to persist.
Still, interesting question, and I wonder if there's some other existing bit of tech that can be repurposed for this.
I wonder if/when this application (LLMs in general) will slow down and stabilize long enough for anything but general purpose components to make sense. Like, we could totally shove model parameters in some sort of ROM and have hardware offload for a transformer, IF it wasn't the case that 10 years from now we might be on to some other paradigm.
>
> TFA says it can bump the spec to 768 GB but that it's then more like > $2500 than $2000. At 768 GB that'd be the full, 8 bit, model.
> Seems indeed like a good price compared to $6000 for someone who wants to hack a build.
> I mean: $6 K is doable but I take it take many who'd want to build such a machine for fun would prefer to only fork $2.5K.
.
I am not sure why TacticalCoder's comment was downvoted to oblivion. I would have upvoted if the comment wasn't already dead.
I've also vouched as it doesn't seem like a comment deserving to be dead at all. For at least this instant it looks like that was enough vouches to restore the comment.
Btw, I agree that that was a good comment that deserved vouching! But of course we have to ban accounts because of the worst things they post, not the best.
Seems indeed like a good price compared to $6000 for someone who wants to hack a build.
I mean: $6 K is doable but I take it take many who'd want to build such a machine for fun would prefer to only fork $2.5K.
Per o3-mini, the blocked gemm (matrix multiply) operations have very good locality and therefore MT/s should matter much more than CAS latency.
I built the machine for $5500 four years ago and it certainly has not paid for itself, but it still has tons of utility and will probably last another four years bringing my monthly cost to ~$50/mo which is way lower than what a cloud provider would charge, especially considering egress network traffic. Instead of paying Discord, Twitter, Netflix/Hulu/Amazon/etc, paid game hosting, and ChatGPT, I can self host Jitsi/Matrix, Bluesky, Plex, SteamCMD, and ollama. In total I end up spending about the same, but I have way more control, better access to content, and can do more when offline for internet outages.
Thanks to CloudFlare Tunnel, I dont have to pay a cloud vendor, cdn or vpn for good routes to my web resources or opt into paid DDoS protection services. It's fantastic.
I have to run the, uh, "crispier" very compressed version because otherwise it'll spill into swap. I use the 212GB .gguf one from Unsloth's page, with a name that I can't remember on top of my head but I think it was the largest they made using their specialized quantization for llama.cpp. Jpeggified weights. Actually I guess llama.cpp quantization is a bit closer analogy to the reducing number of colors rather than jpeg-style compression crispiness? Gif had reduced colors (256) IIRC. Heavily gif-like compressed artificial brains. Gifbrained model.
Just like you, I use it for tons of other things that have nothing to do with AI, it just happened to be convenient that Deepseek-R1 came out and just about barely is able to run it on this thing, with enough quality to be coherent. My use otherwise is mostly hosting game servers for my friend groups or other random CPU-heavy projects.
I haven't investigated myself but I've noticed in passing: There is a person on llama.cpp and in /r/localllama who is working on specialized CPU-optimized Deepseek-R1 code, and saw them asking for an EPYC machine for testing, with specific request for a certain configuration. IIRC also said that the optimized version needs new quants to get the speeeds. So maybe this particular model will get some speedup if that effort succeeds.
This rig does >4 tok/s, which is ~15-20 ktok/hr, or $0.04/hr when purchased through a provider.
You're probably spending $0.20/hr on power (1 kW) alone.
Cool achievement, but to me it doesn't make a lot of sense (besides privacy...)
I would argue that is enough and that this is awesome. It was a long time ago I wanted to do a tech hack like this much.
A) somehow continuously interact with the running model, ambient-computing style. Say have the thing observe you as you work, letting it store memories.
B) allowing it to process those memories when it chooses to/whenever it's not getting any external input/when it is "sleeping" and
C) (this is probably very difficult) have it change it's own weights somehow due to whatever it does in A+B.
THAT, in a privacy friendly self-hosted package, i'd pay serious money for
Quite scary. As the meme has it, it seems that we're getting ready to create the Torment Nexus from classic sci-fi novel Don't Create The Torment Nexus.
Privacy is worth very much though.
It’s cool to run things locally and it will get better as time goes on but for most use cases I don’t find it worth it. Everyone is different and folks that enjoy the idea of local network secure can run it locally.
Wouldn't that be much more cost-effective?
Especially when you inevitably want to run a better / different model in the near future that would benefit from different hardware?
You can get similar Tok/sec on a single RTX 4090 - which you can rent for <$1/hr.
Is it possible that this is an AI bubble subsidy where we are actually getting it below cost?
Of course for conventional compute cloud markup is ludicrous, so maybe this is just cloud economy of scale with a much smaller markup.
Of course that doesn't map well to an individual chatting with a chat bot. It does map well to something like "hey, laptop, summarize these 10,000 documents."
1. Economies of scale. Cloud providers are using clusters in the tens of thousands of GPUs. I think they are able to run inference much more efficiently than you would be able to in a single cluster just built for your needs.
2. As you mentioned, they are selling at a loss. OpenAI is hugely unprofitable, and they reportedly lose money on every query.
He uses old, much less efficient GPUs.
He also did not select his living location based on the electricity prices, unlikely the cloud providers.
that's the whole point of local models
lol.
Yeah, just besides that one little thing. We really are a beaten down society aren't we.
I am as pro-privacy as they come, but let’s not pretend that government and corporate surveillance is some wild new thing that just appeared. Read Horace’s Satires for insight into how non-private private correspondence often was in Ancient Rome.
Most of us have more privacy than 200 years ago in some ways, and much less privacy in other ways.
The odds of a cloud server leaking my information is non-zero, but it’s very small. A government entity could theoretically get to it, but they would be bored to tears because I have nothing of interest to them. So practically speaking, the threat surface of cloud hosting is an acceptable tradeoff for the speed and ease of use.
Running things at home is fun, but the hosted solutions are so much faster when you actually want to get work done. If you’re doing some secret sensitive work or have contract obligations then I could understand running it locally. For most people, trying to secure your LLM interactions from the government isn’t a priority because the government isn’t even interested.
Legally, the government could come and take your home server too. People like to have fantasies about destroying the server during a raid or encrypting things, but practically speaking they’ll get to it or lock you up if they want it.
I believe it was an 850w PSU on the spec sheet?
Not sure about where you are but where I am a 2kW plus li-ion batteries is about 2months of the average salary here, not for tech, average salary, to put it into perspective converted to USD that is 1550 usd. Panels is maybe 20% of that cost, you can add 4kW of panels for 450 USD where I am.
So for less than the price of that PC I would be able to do 2kW of solar with li-ion batteries and overspecing panels by double. None of that cheaping out on components, can absolutely get lower than that if cheaping out. Installation will be maybe another 500-600 USD here, likely to be much higher depending on region. Also to put it into perspective we pay about 0.3 USD cents per kWh for electricity and this would pay for itself in between a year and two in savings.
By the time it needs to be replaced which is from 5-7 years on the stuff I just got pricing on it would have 100% offset the cost of running.
Again I am lucky and we effectively get 80-100% output year round even with cloud cover, you might be pretty far north and that doesn't apply.
TLDR: it depends but if you are in the right region and this setup generates even some income for you the cost to go solar is negative, it would actually not make financial sense to not do it, concidering a 2K USD box was in your budget.
And I think your math is off, $0.20 per kWh at 1 kW is is $145 a month. I pay $0.06 per kWh. I've got what, 7 or 8 computers running right now and my electric bill for that and everything else is around $100 a month, at least until I start using AC. I don't think the power usage of something like this would be significant enough for me to even shut it off when I wasn't using it.
Anyway, we'll find out, just ordered the motherboard.
That is like, insanely cheap. In Europe I'd expect prices between $0.15 - 0.25 per kWh. $0.06 sounds like you live next to some solar farm or large hydro installation? Is that a total price, with transfer?
For those that aren't following - means you're spending ~$10/MTok on power alone (compared to $2/MTok hosted).
But the newest AI models require an order of magnitude more RAM than my system or the systems I typically rent have.
So I’m curious to people here, has this in the history of software happened before? Maybe computer games are a good example. There people would also have to upgrade their system to run the latest games.
If anything this isn’t so bad: $4K in 2025 dollars is an affordable desktop computer from the 90s.
But it's hard to tell because most of the stuff posted is people trying to do duct tape and bailing wire solutions.
I think the problem is thinking that you always need to use the best LLM. Consider this:
- When you don't need correct output (such as when writing a blog post, there's no right/wrong answer), "best" can be subjective.
- When you need correct output (such as when coding), you always need to review the result, no matter how good the model is.
IMO you can get 70% of the value of high end proprietary models by just using something like Llama 8b, which is runnable on most commodity hardware. That should increase to something like 80% - 90% when using bigger open models such as the newly released "mistral small 3"
> When you don't need correct output (such as when writing a blog post, there's no right/wrong answer), "best" can be subjective.
This is like, everything that is wrong with the Internet in a single sentence. If you are writing a blog post, please write the best blog post you can, if you don't have a strong opinion on "best," don't write.
for rapidly developing prototypes or working on side projects, i find llama 8b useless. it might take 5-6 iterations to generate something truly useful. compared to say 1-shot with claude sonnet 3.5 or open ai gpt-4o. that’s a lot less typing and time wasted.
Maybe a better comparison would be weather simulations in the 90s? We had access to their outputs in the 90s but running the comparable calculations as a regular Joe might've actually been impossible without a huge bankroll.
https://www.logicalincrements.com/
Still accessible but only for dedicated hobbyists with deeper pockets.
= "When you needed to run common «building blocks» (such as, in other times, «Python, Docker, or C++» - normal fundamental software you may have needed), even scrappy hardware would suffice in the '90s"
As a matter of facts, people would upgrade foremostly for performance.
> OpenAI's nightmare: DeepSeek R1 on a Raspberry Pi
https://x.com/geerlingguy/status/1884994878477623485
I haven't tried it myself or haven't verified the creds, but seems exciting at least
or well, I guess you save a bit on power usage.
Not very cheap though! But you get a quite usable personal computer with it...
Are there any other laptops around other than the larger M series Macs that can run 30-70B LLMs at usable speeds that also have useful battery life and don’t sound like a jet taxiing to the runway?
For non-portables I bet a huge desktop or server CPU with fast RAM beats the Mac Mini and Studio for price performance, but I’d be curious to see benchmarks comparing fast many core CPU performance to a large M series GPU with unified RAM.
There is only one model of Deepseek (671b), all others are fine-tunes of other models
If you're paying that much you're being ripped off. They're $800-900 on eBay and IMO are still overpriced.
Idle wattage: 60w (well below what I expected, this is w/o GPUs plugged in)
Loaded wattage: 260w
RAM Speed I am running currently: 2400 (V likely 3200 has a decent perf impact)
I was an AI sceptic until 6 months ago, but that’s probably going to be my dev setup from spring onwards - running DeepSeek on it locally, with a nice RAG to pull in local documentation and datasheets, plus a curl plugin.
It's just vaporware until then.
It’s also a more general comment around „AI desktop appliance“ vs homebuilts. I’d rather give NVIDIA/AMD $3k for a well adjusted local box than tinkering too much or feeding the next tech moloch, and have a hunch I’m not the only one feeling that way. Once it’s possible of course.
Just because Jensen calls it a super computer and gives it a DGX-1 design, doesn't make it one.
In the Cleo Abram interview [1], Jensen said that DIGITS is 6 times more powerful than the first DGX-1.
According to this PDF [2], DGX-1 had 170 TFLOPS of FP16 (half precision). 170x6=1020 TFLOP (~1 PFLOP). Yes DIGITS is suppose to have 1 PFLOP, but according to the presentation, it should be in FP4...
He also said that it will draw 10k times less power. But DGX-1 had a TDP of 3.5kW [3] and I highly doubt DIGITS will draw 3500/10000=0.35W... the GPU alone will have a peak TDP that is more like 200 times higher than that.
I mean, we all know that NVIDIA does fudge the numbers in charts. Like comparing FP8 from last generation, to FP4 on this. But this is extreme.
Having said that. Do I believe that they can deliver a laptop (in another form factor) and it will perform 1 PFLOP of FP4. Of course! Like I said, it is nothing special. Both Apple and AMD have unified memory in relatively cheap systems.
1. https://youtu.be/7ARBJQn6QkM
2. http://images.nvidia.com/content/technologies/deep-learning/...
3. https://images.nvidia.com/content/pdf/dgx1-v100-system-archi...
My guess is that they will use the RTX 5070 Ti laptop version (992 TFLOPS, slightly higher clocked to reach 1000 TFLOPS/ 1 PFLOP).
Their big GB200 chips have 546 GB/s to their LPDDR memory, they could use the same memory controler on the GB10. They don't need to design a new one. It would still be slower than what they are currently using on the RTX 5070 Ti laptop GPU, but any slower than that, and there is no chance that they could argue that it would hit anywhere near 1 PFLOP of FP4. It would only be possible in extreme edge case scenarios when all data will fit in it's 40MB L2 cache.
We kind of know what storage cost in a store and we know that Apple (Mac computers) and every phone manufacturer adds a ton of cost for a small increase. NVIDIA will probably do the same.
I have no idea what the cost for their cabling would be, but they exist in 100G, 200G, 400G and 800G speeds and you seem to need two of them.
If you are only going to use one DIGITS, and you can make do with whatever is the smallest storage option, then it is $3000. Many people might have another computer (set up FTP/SMB or similar solution), NAS or USB thumbdrive/external hardrive where they can stor extra data, and in that case you can have more storage without paying for more.
We already know that it is going to be one single CPU and GPU and fixed memory. The GPU is most likely the RTX 5070 Ti laptop model (992 TFLOPS, clocked 1% higher to get 1 PFLOP).
Any suggestions on building a low-power desktop that still yields decent performance?
You don't for now. The bottleneck is mem throughput. That's why people using CPU for LLM are running xeon-ish/epyc setups...lots of mem channels.
The APU class gear along the lines of Halo Strix is probably the path closest to lower power but it's not going to do 500gb of ram and still doesn't have enough throughput for big models
I think your best bet might be a Ryzen U-series mini PC. Or perhaps an APU barebone. The ATX platform is not ideal from a power-efficiency perspective (whether inherently or from laziness or conspiracy from mobo and PSU makers, I do not know). If you want the flexibility or scale, you pay the price of course but first make sure it's what you want. I wouldn't look at discrete graphics unless you have specific needs (really high-end gaming, workstation, LLMs, etc) - the integrated graphics of last few years can both drive your 4k monitors and play recent games at 1080p smoothly, albeit perhaps not simultaneously ;)
Lenovo Tiny mq has some really impressive flavors (ECC support at the cost of CPU vendor-lock on PRO models) and there's the whole roster of Chinese competitors and up-and-comers if you're feeling adventerous. Believe me you can still get creative if you want to scratch the builder itch - thermals is generally what keeps these systems from really roaring (:
https://www.maginative.com/article/nvidia-leverages-ai-to-as...
> NVIDIA researchers customized LLaMA by training it on 24 billion tokens derived from internal documents, code, and other textual data related to chip design. This advanced “pretraining” tuned the model to understand the nuances of hardware engineering. The team then “fine-tuned” ChipNeMo on over 1,000 real-world examples of potential assistance applications collected from NVIDIA’s designers.
2023 paper, https://research.nvidia.com/publication/2023-10_chipnemo-dom...
> Our results show that these domain adaptation techniques enable significant LLM performance improvements over general-purpose base models across the three evaluated applications, enabling up to 5x model size reduction with similar or better performance on a range of design tasks.
2024 paper, https://developer.nvidia.com/blog/streamlining-data-processi...
> Domain-adaptive pretraining (DAPT) of large language models (LLMs) is an important step towards building domain-specific models. These models demonstrate greater capabilities in domain-specific tasks compared to their off-the-shelf open or commercial counterparts.
I spent extra on the 9274F because of some published benchmarks [1] that showed that the 9274F had STREAM TRIAD results of 395 GB/s (on 460.8 GB/s of theoretical peak memory bandwidth), however sadly, my results have been nowhere near that. I did testing with LIKWID, Sysbench, and llama-bench, and even w/ an updated BIOS and NUMA tweaks, I was getting <1/2 the Fujitsu benchmark numbers:
Results for results-f31-l3-srat:
{
"likwid_copy": 172.293857421875,
"likwid_stream": 173.132177734375,
"likwid_triad": 172.4758203125,
"sysbench_memory_read_gib": 191.199125,
"llama_llama-2-7b.Q4_0": {
"tokens_per_second": 38.361456,
"model_size_gb": 3.5623703002929688,
"mbw": 136.6577115303955
}
}
For those interested in all the system details/running their own tests (also MLC and PMBW results among others): https://github.com/AUGMXNT/speed-benchmarking/tree/main/epyc...[1] https://sp.ts.fujitsu.com/dmsp/Publications/public/wp-perfor...
Since the CPU is clocked quite high, figures you should be getting are I guess around ~100ns, but probably less than that, and 40-ish GB/s of BW. If those figures do not match then it could be either a motherboard (HW) or BIOS (SW) issue or RAM stick issue.
If those figures closely match then it's not a RAM issue but a motherboard (BIOS or HW) and you could continue debugging by adding more and more cores to the experiment to understand at which point you hit the saturation point for the bandwidth. It could be a power issue with the mobo.
mlc --idle_latency
Intel(R) Memory Latency Checker - v3.11b
Command line parameters: --idle_latency
Using buffer size of 1800.000MiB
Each iteration took 424.8 base frequency clocks ( 104.9 ns)
As does the per-channel --bandwidth_matrix results: Numa node
Numa node 0 1 2 3 4 5 6 7
0 45999.8 46036.3 50490.7 50529.7 50421.0 50427.6 50433.5 52118.2
1 46099.1 46129.9 52768.3 52122.3 52086.5 52767.6 52122.6 52093.4
2 46006.3 46095.3 52117.0 52097.2 50385.2 52088.5 50396.1 52077.4
3 46092.6 46091.5 52153.6 52123.4 52140.3 52134.8 52078.8 52076.1
4 45718.9 46053.1 52087.3 52124.0 52144.8 50544.5 50492.7 52125.1
5 46093.7 46107.4 52082.0 52091.2 52147.5 52759.1 52163.7 52179.9
6 45915.9 45988.2 50412.8 50411.3 50490.8 50473.9 52136.1 52084.9
7 46134.4 46017.2 52088.9 52114.1 52125.0 52152.9 52056.6 52115.1
I've tried various NUMA configurations (from 1 domain to a per-CCD config) and it doesn't seem to make much difference.Updating from the board-delivered F14 to the latest 9004 F31 BIOS (the F33 releases bricked the board and required using a BIOS flasher for manual recover) gave marginal (5-10%) improvement, but nothing major.
While 1DPC, the memory is 2R (but still registers at 4800), training on every boot. The PMBW graph is probably the most useful behavior chart: https://github.com/AUGMXNT/speed-benchmarking/blob/main/epyc...
Since I'm not so concerned with CPU inference, I feel like the debugging/testing I've done is... the amount I'm going to do, which is enough to at least characterize, if not fix the performance.
I might write up a more step-by-step guide at some point to help others but for now the testing scripts are there - I think most people who are looking at theoretical MBW should probably do their own real-world testing as it seems to vary a lot more than GPU bandwidth.
lkwid -t load -i 100 -w S0:5GB:8:1:2
and see what you get. I think you should be able to get somewhere around ~200 GB/s.Can you repeat the same lkwid experiment but with 1, 2 and 4 threads? I'm wondering when is it that it begins to detoriate quickly.
Maybe also worth doing is repeating the 8 threads but forcing lkwid to pick every third physical core so that you get 1 thread per CCD experiment setting.
With `likwid-bench -i 100 -t load -w M0:5GB:1 -w M1:5GB:1 -w M2:5GB:1 -w M3:5GB:1 -w M4:5GB:1 -w M5:5GB:1 -w M6:5GB:1 -w M7:5GB:1` we get 187976.60
Obvious there's a bottleneck either going on somewhere - at 33.5GB/s per channel, that would get close to 400GB/s, what you'd expect, but the reality is that it doesn't get to half of that. Bad MC? Bottleneck w/ the MB? Hard to tell, not sure that without swapping hardware there's much more that can be done to diagnose things.
At a quick glance, some of them look interesting such as "Workload tuning" where you can pick different profiles. There is "memory throughput intensive" profile. You can also try to explicitly disable DIMMs that are not in use given you use only half of them. I wouldn't hold my breath that any of these will make a big difference but you can give it try.
The bug report used AMDuProf to confirm that the bandwidth is actually ~2x than what likwid reported. You could try the same.
Still, 8x3090 gives you ~2.25 bits per weight, which is not a healthy quantization. Doing bifurcation to get up to 16x3090 would be necessary for lightning fast inference with 4bit quants.
At that point though it becomes very hard to build a system due to PCIE lanes, signal integrity, the volume of space you require, the heat generated, and the power requirements.
This is the advantage of moving up to Quadro cards, half the power for 2-4x the VRAM (top end Blackwell Quadro expected to be 96GB).
Benchmarks: https://github.com/ggerganov/llama.cpp/issues/11474#issuecom...
You can rent a single H200 for 3$/hour.
Yesterday I was comparing DeepSeek-R1 (NVidia hosted version) with both Sonnet 3.5 (regarded by most as most capable coder) and the new Gemini 2.0 flash, and the wait was worth it. I was trying to get all three to create a web page with a horizontally scrolling timeline with associated clickable photos...
Gemini got to about 90% success after half a dozen prompts, after which it became a frustrating game of whack-a-mole trying to get it to fix the remaining 10% without introducing new bugs - I gave up after ~30min. Sonnet 3.5 looked promising at first, generating based on a sketch I gave it, but also only got to 90%, then hit daily usage limit after a few attempts to complete the task.
DeepSeek-R took a while to generate it, but nailed it on first attempt.
Lets say I ask for some function that calculates some matrix math in python. It will spit out something but I dont like what it did. So I will say, now dont us any calls to that library you pulled in, and also allow for these types of inputs. Add exception handling...
So response time is important since its a conversation, no matter how correct the response is.
When you say deep seek "nailed it on the first attempt" do you mean it was without bugs? Or do you mean it worked how you imagined? Or what exactly?
With Sonnet 3.5, given the same brief prompt I gave DeepSeek-R, it took a half dozen feedback steps to get to 90%. Trying a hand drawn sketch input to Sonnet instead was quicker - impressive first attempt, but iterative attempts to fix it failed before I hit the usage limit. Gemini was the slowest to work with, and took a lot of feedback to get to the "almost there" stage, after which it floundered.
The AI companies seem to want to move in the direction of autonomous agents (with reasoning) that you hand a task off to that they'll work on while you do something else. I guess that'd be useful if they are close to human level and can make meaningful progress without feedback, and I suppose today's slow-responding reasoning models can be seen as a step in that direction.
I think I'd personally prefer something fast enough responding to use as a capable "pair programmer", rather than an autonomous agent trying to be an independent team member (at least until the AI gets MUCH better), but in either case being able to do what's being asked is what matters. If the fast/interactive AI only gets me 90% complete (then wastes my time floundering until I figure out it's just not capable of the task), then the slower but more capable model seems preferable as long as it's significantly better.
The AI companies seem to be pushing AI-assisted software development as an early use case, but I've always thought this is one of the more difficult things for them to become good at, since many/most development tasks require both advanced reasoning (which they are weak at) and ability to learn from experience (which they just can't do). The everyday, non-development tasks, like "take this photo of my credit card bill and give me category subtotals" are where the models are now actually useful, but software development still seems to be an area where they are highly impressive but ultimately not capable enough to be useful outside of certain narrow use cases. That said, it'll be interesting to see how good these reasoning models can get, but I think that things like inability to learn (other than in-context) put a hard limit on what this type of pre-trained LLM tech will be useful for.
IO takes CPU cycles but I’ve not seen evidence that striping impacts that. Memory overhead is minimal, as the stripe to read from is done via simple math from a tiny data structure.
I've written this more that it works, but its not a put another drive in and done solution. but if you just want to dumb a second drive into it and use mdraid/zfs you will have an overhead. of course if somebody tunes it and builds the application around it you can trim down the overhead significantly.
I'd love so much if quad-channel Strix Halo could get up to 256gb of memory. 192GB (4x48) won't be too bad, and at 8533MT/s, should provide competitive-ish throughput to these massive Epyc systems. Of course the $6k 24 channel 3200MHz is 4.4x more throughput but it does have a field of dimms to get there, and high power consumption.
Is this really that cheap? Looking at several local (CZ) eshops, i cannot find 32 GB DDR4 ECC RDIMM cheaper than $75, which will be $1200 for 512 GB.
So, when I use a "full" AI like chatGPT4o, I ask it questions and it has a firm grip on a vast amount of knowledge, like, whole-internet/search-engine scope knowledge.
If I run an AI "locally", on even a muscular server, it obviously does NOT have vast amounts of stored information about everything. So what use is it to run locally? Can I just talk to it as though it were a very smart person who, tragically, knows nothing?
I mean, I suppose I could point it to a NAS box full of pdf's and ask questions about that narrow range of knowledge, or maybe get one of those downloaded wikipedia stores. Is that what folks are doing? It seems like you would really need a lot content for the AI to even be remotely useable like the online versions.
Some AI services allow the use of 'tools', and some of those tools can search the web, calculate numbers, reserve restaurants, etc. However, you'll typically see it doing that in the UI.
Local models can do that too, but it's typically a bit more setup.
So I would assume the locally queried output to be comparable with the output you get from an online service (they probably use slightly better models, I don't think they release their latest ones to the public).
...the full version needs over 700GB of disk space.
THAT is rather shocking. Vastly smaller than I would expect.This is probably one of the most confusing things about LLMs. They are not vast archives of information and the models do not contain petabytes of copied data.
This is also why LLMs are so often wrong. They work by association, not by recall.
See it demonstrated in a <7 minute video here: https://www.youtube.com/watch?v=d1Fnfvat6nM
The video explains that you can download the larger models on that Github page and use them with other command line parameters, and shows how you can get a Windows + nVidia setup to GPU accelerate the model (install CUDA and MSVC / VS Community edition with C++ tools, run for the first time from MSVC x64 command prompt so it can build a thing using cuBLAS, rerun normally with "-ngl 35" command line parameter to use 3.5GB of GPU memory (my card doesn't have much)).
"IMPORTANT: This video is obsolete as of December 26, 2023 GPU now works out of the box on Windows. You still need to pass the -ngl 35 flag, but you're no longer required to install CUDA/MSVC."
So that's convenient.
When you quantize, you sacrifice model performance. In addition, a lot of the models favored for local use are already very small (7b, 3b).
What OP is pointing out is that you can actually run the full deepseek r1 model, along with all of the ‘knowledge’ on relatively modest hardware.
Not many people want to make that tradeoff when there are cheap, performant APIs around but for a lot of people who have privacy concerns or just like to tinker, it is pretty big deal.
I am far removed from having a high performance computer (although I suppose my MacBook is nothing to sneeze at), but I remember building computers or homelabs back in the day and then being like ‘okay now what is the most stressful workload I can find?!’ — this is perfect for that.
[0] not true but let's ignore that for a second
Another thing I wonder is whether using a bunch of Geforce 4060 TI with 16GB could be useful - they cost only around 500 EUR. If VRAM is the bottleneck, perhaps a couple of these could really help with inference (unless they become GPU bound, like too slow).
hit between 4.25 to 3.5 TPS (tokens per second) on the Q4 671b full model ollama pull deepseek-r1:671b
This will pull down 400GB: https://ollama.com/library/deepseek-r1:671bBut the Huggingface repo has 163 files of ~4.3GB each, so around 700GB: https://huggingface.co/deepseek-ai/DeepSeek-R1/tree/main
Anything lower than 10 t/s is going to give lot of wait time considering the reasoning wait time
Can you describe for me what the product does?
It's not that hard to make a turnkey "just add power" appliance that does nothing but spit out tokens. Some sort of "ollama appliance", which just sits on your network and provides LLM functionality for your home lab?
But beyond that, what would your mythical dream product do?
automatic updates etc etc
the AI home appliance
great for privacy
Also barely anyone can actually run the real R1 locally.
Is the setup disingenuous to get people excited about the post or what is going on here?
ONNX version: https://huggingface.co/onnx-community/DeepSeek-R1-Distill-Qw...
Someone getting a dollar or two in return for you following an affiliate link after you read something they put real time and effort into to make it valuable info for others is not “affiliate link spam”.
Also, this whole self hosting of LLMs is a bit like cloud. Yes, you can do it, but it's a lot easier to pay for API access. And not just for small users. Personally I don't even bother self hosting transcription models which are so small that they can run on nearly any hardware.
Why not? While 3-4 tok/s is still on the lower end, it is still usable to the point that I can use it for any task that doesn't require me to get into a real-time communication with the model.
In other words, I don't mind waiting a 1-minute for good-enough response from the model for topic that would take me multiples of that to compile and research on my own. It's a clear net win.
I would love to see Deepseek running on premise with a decent TPS.