From the creator of Redis; run LLM locally with ds4
dwarfstar.sh
dwarfstar.sh
The project GitHub page is a much better introduction for the hn crowd.
For this project in particular I've known about ds4 for a while and the landing page feels like it's doing a gigantic disservice other than providing a download link. IMO, ds4 is far more interesting than this landing page would suggest!
In addition to the library bindings, we have a small library of tools (workspace for view/edit, scratchpad for persistence) and making your own is registering a Go function. And in recent weeks, I added the Vision and Qwen support, as ds4 added them.
Even if you don't use the Go library, the ds4go binary makes it really easy to download the libraries off of HuggingFace with a TUI available vie Homebrew.
Here's some TUI toy screenshots, sorry I still haven't released that code; it's of different quality than the others. [3]
EDIT: add ds4go TUI screenshot gist [4]
[1] https://github.com/NimbleMarkets/ds4/releases/tag/v0.8.20260...
[2] https://github.com/nimblemarkets/ds4go#install
[3] https://gist.github.com/neomantra/ae47422c8daf7a458212c93992...
[4] https://gist.github.com/neomantra/40180ade13df93290250ce8c6d...
I'm also looking into expanding the protocol and the engine to support various steering techniques.
Maybe Intel and AMD should help them with that.
I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone's specific needs. Do you have a specific machine/spec in mind?
I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR.
On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.
I'm pretty sure 2027 will be a very interesting year for local models and inference.
If your GPU supports XMX we could also explore using it to improve the prefill kernel, but I don't have the hardware to test it myself.
As I said, I don't have the hardware to test it myself, so let me know how it goes!
I've submitted it as PR, as well as an initial implementation for an OpenAI like API
Now Ive been running qwen 3.8 flash next for more than a week and it’s doing great, really fast and super long context windows. Sometimes the model is behaving stupidly by not remembering something I said earlier but it could be also a problem from the agentic AI harness. Im using oh my pi but Im wondering what people are using ds4 with here ?
small native inference engine optimized first for DeepSeek V4 Flash (including the experimental vision model), DeepSeek V4.1 Flash (Metal, and text inference on CUDA), and additionally GLM 5.2 and 5.3, GLM 5.3 Flash and DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA)
This is local targeting high end consumer hardware like DGX Spark or AMD Ryzen AI Halo.For our mere mortals that were kids not long ago and can't really believe we've got our hands on a x090 series targeting Qwen3.8 27b, https://github.com/noonghunna/club-3090 is the way to go.
I'm maintaining a web frontend for this, trying to at least. You can follow it here: https://github.com/gchamon/club-3090-server
Does not exist. You're thinking of Qwen3.6 35ba3b
The other inference engine are also model by model with a huge switch statement deciding which part to load for which model or are they very generic?
But what really made the difference was his understanding of hardware and systems programming in general and low level or architectural tricks to pull.
From the github repo it seems like you really don't need a big Mac with huge amounts of RAM but SSD is sufficient.
If this is anywhere near 50 TPS, that would be a game changer in the personal LLM space!
https://gist.github.com/neomantra/d49df05d6b137b9e6844186499...
I started playing with local LLM+MCP in April 2025... I had to beg Qwen to look at the tool list and try anything.
These ds4 models, will happily call tools and all those harnesses I've made are composed of custom tools.
Once the HuggingFace+OpenAI showed how powerful notes are, I added a scratchpad tool to ds4go to improve self-improvement.
While you can do 64G/96G with the Qwen3.8 model, realistically you need 128G. Also, despite tons of playing with local models, the cloud-hosted models on bigger iron are smarter and faster. I don't truly code with my local models and don't recommend this path right now to replace something like Opus/Astra or full-brain DeepSeek4.
The "frontier-ness" of ds4 is great though! It has vast knowledge and thinking capability. Look at that steering video especially. I'm now exploring using ds4 for high-level thinking to create prompts for denser coding models.
Definitely worth looking at if you have only a single 5090 is ninfer, and various hardware specific forks (3090, 4090).
DwarfStar's selling point is that it only supports a small set of carefully chosen models, but it supports them really well.
With llama.cpp once you have the model + the runner on the same box you have from 1 ... 40+ configurations of GPUs (depending on how many gpus you have and how many sub-classes of GPUs) for basic options that will load the model. Then you can start multiplying by different batch sizes (1024, 4096, 8192), CPU thread counts (4, 8, 16), tensor splits (this can get really hairy), P2P enabled/disabled, various caching options, speculative decoding options and so on.
The effect is that you can easily spend a day or more benchmarking. On first run of a new model the software should figure this out by itself.
llama-bench is next to useless for this purpose.
Metal, the primary target, on Macs with 96 GB or more. Smaller machines can use SSD streaming. SSD streaming is also needed in order to run very large models such as full GLM 5.x (not Flash) on 128GB systems
Anyone tested token speeds at less than 96gb RAM on apple?> It must be more a working template for the biggest use cases, without trying to cover every possible setup.
this is such a great quote about building software in the AI age. given that everyone has a different use case, the best way for open source software to be built is to build it for the general use case and specialization could be done afterwards.
ds4 is referring to “dwarfstar” “4” and references DeepSeek V4 most of the time
but its model agnostic-ish
and benchmarks compared to what? what do these large MoE models typically get in tokens per second?
I’m garnering this is just an easier way to load large models per expert on consumer hardware? as opposed to the hackier solutions?
I’m intruiged. Note that the blogpost says 64gb Macs are good minimums while the github says 96gb is a minimum
As Antirez is using a non-frontier model for his work (via locally-running), then I think that is further proof that AI is ready for widespread use in software engineering.
The risks are:
- creating a lot of dead code - ending up with substantial repetitions (this is getting better over time but the risk is definitely still there) - testing only on the happy path rather than all execution paths - mixing current and outdated information resulting in subtly broken code - inability to reason past a certain level of complexity, but no signal that this is the case
For each of these risks there are remedies, one of the more powerful ones for me is the ability to just roll back when things have gone too far off the rails, realize in hindsight what caused it and to retry with a much better initial prompt.
Also the goal of the project is to squeeze the absolute maximum performance and capability possible out of limited hardware resources (compared to clusters of B200s or something).
Does Rust even give you good access to low-level code on different platforms? And if so, how much extra work do you need to do to make it acceptable to the compiler? And is that work worthwhile if you are not going to get the security guarantees of normal Rust code? Is it a worthwhile tradeoff when the goal is performance?
Those are real questions by the way, not rhetorical. If Rust could work well for this type of project then I would like to know.
It's about as good as it can get for this kind of code.
[1] https://www.reddit.com/r/rust/comments/1ixt1ei/zlibrs_is_fas...
This is explicitly an AI-coded project, Antirez argues that LLMs are worse at writing Rust than C because so much high quality systems code (think e.g. sendmail) that ends up in AI training sets is C, not Rust. Another related argument is that the more detailed syntax and compiler feedback found in Rust compared to C are really a negative for LLM workflows.
There's plenty of room to disagree wrt. this of course: without the strong typing checks of Rust around e.g. indirect references, safety and correctness ends up being a global property in typical C programs, and LLMs are terrible wrt. reasoning about global properties. You're better off forcing them to adapt to a different local syntax that does a more complete job of enforcing modularity, since this is comparatively foolproof.
Much easier to work with a language you are most comfortable with right?
The video is in Italian but has an auto-dubbed English audio track: https://www.youtube.com/watch?v=sOt0WpQG5eU\&t=526s
I like C's simplicity so much. The only other language that comes close in simplicity and minimalism is go.
I don't know Rust well enough to understand what else it would provide over the safe memory guarentees?
If.
(Mind you, fuzzing a program with any non-trivial input space can only ever prove that it is unsound. It can never prove that the program is sound.)
"Compressed, not lobotomized."
"Dense, resident, yours."
"Local frontier inference, narrow on purpose."
"ds4 hardware fit: local, streamed and distributed."
The whole site says nothing with so many words. It's also got all the hallmarks of a typical vibe-coded web site (small all-caps text, highly sectioned content, silly animations). Why do people do this? It doesn't impress. In a few years, we'll look back on sites like this like we look at geocities sites today.
We look at geocities with nostalgia, I guess. Ugly as fuck but made with heart when all this thing of the internet was growing.
This slop shit on the other hand... It's cringe right now.
https://github.com/alainnothere/llama.cpp/commits/disk-cache...
And you know it's load bearing each of the load baerings parts that bear some load and load a bear... you fight a bear because it took a load... or something like that...
llama_state_save_file
or llama_state_seq_save_file
and the load equivalents.That was a year or so ago though...
Tensorfold is getting 100%+ speed increases on both prefill and decode for models like Qwen 27B. oMLX has followed them and have had similar improvements in the past week.
There's lots of different techniques like letting CPU help with prefill, DFlash specualtive decoding etc.
I'm really excited for this as I'll be receiving an M5U in about a month. Expect to be running Qwen 4 27B or Flash (it's a 96gb machine), and they may come close in performance to DS 4/4.1 Flash, and should be able to hit 100 tps. Local is really becoming viable, especially considering that GPT 6.1 has been running at 20ish tps the past week.
the ds4 quants were very good beating the unsloth quants https://github.com/michaelasper/benchmarks/blob/main/deepsee...
I mean, I've been using local models on vscode right next to frontier models with ollama for a few months. What's new?
https://aphyr.com/posts/283-jepsen-redis
finally
https://aphyr.com/posts/307-jepsen-redis-redux (see his comments there too)
He had an awesome opportunity to do high concurrency synchronous replication (raft) on top of in-memory databases at a time where ssds were still uncommon, but instead chose to redneck-engineer his own protocol, then double-down that he knows best.
Not that dissimilar to choosing C over rust for familiarity.
Yeah, as usual with devs; pick a tribe then go to insufferable lengths with the newfound and completely unearned superiority complex.
antirez: no, u!
elktown: both sides are tribals with superiority complexes!
Cause he did fuck up, then changed the subject, then claimed the criticism was unfair because he was clearly talking about something else.
It's LLM age. Just port it to Rust if that is so wrong for you. He is doing it for free, no need for arguing about the language he wants to use
Dude asked what's wrong with antirez, I answered
Grandparent wondered whether he was the only one that mistrusts antirez, parent asked for reasons why one would do that, I provided my own.
I used redis multiple times in the past, I have no resentment towards the project or its author(s).
That does not mean that either is above criticism, and I know my freedoms quite well thank you very much.
Feel free to feel insulted!
After I read the discussion between Martin Kleppmann and antirez in which antirez showed comical level of overconfidence while simultaneously demonstrating the complete lack of understanding of the subject (nature of distributed systems), I simply cannot take anything he says or does seriously. For me Redis is just a toy software, which we occasionally use for unimportant features.
- https://martin.kleppmann.com/2016/02/08/how-to-do-distribute... - https://antirez.com/news/101 - https://medium.com/@talentdeficit/redlock-unsafe-at-any-time...
Previous discussions:
- https://news.ycombinator.com/item?id=11065933 - https://news.ycombinator.com/item?id=41315621
@za_creature Provided links to other similar discussions in this thread, so it's not a one-off thing for antirez.
What are we going to name the company, how about Dwarfism 2.0? What happened to 1.0 Jared?