Running Kimi K3 on a M1 Max
github.com
github.com
Soon decent speed across two Mac Studios with 512GB of RAM.
And it gets worse with every token.
192GiB Gorgon Halo systems will be an interesting future target for this model, the best you can do with 128GiB or less is probably to push batching higher in order to amortize the weights traffic over multiple inferences - which of course will sink single-session speeds even lower for a modest gain in total throughput.
For AI agents this would take a year
This is like 98%.
If you already committed to your hardware, huge models at single or sub-digit TPS are still useful. LLM-as-judge is a good use case.
If your machine is gonna be idle overnight (and the power efficiency is excellent here), why pay openrouter if it’s not an interactive workload?
To really go fast you’d probably have to do PCB layout and do like 256 or 1024 chips in parallel with a fast SRAM aggregation buffer feeding a GPU or TPU rig.
Or could you do the same with custom layout of cheap slower RAM?
I wonder if anyone is doing this? You would flash in a model and then just run it. It would need RAM for context but much less of it.
Optane would actually be useful in this era. Intel was ahead of their time.
We've made some updates to deltafin since this was shared 5 days ago. I wonder if you'd be able to get better speeds now? https://github.com/gavamedia/deltafin
It reminds me of when Willow Garage chose to name their bot the TurtleBot, because if they named it anything else, people would think it was fast and capable. But when they called it Turtle Bot, people just kind of liked it and were satisfied with what it did.
At the level of Kimi 3, I probably can code only about 1,000 good tokens per day, too. (thankfully coding isn't my job)
Kimi Pen Pal. Bring back lettets and postcards. Do OCR, and use one of those 3D printer-like pen plotters write the model output as a letter.
Challenge would be automating the opening and OCR preparation, and the folding and mailing of the return letter. But given it's done commercially it should be possible.
Works in mutt; other MUAs may vary.
For example there's a billboard on 101 for Poppy Bank offering 4% interest and I asked my daily LLM to look into it through a voice note. The next morning, lo and behold it says it's an advertised rate, hard to actually get, and businesses aren't eligible. OK, done. Better than getting the response while I'm driving. And even if I wasn't driving, there's a level of, how do you say, it's easy to drop it the next day when I see it vs getting engrossed in the research.
Also its answering the question of what gonna happen if you wake up tomorrow and datacenters are gone. Or internets are gone.
Some people on our globe live in countries with no internet whatsoever. Of course most of them dont have Macbook with 64GB RAM either, but it's much much easier to get than internet connection or rack of GB200.
SOTA LLMs are efficiently compression of all the knowkedge humanity has built. Having ability to run it at home to extract said knowledge is important no matter the speed.
But also that's a pretty extreme hypothetical. Imagine the polymarket on that.
People tend to think they are the thriving hero but most won't survive the first year.
I'd much rather have a USB stick of prepper disk. It works on any shitty machine I find, and it's not going to take days for an answer on how to make potash or fix a solar panel.
16tk/s... Then 3 tks per minute. Then someone else posted 0.3tk/s.
(note. I am not a believer in AGI)
"useful" is highly contextual. The clock of the long "now" is not useful in the sense you mean, to synchronise your wristwatch. I'm still glad it exists.
Need to justify buying an expensive rig that doesn't do what you expected.
Specifically thinking the people they could do something AI with cpu, and realizing it isn't feasible. Happened at my fortune 20 company. They had to get approvals and ofc it was useless. Plenty people tried to explain, but they were the principle engineer, and out ranked everyone.
"It's not going to work", the topic changed, and we never spoke about it again.
- bunch of people only ~4 years ago
Consider this like if it were the first test
The point of these engineering tricks is to see the envelope of what's possible. You can use these tricks to both run a bigger model on smaller hardware or run a smaller model on smaller hardware.
idk how people access (soldout) and even afford 512GB RAM MacStudio's. Isn't it $40k or so?
I do really care to know: "Is there a new and more beefy solution we don't know about yet?"
I don't know if I'd call this "running"
UPD: I know it's not the same at all, just the reversal of units that gets me
It’d be like thinking as slowly as Ents talk to each other in Lord of the Rings.
Local AI on your device seems like a much more likely future to me than datacenters in space. For inference at least, training is another story.
0.01 tokens per second means 1 million tokens ($3 worth of API usage [1]) takes 3.2 YEARS.
Looking more broadly though, a model I can run on my laptop (Gemma 4) is ~4 points away from GPT-5.3 codex or Sonnet 4.5 on arena.ai LLM leaderboard. Those models were SOTA less than a year ago.
Quick mission statement, since it explains everything below: Deltafin runs the full, never-pruned, 2.8-trillion-parameter Kimi K3... all 16 experts, every token, no shortcuts, nothing re-encoded down to "close enough." Speed is not allowed to come from quality.
A few similar K3 projects have shown up claiming to be the same thing, just faster. But they all get there by shrinking the model - quantizing it down, trading fidelity for speed, calling the results close enough. Those are great projects, and have real value - but they're very different projects. Deltafin's purpose remains the real full K3, unaltered and uncut, as fast* as possible.
* it's not fast
Next, an update since this was posted five days ago. Thanks to everyone who filed issues, including a couple of you running M3 Ultra and M5 Max boxes, which helped a lot!Since this was posted, throughput on the M1 Max laptop went from 0.0141 tok/s to 0.2847 tok/s - about 20x, and still climbing. Most of that came from routine hardening: fixing a storage race, a rust/c++ rewrite, and tolerating a float32 rounding drift that looked like a routing bug until it wasn't.
None of that is a moonshot (get it?). It's the boring, necessary work of making something this huge survive on consumer hardware.
Now one huge thing to get out of the way, because this is about half the comments: is this a viable business product, the kind of thing that makes people in suits happy because it helps them sell widgets? No. My interest in computing goes beyond that, and I suspect a lot of yours does too. I'd like to think that's still what hacker news is largely about.
A research project usually isn't the useful thing itself. It's what falls out of chasing it.
It is genuinely hard to squeeze even a small speed gain out of a model this size on a machine this small. Yes, an M1 Max is a good laptop. But Kimi's own recommended setup is 16 nodes at ~300GB of VRAM each. That's $2,000,000 of hardware. Next to that, this laptop is an ant lifting a car. And that is the point of this project - the insane challenge.
My hope is that if we push as hard as we can to get the largest public model ever released running on the smallest machine that can plausibly hold it, some of what falls out along the way could be useful elsewhere... maybe to smaller models, on smaller machines, running faster than they do today.
But first I wanted to find out whether this was even possible at all. K3 is the largest public model that exists. Nobody had run something this size on a laptop before. The closest comparable projects like colibri and DwarfStar ran quantized models at half this size. Turns out: yes, it's doable, but slow af.
But I've spent a long time frustrated with the quality ceiling on local models, and quantization is a big part of that ceiling. The hallucinations. The reasoning that almost holds together, but then doesn't.
So I'm done compromising on that axis. This isn't going to be another quantized, slightly-worse copy of the real model wearing its name.
Maybe no techniques we’re using are groundbreaking at all in the end. Bummer. But, at least I tried. Try and fail at life, people. That’s okay. Just try, and at least you’ll have lived a little. It's a decent way to spend the time you get.
If people want to keep following this, I'll post updates here instead of just on Reddit like I have been.
> It is not fast — about 16 seconds per token on our M1 Max