Thank you for your work!
Thank you for your work!
For me the nvidia driver just keeps waking up the system instantly - but my setup is deviating from the upstream flake in a few ways, so I'm just wondering if it's worth setting up the system from scratch if it's working for other people.
Other than that, can fully second that the flake is working great. Only gotcha is that CUDA-enabled packages (including Firefox) require using the flox binary cache unless you want to compile them from source, but then the package versions can lag behind a bit (and debugging nix cache issues is surprisingly difficult).
I’ve been running a custom VLLM image with b12x as well as nvfp4_ds_mla.
I would say it’s quite fantastic in day to day, I use it mostly in Hermes and sometimes for coding.
I have qwen 3.6 27b on an rtx 6000 pro as well so I use that as a workhorse in pi with DS as a reviewer/planner.
[0] https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Edit: I think you may have misread my post. k3s is NOT kimi k3, and I did mention I was running deepseek.
I have to do like, low paying web dev to fund it, so I can't promise timelines, but it's coming and it will be free to anyone and fast as fuck.
I estimate about 20 bucks an hour at current spot rates in the hundreds if not thousands of tokens per second.
you're asking me how a watch works. let's just try to keep an eye on the time.
federal felony prosecution.
i'm not quite ready to OSS the whole thing, it's got a few rough edges, but if anyone wants to alpha test, caveat emptor and it's yours.