A few hours ago GH pages just would not publish at all.. and the status page showed as fine, I guess it's now cascaded into a proper event.
7,439 karma · joined March 31, 2016
A few hours ago GH pages just would not publish at all.. and the status page showed as fine, I guess it's now cascaded into a proper event.
I didn't expect it - but the majority of folks building recipes for DGX Sparks on X/Twitter are now adding/using RigMark as one of their primary benchmarks when sharing figures.
We've gone from a random claim of "I get 65-80 tok/s with this model" to actual receipts that can be traced back to a particular SHA/version of RigMark.
We're on the third version now, and PRs / feedback is welcome.
Where I think we could all use some help going forward is on quality - running local AI sometimes means using compression of weights, which is lossy.
Also: it has to be free.
So yes you're right, people confuse VC backed companies, and vibe-coded pet-projects for sustainable software.
SlicerVM was started in 2022 and internal only, plenty of YouTube videos and such about it - written completely manually from our Actuated work.
There's a free trial for anyone who wants to play about on their Mac or Linux computer, the comment here is from a real user (unprompted) that knew and used free alternatives previously.
No flashy clones of Diablo, or CoD here - just hardware deployed for local AI, in a business, with use cases explained, and how a GPU became 2x Sparks, then 4x.
The switchless NCCL finding in this post is worth 1200-1500 GBP alone, and saves a lot of heat/noise.
Hope the right folks find it here, will try to post again in a day or two if not.
So, the interesting part for me was:
"Who runs the fleet Anthropic, its own E2B, rented, third party"
Like that's the only option, a bespoke platform or a rented SaaS for agent sandboxes.
Disclaimer: I am tooting my own horn here - we built SlicerVM.com (Firecracker + Apple Virtualization framework) so that people don't have to pick between those two options.
Especially when you have your own 1500-7000 USD MacBook right under your fingers.
For remote access, (and this happened by accident one night in January) we built superterm.dev - initially as a tmux manager, then as a way to run and monitor agents. Our team sits with 6-24 long running tasks with dictation available, great mobile experience, and can combine both for secure agents + high amounts of feedback.
If anyone's curious - Superterm has a CE edition - free for unlimited personal use. Slicer has a free trial and then a personal edition is available after that.
Could a 200-400G switch be faster? Potentially. Sparks are compute/memory bound and we've had at least one person with said 1300GBP switch comment and say how our numbers matched his on X.
No "count to 200" as a benchmark, we've had this hooked up to opencode for Slicer/OpenFaaS/inlets development over the course of the week along with various OSS contributions.
And my current favourite for a 2x Spark setup is DS4F all day long NVFP4 (which is a mixed quant, of higher quality than straight 4-bit)
https://x.com/MiaAI_lab/status/2090736338328748220?s=20
> "I've got a confirmation on what model is Ox Alpha, but I can't share it yet. What I can say is this: You should ALL get really excited for this one!!! And it’s NOT what you think it is"
And others have said they have done analysis and found it to be GLM 5.x related.
That said, Mia said "it will be OSS" and "it'll run on 2x DGX Sparks" - well GLM 5.2 can run on 2x Sparks, but slowly and heavily degraded (quant). So doesn't really confirm/deny that suspicion.
https://blog.alexellis.io/the-90s-unix-command-fell-out-of-f...
Can't claim originality though - it was inspired by Sesame - where their models will invoke a search, or check the weather etc, and make a vocalisation to keep you engaged.
Turn taking is one of the hardest things to get right for the exact reasons mentioned - but does seem to be the way that Claude.ai's voice works - in a very obvious way.
Anthropic + OpenAI both rug-pulled voices I liked and got used to and OpenAI really dumbed down their voices at the same time - Arbor went from Estuary English and almost "jack the lad" to some generic English accent. Claude had a Birmingham accent and said things like "shit", ending sentences like "So you're telling me that they asked for a 90% discount yeah?" - then it changed overnight to a mock Derbyshire accent with a dull tone.
ChatGPT's voice also gaslights me for conventional opinions - "my Eastern European neighbour helped me lift a wardrobe upstairs - something you just can't ask your typical neighbour neighbour"... then you get a full on left-leaning lecture from the safety layers rather than a head nod or "what luck!"
Claude + Sesame are nowhere near as overbearing.
In both cases - from edgy and engaging to something that just didn't gel.
The point of making my own assistant is that I can talk for as long as I want, episodic memory is personal and private, there's no "trust me bro, we're a big corporation" vibes.
This was not my first attempt - when I had a bunch of Opus credit around Jan/Feb - I tried really hard and created something that was not good enough. What I have now, is working, and each session is training Claude/Codex on what to tune, and to fix.
"Just had a convo - can you look into what happened?" And if it's one I don't mind sharing with the model - I'll say, "and what did you think of the questions I asked?" Sometimes it'll give a lovely commentary on how the model did.
af_heart is probably the smoothest voice - but yes it's more like another commented - more "StarTrek" than "telesales assistant that pauses and laughs at your jokes".
If you're on a similar path and want something full duplex - the go to solution is PersonaPlex from Nvidia based upon Moshi.
We run quite a few Slicer instances on mini PCs and Ryzen builds - also on Hetzner (and yes ouch 120 EUR / mo up to ~ 550 EUR / mo for 16core / 128GB RAM feels almost unfair)
We’ve been on a similar journey, but came at it from the opposite direction. We started SlicerVM in 2022 after seeing how slow Multipass felt when launching more than one Linux VM, even though it is relatively lean. Tearing them down was slower.. we made it seconds either way for a 30 node cluster and kept it internal until August last year.
With Slicer, microVMs are the native primitive: API launch, guest-agent exec/shell/cp/forward workflows, isolated networking, and agent sandboxes are built into the control plane.
That was not our first use case. Back then we were standing up Kubernetes clusters quickly for OpenFaaS e2e testing and customer scale-out support across multiple machines. The agent/sandbox workflows came naturally after that.
We do see people come over from Proxmox when they want something more directly driven from code, especially with a deeper guest-agent model: exec, file copy, port forwarding, fs watches, etc. When you string it all together it becomes very powerful and what we've gradually dogfooded for our code review bot that started out by using SSH/SFTP to completely native SDK (Go/TS).
One thing I’d separate in the benchmarks is in-guest boot time vs. actual time-to-interactive/useful. For agent-style workloads, the number that tends to matter is: API request made -> VM created/cloned -> network policy applied -> guest agent reachable -> exec/shell/cp/forward works. Snapshot cloning, network device setup, and control-plane readiness all show up there.
TTI can also be moved around depending on tradeoffs: no real init system, snapshot resume, CrosVM-style lower-level primitives, or a VMM built for one narrow job. We use systemd in the guest, so we’re intentionally carrying some weight there.
I also liked that you retained module support for Docker. Supporting Docker, Kubernetes-ish workloads, and eBPF tends to add a lot of useful weight back in.
There’s room for several tools here. The space is moving quickly, and I’m looking forward to seeing which approaches consolidate.
If folks are looking to scratch that microVM, or programmable / bash / agent / SDK driven primitive, you're welcome to check us out and join the Discord.
And ZDR is still data sharing with a third party. This is the essence of an enterprise agreement, it's not allowed, even if they pinkie promise not to store it.
If your customers allow you to share their data with third parties, then ZDR may be an option for you. I am not a laywer.
Where I see ZDR as being more relevant is in protecting your employer's IP - not allowing a missed setting to mean AI labs can train, retain, and publish/resell your work. It's what we'll consider when the subsidies stop being available - open-router, ZDR - but for coding - not for customer data. Very important distinction.
The cache only makes generation fast, it doesn't influence what gets chosen next. The loops that hurt the most (point 2 below) are when the model re-decides to do the same thing in different words, which is much harder to detect automatically. We're experimenting with repetition penalty and turning thinking off to solve for the 1st kind of looping (below)
2. On "why is looping a problem" for us
Practical example, which I covered in the post: "add --json to every command that does a get or list in faas-cli" - this was a small-ish, open source CLI written with Cobra a very common framework.
If I send that to Claude (any of their models) or Codex (GPT), I would have a fully working solution the next time I opened that terminal - a few seconds - a few minutes.
With the local model, when it loops, you get some progress and start working on something else. Come back, maybe even 30 minutes later and see it's been printing the same 5 lines over and over constantly.
Trust is important for a tool like this, that eroded it.
The other type of loop I mention in the blog post is "unable to solve it" loop - Han ran into that more.
"Oh I need to fix the indent from 8 to 5 characters in main.py" "Wait I don't know how to write Python code" "Oh now it's broken and I don't know what to do, maybe I should stop" "Let me edit ... " etc, etc
Dismissed is a strong term, but let me give you some more details.
It took a good 4 minutes plus to load up on the 2x 3090 rig, and served a single request 3 tokens/second slower.
And the worst bit? With all that work - setting it up and tuning it - it still looped. I was hoping "use just vLLM" advice that we get touted everywhere was the silver bullet.
The only thing I'd caution here is that we don't start bashing on llama.cpp like people did with Ollama. It's a very capable tool and for the use-cases we actually want the card for makes more sense.
For a large team replacing their Claude Subs perhaps vLLM is the only option, but you really need to add about 5 more RTX 6000 cards into the mix, so you can load something like GLM 5.2.
It's the right call for concurrent batched serving (barrkel's point downthread is spot on), but for how we use it llama.cpp is still better for us.
The Spark/GX10 route is a genuinely different bet though and appreciate you sharing your numbers. At the time (several months ago) the consensus was that GX10s were for fine-tuning only, and the numbers were severely low.
..and the card was never about replacing a Claude Max sub. For the workloads we actually bought it for, it's giving us 140-200 tok/s (which matters).
The post is not AI generated, I use AI for code generation and write my own articles.
Which part of the post are you struggling with? This is a post describing our own experience and journey. Happy to back up any specific claim.
https://x.com/ggerganov/status/2067539416436867230?s=20
We also use it with 200-256k (native) context length.
The issue could be that folks that don't see looping aren't pushing the model as hard, or as enthusiastically.
We also had far fewer issues when thinking was turned off, than with a reasoning budget capped at 2048.
Some fine-tunes like Qwopus-Coder just seem prone to looping - google it, you'll see plenty of reports, even on Reddit.
For what it's worth seen the RTX 6000 Pro loop even at fp16 on the KV cache - and with vLLM.
The RTX 3090 in question was used from eBay, no way to return it. The RTX 6000 Pro is the "new card" in question here. The 3090s remain an interesting playground for testing things like VFIO passthrough for SlicerVM and other models whilst not interrupting people on the newer card.
In the end, the most stable fix I've found is to install the older proprietary driver and disable the GSP firmware. Have had no issues since.
So "clearly defective hardware" seems like it may not be quite correct. And the thing that kept me coming back - along with not having a suitable replacement - or having to gamble on eBay again was the reliability once it showed up in nvidia-smi.
35B-A3B is what we started out with in the days of only having the 3090, but the quality is not as good, and the speed from the cards we have now can blaze at 130-200 tokens per second of generation with q5 and a full context in fp16.
Not to say that MoEs don't have their place. For people running on unified RAM, they're sometimes the only viable option due to the slowness of dense models.
Why is a dense model slower? All model weights have to be loaded and exercised. Passing through 27B vs 3B (active) is maths. So yes you will always get more tokens per second of generation.
You must (just as we did) evaluate on your own products and daily work. If the MoE gives the results you need with only 3B parameters then you have your answer.
Not prescriptive at all. This is experience based, from the trenches of a actual software business so hopefully a different perspective for folks than "Ran Qwen on my macbook, generated a great python script for me"
> Local models can quickly read and explain codebases, even if they can't write them - this is a superpower
Might have been buried lower down.
And yes latency of local on a fast card with MTP enabled can be blistering 130-200 tokens per second sustained at full context on Q5. About 100+ on Q8.
On tool calling
> Agent Skills can help immensely - we had a local agent set up Slicer completely from scratch on a new mini PC. It even gave feedback on the usability of slicer CLI which we integrated
There's a link to a post showing some examples.
Occasionally, we'll also have the local model _review_ the changes of GPT/Opus - and it can return duds, but also insights the larger model overlooked, or was too intelligent to pick out.
So yes - absolutely blazing fast at understanding a codebase, very good at running skills "cheaply" and could be used with larger models as a "helper" / sub-agent.
As explained in the post - the 3090s were what were the test bed that proved the investment was worth it. Customer support, architecture reviews, telemetry to check license compliance. None of that could be done with online models. The amount of time we can spend going backwards and forth with enterprise customers over email can really amplify costs to our team. A few actual issues we found and fixed were listed on the linked blog post: https://www.openfaas.com/blog/painless-support-with-diag/
Having recovered revenue using it in an airgap, to preserve data agreements was more of a cherry on the cake. No need to worry about the investment, it's covered itself.
Hope that helps.
Then I actually couldn't set the thing up because of the mini HDMI connection - I have a mini to HDMI cable, but to use my portable screen with it I need mini HDMI to MINI HDMI. Don't get me started on micro HDMI - almost everyone of of those connectors I've bought slips off or breaks in the device. Every time I go to set up an RPi5 I end up having to order another one of those tiny connectors.
Full HDMI for all new devices please. Even if the second display can't be connected.
These days a 175 GBP N95 from a no-name Chinese OEM on Amazon, with 16 GB of RAM and a 500GB SATA SSD is way better value and performance - and importantly - zero fuss - standard setup.
If you look at https://slicervm.com you'll see he's copied our terminal animation from the top of the website. Took out a monthly subscription for 1x month, cloned the majority of the UX/DX and way the guest agent works.
Had people reach out and flag it to me and I'm like "yes there's a reason for that"..
I think this is just par for the course in an AI slop world. Nothing to stop people imitating, copying, cloning with a good prompt and partial source / detailed docs available.
SlicerVM (est. 2022) is already used for prime time, not "free as in beer" but has pretty reasonable individual plans that include all features. Shares the core code with actuated. (Creator of both speaking here)
Feel free to take a look and see if gives you a little more than the others you mentioned. If not no problems, I realise some folks prefer free stuff.
I hadn't heard of "Sun Ray" until today, but it reminds me a lot of the idea behind Linux Terminal Server Project (LTSP) - which I used on our school's IT lab back then at a teen. Set up an old i386 machine with the various netbooting daemons. Then on each host - boot from floppy disk, remove disk, insert in next machine until 20 hosts were running from that poor old hard drive.
The nice thing was that the installed OS on each was unaffected, and each machine was running X11 over the network.
Seems like those solutions were optimising for a time where hardware was overly expensive.
The team say it beats 3.5 27B in numerous benchmarks. But it's still a MoE with 3B active per token, so we'll see!
I hope there's a 3.6 27B coming in short order, or even a 31B.
Gemma 4 looked great, but I've not got it working well with Claude code (Qwen just works) - Gemma 4 _does_ work with opencode though.