HNHacker News
TopNewBestAskShowJobs

alexellisuk

7,439 karma · joined March 31, 2016

Founder OpenFaaS + inlets + actuated. https://www.alexellis.io/
submissionscomments
alexellisuk··on GitHub Incident with Git Operations, Pull Requests and Actions
I thought it was odd that we'd not had one of these for a few weeks.

A few hours ago GH pages just would not publish at all.. and the status page showed as fine, I guess it's now cascaded into a proper event.

alexellisuk··on Show HN: RigMark benchmarks local AI the way coding agents use it
Having been frustrated with benchmaxxing figures shared on X (as described here: https://blog.alexellis.io/how-and-why-we-bought-4-dgx-sparks...), I developed an independent benchmark to get stable and comparable numbers for Prose, Code and Structured output, plus Prefill.

I didn't expect it - but the majority of folks building recipes for DGX Sparks on X/Twitter are now adding/using RigMark as one of their primary benchmarks when sharing figures.

We've gone from a random claim of "I get 65-80 tok/s with this model" to actual receipts that can be traced back to a particular SHA/version of RigMark.

We're on the third version now, and PRs / feedback is welcome.

Where I think we could all use some help going forward is on quality - running local AI sometimes means using compression of weights, which is lossy.

alexellisuk··on 5x faster Edge Functions: V8 isolates to Firecracker MicroVMs
The world: "build a secure, enterprise-ready microVM automation solution - work on it full time, and pay salaries for the staff that work on it"

Also: it has to be free.

So yes you're right, people confuse VC backed companies, and vibe-coded pet-projects for sustainable software.

SlicerVM was started in 2022 and internal only, plenty of YouTube videos and such about it - written completely manually from our Actuated work.

There's a free trial for anyone who wants to play about on their Mac or Linux computer, the comment here is from a real user (unprompted) that knew and used free alternatives previously.

alexellisuk··on How and Why We Bought 4x DGX Sparks
Author here:

No flashy clones of Diablo, or CoD here - just hardware deployed for local AI, in a business, with use cases explained, and how a GPU became 2x Sparks, then 4x.

The switchless NCCL finding in this post is worth 1200-1500 GBP alone, and saves a lot of heat/noise.

Hope the right folks find it here, will try to post again in a day or two if not.

alexellisuk··on The VMs Powering Mobile Agents (Instinct, Claude Code)
I was hoping for a whole catalogue of providers - doesn't amp have their own offering here "Orbs"?

So, the interesting part for me was:

"Who runs the fleet Anthropic, its own E2B, rented, third party"

Like that's the only option, a bespoke platform or a rented SaaS for agent sandboxes.

Disclaimer: I am tooting my own horn here - we built SlicerVM.com (Firecracker + Apple Virtualization framework) so that people don't have to pick between those two options.

Especially when you have your own 1500-7000 USD MacBook right under your fingers.

For remote access, (and this happened by accident one night in January) we built superterm.dev - initially as a tmux manager, then as a way to run and monitor agents. Our team sits with 6-24 long running tasks with dictation available, great mobile experience, and can combine both for secure agents + high amounts of feedback.

If anyone's curious - Superterm has a CE edition - free for unlimited personal use. Slicer has a free trial and then a personal edition is available after that.

alexellisuk··on Show HN: Our GLM-5.3 Flash Switchless recipe is now out for 4x DGX Sparks
This is the culmination of several weeks' worth of R&D and testing. It builds on OSS components and we credit others.

Could a 200-400G switch be faster? Potentially. Sparks are compute/memory bound and we've had at least one person with said 1300GBP switch comment and say how our numbers matched his on X.

No "count to 200" as a benchmark, we've had this hooked up to opencode for Slicer/OpenFaaS/inlets development over the course of the week along with various OSS contributions.

And my current favourite for a 2x Spark setup is DS4F all day long NVFP4 (which is a mixed quant, of higher quality than straight 4-bit)

alexellisuk··on Ox Alpha
The "mia" persona on X has a specific vaguepost:

https://x.com/MiaAI_lab/status/2090736338328748220?s=20

> "I've got a confirmation on what model is Ox Alpha, but I can't share it yet. What I can say is this: You should ALL get really excited for this one!!! And it’s NOT what you think it is"

And others have said they have done analysis and found it to be GLM 5.x related.

That said, Mia said "it will be OSS" and "it'll run on 2x DGX Sparks" - well GLM 5.2 can run on 2x Sparks, but slowly and heavily degraded (quant). So doesn't really confirm/deny that suspicion.

alexellisuk··on Finger: the 1971 social network that never died
I also very fond memories of the finger command - though mainly used on MUDs and the odd Linux host:

https://blog.alexellis.io/the-90s-unix-command-fell-out-of-f...

https://news.ycombinator.com/item?id=44943313

alexellisuk··on Januscape vulnerability CVE-2026-53359 mitigations available (KVM breakout)
Additionally: https://www.openwall.com/lists/oss-security/2026/07/06/7
alexellisuk··on GPT‑Live
Funnily enough - I built this (delegation) over the weekend with Fable for a local voice chat running 100% on local LLMs, Parakeet and Kokoro. I say "...ask the thinking model..." and that redirects it to Qwen 3.6 27B on vLLM.

Can't claim originality though - it was inspired by Sesame - where their models will invoke a search, or check the weather etc, and make a vocalisation to keep you engaged.

Turn taking is one of the hardest things to get right for the exact reasons mentioned - but does seem to be the way that Claude.ai's voice works - in a very obvious way.

Anthropic + OpenAI both rug-pulled voices I liked and got used to and OpenAI really dumbed down their voices at the same time - Arbor went from Estuary English and almost "jack the lad" to some generic English accent. Claude had a Birmingham accent and said things like "shit", ending sentences like "So you're telling me that they asked for a 90% discount yeah?" - then it changed overnight to a mock Derbyshire accent with a dull tone.

ChatGPT's voice also gaslights me for conventional opinions - "my Eastern European neighbour helped me lift a wardrobe upstairs - something you just can't ask your typical neighbour neighbour"... then you get a full on left-leaning lecture from the safety layers rather than a head nod or "what luck!"

Claude + Sesame are nowhere near as overbearing.

In both cases - from edgy and engaging to something that just didn't gel.

The point of making my own assistant is that I can talk for as long as I want, episodic memory is personal and private, there's no "trust me bro, we're a big corporation" vibes.

This was not my first attempt - when I had a bunch of Opus credit around Jan/Feb - I tried really hard and created something that was not good enough. What I have now, is working, and each session is training Claude/Codex on what to tune, and to fix.

"Just had a convo - can you look into what happened?" And if it's one I don't mind sharing with the model - I'll say, "and what did you think of the questions I asked?" Sometimes it'll give a lovely commentary on how the model did.

af_heart is probably the smoothest voice - but yes it's more like another commented - more "StarTrek" than "telesales assistant that pauses and laughs at your jokes".

If you're on a similar path and want something full duplex - the go to solution is PersonaPlex from Nvidia based upon Moshi.

alexellisuk··on MicroVMs: Run isolated sandboxes with full lifecycle control
Yeah, I'm surprised Justin posted this like it was new(s). Wasn't it doing the rounds on the 22nd when it launched?
alexellisuk··on MicroVMs: Run isolated sandboxes with full lifecycle control
For self-hosting, have a look at what we're building with SlicerVM.com (disclosure: I'm the founder). Also runs just as well on Apple Silicon.

We run quite a few Slicer instances on mini PCs and Ryzen builds - also on Hetzner (and yes ouch 120 EUR / mo up to ~ 550 EUR / mo for 16core / 128GB RAM feels almost unfair)

alexellisuk··on Running MicroVMs in Proxmox VE, the Easy Way
This is clever work, especially given that Proxmox is already a very viable VMware replacement and wasn’t originally designed around microVMs as the primary abstraction. I’m glad this is working well for you.

We’ve been on a similar journey, but came at it from the opposite direction. We started SlicerVM in 2022 after seeing how slow Multipass felt when launching more than one Linux VM, even though it is relatively lean. Tearing them down was slower.. we made it seconds either way for a 30 node cluster and kept it internal until August last year.

With Slicer, microVMs are the native primitive: API launch, guest-agent exec/shell/cp/forward workflows, isolated networking, and agent sandboxes are built into the control plane.

That was not our first use case. Back then we were standing up Kubernetes clusters quickly for OpenFaaS e2e testing and customer scale-out support across multiple machines. The agent/sandbox workflows came naturally after that.

We do see people come over from Proxmox when they want something more directly driven from code, especially with a deeper guest-agent model: exec, file copy, port forwarding, fs watches, etc. When you string it all together it becomes very powerful and what we've gradually dogfooded for our code review bot that started out by using SSH/SFTP to completely native SDK (Go/TS).

One thing I’d separate in the benchmarks is in-guest boot time vs. actual time-to-interactive/useful. For agent-style workloads, the number that tends to matter is: API request made -> VM created/cloned -> network policy applied -> guest agent reachable -> exec/shell/cp/forward works. Snapshot cloning, network device setup, and control-plane readiness all show up there.

TTI can also be moved around depending on tradeoffs: no real init system, snapshot resume, CrosVM-style lower-level primitives, or a VMM built for one narrow job. We use systemd in the guest, so we’re intentionally carrying some weight there.

I also liked that you retained module support for Docker. Supporting Docker, Kubernetes-ish workloads, and eBPF tends to add a lot of useful weight back in.

There’s room for several tools here. The space is moving quickly, and I’m looking forward to seeing which approaches consolidate.

If folks are looking to scratch that microVM, or programmable / bash / agent / SDK driven primitive, you're welcome to check us out and join the Discord.

alexellisuk··on Local Qwen isn't a worse Opus, it's a different tool
Thanks for the comment ZDR is mentioned in the post - in particular many the coding plans that are not from the two major leaders have questionable IP/ownership claims on inputs/outputs :)

And ZDR is still data sharing with a third party. This is the essence of an enterprise agreement, it's not allowed, even if they pinkie promise not to store it.

If your customers allow you to share their data with third parties, then ZDR may be an option for you. I am not a laywer.

Where I see ZDR as being more relevant is in protecting your employer's IP - not allowing a missed setting to mean AI labs can train, retain, and publish/resell your work. It's what we'll consider when the subsidies stop being available - open-router, ZDR - but for coding - not for customer data. Very important distinction.

alexellisuk··on Local Qwen isn't a worse Opus, it's a different tool
1. On the technical:

The cache only makes generation fast, it doesn't influence what gets chosen next. The loops that hurt the most (point 2 below) are when the model re-decides to do the same thing in different words, which is much harder to detect automatically. We're experimenting with repetition penalty and turning thinking off to solve for the 1st kind of looping (below)

2. On "why is looping a problem" for us

Practical example, which I covered in the post: "add --json to every command that does a get or list in faas-cli" - this was a small-ish, open source CLI written with Cobra a very common framework.

If I send that to Claude (any of their models) or Codex (GPT), I would have a fully working solution the next time I opened that terminal - a few seconds - a few minutes.

With the local model, when it loops, you get some progress and start working on something else. Come back, maybe even 30 minutes later and see it's been printing the same 5 lines over and over constantly.

Trust is important for a tool like this, that eroded it.

The other type of loop I mention in the blog post is "unable to solve it" loop - Han ran into that more.

"Oh I need to fix the indent from 8 to 5 characters in main.py" "Wait I don't know how to write Python code" "Oh now it's broken and I don't know what to do, maybe I should stop" "Let me edit ... " etc, etc

alexellisuk··on Local Qwen isn't a worse Opus, it's a different tool
vLLM is great at continuous batching and model serving in production, but it's a very different beast and much less versatile for the prosumer category (where we sit for our usage)

Dismissed is a strong term, but let me give you some more details.

It took a good 4 minutes plus to load up on the 2x 3090 rig, and served a single request 3 tokens/second slower.

And the worst bit? With all that work - setting it up and tuning it - it still looped. I was hoping "use just vLLM" advice that we get touted everywhere was the silver bullet.

The only thing I'd caution here is that we don't start bashing on llama.cpp like people did with Ollama. It's a very capable tool and for the use-cases we actually want the card for makes more sense.

For a large team replacing their Claude Subs perhaps vLLM is the only option, but you really need to add about 5 more RTX 6000 cards into the mix, so you can load something like GLM 5.2.

alexellisuk··on Local Qwen isn't a worse Opus, it's a different tool
We did run vLLM on the 3090s — measured ~3 tok/s slower on generation for our single-to-few-user pattern, plus less flexibility on quant and slower startup (actual minutes vs single digit seconds). We may do more with it again in the future - there isn't unlimited time for us to tinker, I'm sharing our journey (so far) and reasoning.

It's the right call for concurrent batched serving (barrkel's point downthread is spot on), but for how we use it llama.cpp is still better for us.

The Spark/GX10 route is a genuinely different bet though and appreciate you sharing your numbers. At the time (several months ago) the consensus was that GX10s were for fine-tuning only, and the numbers were severely low.

..and the card was never about replacing a Claude Max sub. For the workloads we actually bought it for, it's giving us 140-200 tok/s (which matters).

alexellisuk··on Local Qwen isn't a worse Opus, it's a different tool
Fair enough, that sentence was fairly compressed. I’ve reworded it - the meaning remains the same.

The post is not AI generated, I use AI for code generation and write my own articles.

Which part of the post are you struggling with? This is a post describing our own experience and journey. Happy to back up any specific claim.

alexellisuk··on Local Qwen isn't a worse Opus, it's a different tool
I think that's quite telling Gorgi replied that he uses Qwen with 131k context.

https://x.com/ggerganov/status/2067539416436867230?s=20

We also use it with 200-256k (native) context length.

The issue could be that folks that don't see looping aren't pushing the model as hard, or as enthusiastically.

We also had far fewer issues when thinking was turned off, than with a reasoning budget capped at 2048.

Some fine-tunes like Qwopus-Coder just seem prone to looping - google it, you'll see plenty of reports, even on Reddit.

For what it's worth seen the RTX 6000 Pro loop even at fp16 on the KV cache - and with vLLM.

alexellisuk··on Local Qwen isn't a worse Opus, it's a different tool
Ha, you underestimate how dogged you need to be to get this stuff working well.

The RTX 3090 in question was used from eBay, no way to return it. The RTX 6000 Pro is the "new card" in question here. The 3090s remain an interesting playground for testing things like VFIO passthrough for SlicerVM and other models whilst not interrupting people on the newer card.

In the end, the most stable fix I've found is to install the older proprietary driver and disable the GSP firmware. Have had no issues since.

So "clearly defective hardware" seems like it may not be quite correct. And the thing that kept me coming back - along with not having a suitable replacement - or having to gamble on eBay again was the reliability once it showed up in nvidia-smi.

alexellisuk··on Local Qwen isn't a worse Opus, it's a different tool
The important thing about MoEs which I mention in the conclusion is that they carry fewer (way fewer) active tokens during inference/generation.

35B-A3B is what we started out with in the days of only having the 3090, but the quality is not as good, and the speed from the cards we have now can blaze at 130-200 tokens per second of generation with q5 and a full context in fp16.

Not to say that MoEs don't have their place. For people running on unified RAM, they're sometimes the only viable option due to the slowness of dense models.

Why is a dense model slower? All model weights have to be loaded and exercised. Passing through 27B vs 3B (active) is maths. So yes you will always get more tokens per second of generation.

You must (just as we did) evaluate on your own products and daily work. If the MoE gives the results you need with only 3B parameters then you have your answer.

Not prescriptive at all. This is experience based, from the trenches of a actual software business so hopefully a different perspective for folks than "Ran Qwen on my macbook, generated a great python script for me"

alexellisuk··on Local Qwen isn't a worse Opus, it's a different tool
One of the things I mentioned in the post:

> Local models can quickly read and explain codebases, even if they can't write them - this is a superpower

Might have been buried lower down.

And yes latency of local on a fast card with MTP enabled can be blistering 130-200 tokens per second sustained at full context on Q5. About 100+ on Q8.

On tool calling

> Agent Skills can help immensely - we had a local agent set up Slicer completely from scratch on a new mini PC. It even gave feedback on the usability of slicer CLI which we integrated

There's a link to a post showing some examples.

Occasionally, we'll also have the local model _review_ the changes of GPT/Opus - and it can return duds, but also insights the larger model overlooked, or was too intelligent to pick out.

So yes - absolutely blazing fast at understanding a codebase, very good at running skills "cheaply" and could be used with larger models as a "helper" / sub-agent.

alexellisuk··on Local Qwen isn't a worse Opus, it's a different tool
Author here. Thanks for the question. I'll answer assuming this is a question you have for me.

As explained in the post - the 3090s were what were the test bed that proved the investment was worth it. Customer support, architecture reviews, telemetry to check license compliance. None of that could be done with online models. The amount of time we can spend going backwards and forth with enterprise customers over email can really amplify costs to our team. A few actual issues we found and fixed were listed on the linked blog post: https://www.openfaas.com/blog/painless-support-with-diag/

Having recovered revenue using it in an airgap, to preserve data agreements was more of a cherry on the cake. No need to worry about the investment, it's covered itself.

Hope that helps.

alexellisuk··on Local Qwen isn't a worse Opus, it's a different tool
Hi - the author of the post here. I wanted to write up something that was a bit more than "Qwen is the goat" or "Cancelled Claude, run everything local now" or even "The model organised my CD collection, so it's great" - this is real from the trenches stuff. And what the model/harness did achieve may surprise you - along with how it ended up paying for itself.
alexellisuk··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
What quant?
alexellisuk··on News about Raspberry Pi 6 and Microcontroller Development
I was thinking about the RPi 6 yesterday whilst realising I couldn't set up my RPi Zero 2W anymore - the OS has become burdensome - tied strictly to an imager, that gives me an allergic reaction. Yes - they did all this for the uninitiated - but for Raspberry Pi OS Lite - bring back this experience: dd the image, write ssh into the boot drive, SSH in - change password, fully set up in almost zero fuss or effort.

Then I actually couldn't set the thing up because of the mini HDMI connection - I have a mini to HDMI cable, but to use my portable screen with it I need mini HDMI to MINI HDMI. Don't get me started on micro HDMI - almost everyone of of those connectors I've bought slips off or breaks in the device. Every time I go to set up an RPi5 I end up having to order another one of those tiny connectors.

Full HDMI for all new devices please. Even if the second display can't be connected.

These days a 175 GBP N95 from a no-name Chinese OEM on Amazon, with 16 GB of RAM and a 500GB SATA SSD is way better value and performance - and importantly - zero fuss - standard setup.

alexellisuk··on Show HN: Tilde.run – Agent sandbox with a transactional, versioned filesystem
Not a surprise at all.

If you look at https://slicervm.com you'll see he's copied our terminal animation from the top of the website. Took out a monthly subscription for 1x month, cloned the majority of the UX/DX and way the guest agent works.

Had people reach out and flag it to me and I'm like "yes there's a reason for that"..

I think this is just par for the course in an AI slop world. Nothing to stop people imitating, copying, cloning with a good prompt and partial source / detailed docs available.

alexellisuk··on Show HN: Tilde.run – Agent sandbox with a transactional, versioned filesystem
It's interesting to see this one launch (yes yet another sandbox.. I was getting worried we'd not seen one for a few days)

SlicerVM (est. 2022) is already used for prime time, not "free as in beer" but has pretty reasonable individual plans that include all features. Shares the core code with actuated. (Creator of both speaking here)

Feel free to take a look and see if gives you a little more than the others you mentioned. If not no problems, I realise some folks prefer free stuff.

alexellisuk··on Setting up a Sun Ray server on OpenIndiana Hipster 2025.10
Interesting to see it all play out through the post.. OpenIndiana is virtualized, the Sun Ray connects to it and runs like a thin client.

I hadn't heard of "Sun Ray" until today, but it reminds me a lot of the idea behind Linux Terminal Server Project (LTSP) - which I used on our school's IT lab back then at a teen. Set up an old i386 machine with the various netbooting daemons. Then on each host - boot from floppy disk, remove disk, insert in next machine until 20 hosts were running from that poor old hard drive.

The nice thing was that the installed OS on each was unaffected, and each machine was running X11 over the network.

Seems like those solutions were optimising for a time where hardware was overly expensive.

alexellisuk··on Qwen3.6-35B-A3B:Now Open-Source
The unsloth quants are already available: https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF

The team say it beats 3.5 27B in numerous benchmarks. But it's still a MoE with 3B active per token, so we'll see!

I hope there's a 3.6 27B coming in short order, or even a 31B.

Gemma 4 looked great, but I've not got it working well with Claude code (Qwen just works) - Gemma 4 _does_ work with opencode though.

Page 1 of 27Next →