If we get a new programming language not in the training dataset, we could give an LLM a decent compiler with compile errors, and some sample code and it would be able to write code in the new language without training.
2,485 karma · joined November 28, 2013
If we get a new programming language not in the training dataset, we could give an LLM a decent compiler with compile errors, and some sample code and it would be able to write code in the new language without training.
Thought experiment: How effective will 2026 LLMs be for humans in 2526?
It's not game over just because 500 years are missing from the training data. The important question is how well can 2526 humans make culture and knowledge navigable to LLMs via tool calls.
Today's LLMs might need for example sub agents to translate to 2526 English, sub agents to read 2526 docs.
It's _really not clear_ whether 2026 LLMs will be useless. To believe that reflects an enormous misunderstanding.
Can you make any similar guarantees about Drop?
isn’t this very straightforward to do..? I thought batching for Qwen models is already proven out.
> but this would decrease single-session performance even further
Well let’s take Qwen 3.8 27B. Throughput for M3 at 8 agents is 4x compared to single agent. [1]
It’s really not clear to me that 8 concurrent agents at half speed will be worse task completion latency than 1 agent.
And that’s M3 studio benchmarks, not even M5 ultra, and without the many software improvements we will see
If you haven’t tried Qwen 3.8 27B xhigh on a task you might not get the hype. Idk.
If you’ve tried doing this and don’t like it sure, and be specific about what isn’t effective, but let’s not speculate.
[1]: https://omlx.ai/benchmarks/performance/69kzkrv8?utm_source=c...
It’s better than Docker, but it can’t be compared to Firecracker at all. Firecracker actually minimizes the attack surface whereas smolvm does not
I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.
It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.
The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.
An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.
https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyb...
Why? Firecracker mounts very few host systems into the VM, exposing minimal host code to malicious guests. Qemu and smolvm expose much more.
So yeah, smolvm is more like a docker or qemu alternative, definitely useful but NOT relevant to the discussion of sandboxing malicious code
Even if you confine yourself to a dev workstation, having 5 agents concurrently building testing deploying code makes your computer loud and/or hot.
I think you could get by with 1 computer, but it’ll have to have a pretty decent machine.
Between agents running tests, CI, docker image builds, an average $400 mini PC won’t cut it.
Don’t forget also many older mini PCs don’t support KVM. Some newer ones don’t support AVX/ mongodb.
It’s not so easy to buy any old hardware sadly.
https://github.com/agent-substrate/substrate
(For context I built something very similar to this the past 2 weeks for my homelab, trying to solve many of these problems. This comment is an edited version of an unreleased blog post I wrote last week.)
- Run code in secure microVMs or gVisor. Docker is not good enough. Qemu is not good enough. A secure environment for running untrusted code is the bare minimum. I don't see Firecracker in the repo yet, but that's ok the idea is there.
- Fast resumption. In my homelab, time-to-first-message is around 11-12 seconds. That's half setting up the pod, and half resuming the CLI (e.g. `codex resume ..`). Why resuming? In my homelab agents are commonly blocked waiting for CI or waiting for me to approve an action, in this case I stop their container to keep resource usage low. Then for resumption, you definitely don't want to waste the agents time by giving a new ephemeral disk and forcing them to re-clone and re-build. For microVMs this is not actually straightforward, for example Firecracker only allows block devices, so re-attaching an agents disk workspace requires a custom storage interface
- Zero Trust. Codex CLI permissions for example are extremely broken. "Can I run this 500 line long command? or allow any command starting with first 100 chars always?" More reasonable grants are needed.
I don't understand yet how they will surface Zero Trust notifications. In my homelab it's a Forgejo comment linking to an auth service, and a ntfy.sh iOS notification which opens up the auth service.
I don't get why they to restore the RAM of the agent env. Maybe to fully optimize resumption. Idk, I don't have that much RAM in my homelab, my agents use a ton, testing stuff in Chromium making screenshots for me. I can't keep RAM for 100 workspaces from the past 24 hours in RAM.
MITM gateway is very cool.
I'm curious how they will integrate with microVMs. I just wrote yesterday[1] about how there are NO GOOD OPTIONS for this atm. Kata is decent but the attack surface it introduces makes me uncomfortable.
[1]: https://srcreigh.ca/posts/auditable-kata/
But anyway, even if this project is abandoned out of the gate by Google, we should be happy, it sets the bar where it should be. I'm excited to learn how they solved these problems differently than I did.
IIUC, Godel's incompleteness is less about theorems and more about axiomatic systems. Given an axiomatic system, there are statements within it which cannot be proven or disproven. It's relatively unrelated to the platonic ideal of the theorem itself. The statements it considers are axiomatic-system-specific.
Another way to view it is, who cares if we can't prove or disprove "This statement is false". Ok, the axiomatic system is incomplete; fine. What's important is can the system prove a real theorem that I care about.
The busy beaver computability argument addresses these issues. The problem format is always "For Turing machine T with no input, does T halt?". This format can encode many math problems. And we know already that BB(432) is independent of ZF, aka, there is a 432-state TMs which ZF can't prove or disprove the halting behaviour of.
So BB looks at real theorems, ranks them, and we can ask what axiomatic systems can solve them or not. Godel looks at 1 axiomatic system and produces a toy theorem which the system can't solve. That's an extremely important difference!
The core issue is that any fixed LLM can only encode so many axiomatic systems in its states, and the fixed systems implies an upper bound in terms of the BB number which it can solve. Godel is only looking at one system at a time, while BB is a way to use a common problem format to rank every axiomatic system on an infinite number line.
For any finite program (eg some LLMs), there is a true math theorem which they cannot prove or disprove (given fixed input of the statement with no other information sources). If that weren’t true, BB would be computable.
Math is beyond computation. Since AI is just bits in bits out, it has this fundamental limitation.
Any magic of AI systems comes from the transformed meaning of its input data. With fixed weights any LLM is just an artifact. For example a human prompting an LLM constitutes an extra information source, which removes the above limitations. In theory any input from the natural world would remove the limitations too. The natural world is a black box and we don't know what kind of meaning or intelligence could underly it.
"Everything that can be proven" is relative. PA can prove some things, ZF more things. In 200 years we could develop more powerful math foundations which can prove more things. Today's proof verifiers could never prove them, but tomorrow's proof verifiers could. And the cycle repeats.
To put another way, if it were true that some fixed LLM could solve every math problem, it would implement Halt, which is impossible.
The LLM itself is finite, the axioms it knows are fixed, there is an N where BB(N) is independent of those axioms, so the LLM cannot solve it.
The proof verifier uses fixed math axioms. The busy beaver function at high enough N cannot be proven with those axioms.
LLMs are computer programs, so there are math problems which they cannot solve. AKA, ideas which are not possible for them to generate.
The argument for this is that Busy Beaver function is uncomputable. More specifically, some N-state Turing machine requires a proof that it doesn't halt. At some point N is too large and LLM being a computer program, it cannot generate the required proof.
See the Busy Beaver Frontier [1]
This is VERY DIFFERENT from the Halting Problem. In the Halting Problem, we see that no computer can decide whether an arbitrary given input program halts. With the argument above, for a fixed LLM, there is specific math problem which is beyond the capability of proof by the LLM (though other LLMs or humans could perhaps prove it).
Humans are not bound by the argument since we aren't finite computer programs (no proof for this anyways). LLMs which "evolve" over time with input from the natural world also aren't bound by this, since their code is effectively infinite. The argument only applies to a static program with fixed input, no dynamic information sources.
Some people believe in divine inspiration. Maybe you could believe that humans incorporate information from the natural world which LLMs don't have access to. Either of these beliefs would imply that humans have an edge.