WASM is a great system, but quite complex -- the spec for Mog is roughly 100x smaller.
140 karma · joined January 25, 2010
AI from first principles.
Previously:
Principal | Martian Engineering (https://martian.engineering)
We solve hard problems in software engineering. Small, tight team of ex-YC founders and a team of 10-20yrs+ senior SWEs with experience in hardware, languages, compilers, network protocols, operating systems, security, data engineering, web development.
Previously:
CTO -- Urbit Foundation
Kernel Engineer -- Tlon Corporation
Tech Lead, Data Acquisition -- 3Scan, Inc.
Cofounder -- Zigfu (YC S11)
https://github.com/belisarius222
https://x.com/rovnys
WASM is a great system, but quite complex -- the spec for Mog is roughly 100x smaller.
It's also designed to be run in an event loop. I've tested this with Bun's event loop that runs TypeScript. I haven't tried it with other async runtimes, but it should be doable.
As for the browser, I haven't tried it, but you might be able to compile it to WASM -- the async stuff would be the hardest part of that, I suspect. Could be cool!
Since it's new, Mog will likely not yet beat existing systems at basically anything. Its potential lies in having better performance and a much smaller total system footprint and complexity than the alternatives. WASM is generally interpreted -- you can compile it, but it wasn't really designed for that as far as I know.
More generally, I think new execution environments are good opportunities for new languages that directly address the needs of that environment. The example that comes to mind is JavaScript, which turned webpages into dynamically loaded applications. AI agents have such heavy usage and specific problems that a language designed to be both written and executed by them is worth a shot in my opinion.
JIT means the code is interpreted until some condition kicks in to trigger compilation. This is obviously common and provides a number of advantages, but it has downsides too: 1) Code might run slowly at first. 2) It can be difficult to predict performance -- when will the JIT kick in? How well will it compile the code?
With Mog, you do have to pay the up-front cost of compiling the program. However, what I said about "no process startup cost" is true: there is no other OS process. The compiler runs in process, and then the compiled machine code is loaded into the process. Trying to do this safely is an unusual goal as far as I can tell. One of the consequences of this security posture is that the compiler and host become part of the trusted computing base. JITs are not the simplest things in the world, and not the easiest things to keep secure either. The Mog compiler is written entirely in safe Rust for this reason.
This up-front compilation cost is paid once, then the compiled code can be reused. If you have a pre-tool-use hook, or some extension to the agent itself, that code runs thousands of times, or more. Ahead-of-time compilation is well-suited for this task.
If this is used for writing a script that agent runs once, then JIT compilation might turn out to be faster. But those scripts are often short, and our compiler is quite fast for them as it is in the benchmarking that I've done -- there are benchmarking scripts in the repo, and it would be interesting to extend them to map out this landscape more.
Also, in my experience, in this scenario, the vast majority of the total latency of waiting for the agent to do what you asked it is due to waiting for an LLM to finish responding, not compiling or executing the script it generated. So I've prioritized the end-to-end performance of Mog code that runs many times.
Of course getting permissions to work well might be easier said than done, but I like this direction.
It's quite possible that's wrong. If so, I would write llm_reduce like this: it would spawn a sub-task for every pair of elements in the list, which would call an LLM with a prompt telling it how to combine the two elements into one. The output type of the reduce operation would need to be the same as the input type, just like in normal map/reduce. This allows for a tree of operations to be performed, where the reduction is run log(n) times, resulting in a single value.
That value should probably be loaded into the LCM database by default, rather than putting it directly into the model's context, to protect the invariant that the model should be able to string together arbitrarily long sequences of maps and reduces without filling up its own context.
I don't think this would be hard to write. It would reuse the same database and parallelism machinery that llm_map and agentic_map use.
Urbit is a long-term project starting from the foundations and working up the stack. It'll be stable soon, but it probably won't be ready to run something as heavy as a modern web browser for a while. That being said, it does work, and it's fun to play with as-is.
Urbit is a program you can run on Linux or MacOS intended to provide a complete personal computing experience on its own. It runs as a virtual machine for now, although it could run as a unikernel on bare metal (good project for a contributor who's interested!).
This VM acts like an operating system, in the sense that it loads and runs other applications within itself, and in the sense that it presents an application switcher and overall system management tools to the user.
This VM is designed from scratch to be as simple as possible, based on the thesis that the reason everyone has thought of a personal server but nobody runs one is that it's too complicated to do your own sysadmin.
Why is it complicated to do your own sysadmin? Because Linux is 15 million lines of code, and then there are tons of layers on top of that. What percentage of programmers even know how the internet works? A fair number of programmers have a decent sense for some corner of the modern computing world, but even seasoned professionals don't usually know the full structure of the digital world. How does BGP interact with the IP protocol? How do you make sure fsync() actually did what you wanted it to do? How does Linux overcommit_memory work? etc.
Urbit is weird, but that's mostly because it's a parallel universe of computing, not because it's inherently crazier than the alternative. We all have Stockholm Syndrome about 'ls -alH', and don't tell me 'grep' is an intuitive name.
In fact, there are very few basic building blocks in Urbit: binary trees of integers, the idea of a persistent event-log-based computer, and cryptographic identities. Pretty much everything is constructed out of those components.
And it's designed for a modern world with billions of users who might not all be completely trustworthy, so whole categories of complexity go away -- such as NATs.
So there's no standard industry term for describing this system, because there are no direct analogs or competitors.
Why would you want to use Urbit? The same reasons you want to use the modern consumer internet, which Urbit intends to pave over and replace with something better.
How is Urbit better? Because you'll have all your data and programs on one machine, which you control, using an open source operating system. You can build a peer-to-peer twitter or facebook clone on Urbit in a day or two, because the OS handles more of the distributed systems and identity problems that have killed most peer-to-peer projects in the past.
It's still in alpha, so it's a bit slow and buggy, and a lot more work has been put into kernelspace than userspace so far. What will make you way "wow this is cool" varies widely, but some good candidates are: - an internet experience not predicated on surveillance capitalism - no ads - control over your UI - a minimal aesthetic - interesting people to talk to on the network - lack of the twitter "thunderdome" feel - if you like to write programs, the Urbit system is fascinating to work with; I've learned a lot more CS from working on it
The other thing that's cool is that this new world isn't yet fully settled. You can write a little talkbot or something and still have a big effect on the culture.
Spoken programming languages FTW.
The pronunciations are no more exclusionary than learning Cyrillic, and the effort pays off.
The OS itself runs as a VM on your machine (either a Linux or MacOS host, for now), and that is the only place your data lives.
Urbit does implement a peer-to-peer encrypted, authenticated network among these VMs. It's not solely socially focused, though.
It's also intended to be your personal archive for things like your personal financial data, pictures, nusic, notes, and private documents like tax records and medical histories.
And not only an archive, but also your personal "agent", in the sense that it's a program that's on all the time, on the network on your behalf, to serve your blog, for example.
You can access Urbit through the command-line or a webpage that it serves for you.
The state of an Urbit VM is a folder on the host OS. You can zip it up, move it to another computer, and restart it there seamlessly.
None of these features by themselves is particularly interesting. What's unique about Urbit is that all of this is accomplished using a very small set of primitives, making it easier to write applications that don't go through some huge company that tracks your every move and has a conflict of interest between serving you and serving advertisers.
If you're interested in learning more details of how it works, there is actually quite a bit of technical documentation at https://urbit.org/docs
We're obviously still trying to figure out how to describe this thing. It doesn't occupy a slot that people already have in their minds for a piece of software, except maybe "personal server", but even that is somewhat vague.
Because there's a different world of computing inside Urbit, with its own libraries, apps, languages, etc., it can take a while to wrap your head around.
What's funny about this is that the Urbit world is orders of magnitude simpler than a standard Unix-based stack. We've just spent so much time learning things like DNS A records vs. CNAMEs and what 'xzvf' do for tar, that a parallel universe of these constructs seems bewildering again.
Got some promising initial results for my app but it's time to try some science...