78 karma · joined February 29, 2020
Writing Rust, Zig, or Go, you would still control memory manually, where it matters.
It would take at least some knowledge to hack, not just a random script from a forum.
Now, with LLMs, it's the '90s all over again.
Search Engine have chicken and egg problem, you can only have a good search if enough people use it to tune the ranking and see enough search spam cases.
“Just compete with a monopoly on their own field with one hand tied behind”
To be fair, I haven't try exactly "rewrite in Rust" prompt, but I cannot imagine for this to work.
Although, in the agentic environment, a benefit of not having GC at all definitely helps.
What is the scale? Because I'm pretty sure it's impossible to one shot 65k LoC with "good luck, make no mistakes" prompt.
I also did this within a subscription. I counted the number of tokens afterwards and calculated the cost as if I'm paying per token. $400 is of course arbitrary, but it's a ballpark number, bun was $165000.
> In my experience it mostly comes down to the harness (or lack of) that you use.
Yep, it is.
Tokens per task is a good proxy measure of skills, 'superpowers' or any other.
Either a skill gets you the thing more efficiently (less tokens), or you don't need to redo the result afterwards (less tokens). I benchmark all my skills that way.
Models need less and less steering at this point, especially frontier ones.
The way rune works with files minimizes chances of silent corruption. I keep original byte blobs immutable, separately there is a journal (kinda WAL) of positional deltas (inserts and deletes).
So I only need to validate that blob + deltas = snapshot.
Disk IO is encapsulated through VFS, and writes are atomics (write to a temp file, then rename).
Separate virtual rendering buffer is built on top of that. Rune, like Obsidian, renders markdown preview inline, so the same chunk could be rendered as "Header" as well as `## Header` when under the cursor.
All of that makes it quite easy to work with text. The core function is to translate offset in a byte array to line and column and back, which is pure math and relatively easy to test.
Another trick that helped a lot is to use sqlite extensively: blobs, deltas, vfs, redo and undo history graph are all sqlite tables.
I noticed that LLMs make stupid decisions when it comes to data structures, but they understand CRUD and SQL, so I turned all Rune's internals into dumb CRUD.
But somehow compiler has decided that i <= limit is always true.
65k LoC of Go without comments resulted in roughly 60k LoC of Rust witout comments (code column of the cloc tool).
The error handling is not so different between Rust and Go, in both cases I cannot panic to avoid the data loss. So it boils down to if (failure) return something for graceful degradation. And generally errors in my case (a text editor) are rare, only disk IO, which is encapsulated in one VFS module, everything else, like non-closed brackets in code is expected behavior.
The biggest differences were in third party libraries, UI, markdown parsing — completely different API and paradigms.
> Performance characteristics of resulting rust
I haven't measured. I don't think there's any significant difference between Go and Rust if app doesn't do allocations on a critical path. The reason I started this project was mainly to experiment (now I use similar approach to refactor much bigger legacy code base), and tree-sitter support is better Rust so it seemed like a good fit.
I did because I have a huge legacy code base, a distributed monolith, a few millions lines of code. Ideally, I want to get rid of it.
At this point, I know a recipe to break the monolith, so I finally could eat the elephant piece by piece.
My point is that one can translate the data flow to another language.
You could imagine any program as input -> [blackbox] -> output. For example, same pixels rendered on the screen provided identical keyboard input.
I propose a way to decompose the blackbox.
In my experience, you can get good results if Fable does't write code itself, only spawn subagents.
I can run Fable for 10 hours, and it would output 50k tokens and read 300k (30% of the context window). The resulting code is okay-ish. I would rarely merge LLM-produced code first try without an adversary review.
There is a way to control tool invocations at the harness level when writing skills or agents.
Example: ~/.claude/agents/critic.md
---
name: critic
description: >
Plan Critic. Reviews an implementation plan. Use before the implementation.
tools: Read, Agent
---
For Claude skills, the frontmatter is different, the keys are: allowed-tools: Read Grep
disallowed-tools: WebSearch Glob
https://code.claude.com/docs/en/tools-referenceYes, I completely forgot to mention, this is exactly my case. Rune is a TUI editor, so I feeded the same terminal sequences to the old and new apps.
It didn't translate 1:1 (I ported core editor first, there were side panels, and different chrome elements) so I instructed LLM to use ttyd (tty -> browser render), Fable then could open both apps with playwright, make and compare screenshots.
To rephrase, one critical component is to establish a feedback loop for the model. This new generation of models: Opus 5, Fable, GLM-5.2, even Qwen3.8-27B can self-correct, provided they know whether they are progressing or not.
A month ago, especially smaller model would fall into a rabbit hole it dug for itself and would never recover. This generation can sometimes run tens of hours without losing track.
I still wouldn't trust a model after 70% context window, but the progress is noticeable.
My all-time favourite example is (again, my memory, I may be a bit wrong):
for (int i = 0; i < (size_t)limit; ++i) {
}
at some point limit could potentially become greater than INT_MAX, the compiler decided that i<limit could never be true because that would cause signed int overflow which is UB, so it "optimized" the loop into while(true)
signed unsigned mismatch makes me shiverIf not an intermediate representation, my first question would be: how to split the work. You cannot just prompt full rewrite of 60K LoC, you cannot do it module by module — modules do translate 1:1, like in my case golang packages did not matched Rust crates. Bun did file by file.
With hierarchical state machines in the middle, I did it almost in one shot.
A lot of infra is vibe coded nowadays too.
Even prototypes are contributing to the speed of software development. Many people vibe code throwaway dashboards around the main platform which gives a lot of insights.
I absorbed so many different models at this point :)
Thank you for the kind words.
My friend once sent me a snippet, maybe 10 lines of C+++, asking "can you spot the UB?".
So I'm staring at these 10 lines, I KNOW there is an UB. I wasn't able to find it without a hint.
I'm sure there are other caveats. Also cost of the more or less straightforward Bun port was $165000 for 500KLoC.
I haven't played with GLM 5.3, only with 5.2, so I cannot say for sure.
GLM is at the level of Opus. Fable is something different entirely. It is capable of tracing the data flows of the app, I even tried it in a huge PHP codebase, it works.
PHP is a weird beast because it allows something like
$v = 'SomeClass' + 'Controller'
... // and later
new($v)
In other words, it could be hard to understand the code without running it, Fable reads this.And another distinct feature of Fable — it is an amazing orchestrator. I prohibit it basically read and write, and it operates a swarm of haiku and sonnet.
Probably I need some fresh air and a good fiction book.
It is probably because I read tons and tons of LLM output.
The whole point of this experiment was to try and make the rewrite as cheap as possible.