500B Tokens Later: Letting AI Agents Decompile a First-Person Shooter
momo5502.com
momo5502.com
I've found this to be robust for decompiling games, while giving the agents enough freedom to write code that is readable and not waste a ton of time making sure e.g. instruction ordering, register assignments, etc. are all exactly the same. For me, having byte-matching decompilation is only one way to produce a decompilation I know is faithful to the original. This "high-level decompilation" process I just described is something agents can do much more quickly.
(For example, your approach would not necessarily catch all the same overflow behaviors; the OP expressly claimed that "replicating all bugs" was also important, and many bugs are caused by certain overflow behaviors)
That said, I would probably follow this same approach if I were to do this, but with extensive randomized testing as well.
So you're looking not just for functional but also dysfunctional equivalence %)
Also, it feels like just fuzzing the function is slow, compared to simply checking bytes. Fast feedback allows agents to iterate much faster. At some point I wrote a tool that simply proves semantic equivalence, while ignoring irrelevant changes, e.g. register selection or stack slot changes. Turns out it was way too slow as a harness. I still used it to verify some of the remaining functions.
It turns out that getting functions byte-matching is actually not that complicated. Most functions are small and can be one shotted, and over time, agents figured out rules how to control code generation for more complex functions. They wrote down techniques and tricks on how to control register selection for example, or how certain control flow operations may impact basic block ordering, etc. Super insteresting what they figured out. So the further we got into the project, the easier it was to get functions exact.
The agents would have had to mess around with compiler versions, optimisation options, and the phase of the moon as well.
If you just went for functional equivalence, it would probably cost 10x or 100x less tokens.
Another false economy was using Sonnet instead of a more intelligent model like Sol 6.1 (1), which would have cost more per token, but is 100x or so better at reverse engineering and coding and therefore can chew through the source code much quicker and make fewer mistakes, meaning less work needing to be scrapped.
In my testing doing a similar task, I ran multiple sonnet for weeks and burnt through ~$1000 in tokens to get 20% completion and output that was pretty bad. After switching to Sol 6.1, it finished the whole task in around 2 days, cost around $50, and it did it with zero supervision and a single /goal.
(1): struggle to use Opus for reverse engineering, too many safeguards. OAI has virtually none, and uses way less tokens so is more economical.
No one is going to use it if it’s a kind of close but not really reimplementation.
Testing functional equivalence is also pretty much impossible. How would you for example test the new one has exactly the same bugs which haven’t been discovered yet. Or doesn’t introduce new ones? This stuff matters for speed runners.
Yet the C code can’t pick what registers to use, so the poor agent is probably shuffling the code around randomly for hours or days until it matches.
That’s probably why the agent dropped down into inline assembly in the first place (the author complained about this), because I bet it’s thinking trace was that this is futile.
Compilers themselves are not even deterministic and running them multiple times creates different assembly.
Humans too do that for matching decompilation. Or at least I imagine so, given that I personally refuse to do that.
People have different goals and will use different techniques to achieve them. The video game reverse-engineering/decompilation community isn't a hive-mind, I went ahead and created ghidra-delinker-extension because I had my own ideas on how to do that.
Due to this strict harness, we were able to use extremely dumb models efficiently. I think at least 200-300B of tokens (maybe even more) were processed by Luna.
Given that Sol is 100x more expensive per token than Luna, a high token consumption doesn't necessarily mean high cost. So these 300B Luna tokens would equate to 3B Sol tokens.
Unfortunately a huge part of the logs is lost, so I can't do a detailed analysis on which models were used, how much they cost, etc.
Call of Duty: Modern Warfare 2 (2009)
https://web.archive.org/web/20260925153118/https://momo5502....
https://web.archive.org/web/20260925153131/https://momo5502....
Come at me, corporate America.
One of the few problem is that the decompiler is not always reliable. That forces you to go read the assembly, and it is not apparent to decipher the right kind of feng shui, not even from human before the LLM era. I used to play CTFs and my conclusion is exactly that.
This is extremely apparent when there are self-modifying code (e.g. JIT) is involved. You need a stepping debugger to read the right control flow, because the code will diverge based on the instruction pointer and regions you jumped into. At this point static analysis like IDA and Ghidra stopped working.
Total Opus 5.5 token usage on OpenRouter last week is 5000B, presumably not counting cached input.
Edit: Here's a possibility.
The future of software is agent only software. We ask for a result, agent asks questions from us, agent uses the specialized software, user gets result. We subscribe to an AI assistant and specialized agents. We are almost there,at least the start. The future of PC's as we know them are numbered. OSs,CLI,compilers and whatever will melt into AI assistants. Say goodbye to writing software for people.
You could just have the LLM agent use the software itself and then duplicate the functionality it finds, don't even need the screenshots or video.
Unless the software is doing something extremely novel, the LLM can just do this all itself after you point it toward a url for the software with instructions to clone what it finds there.
It'll be interesting to see if SaaS companies and/or providers of internet services like Cloudflare try to stop this sort of agentic cloning of SaaS products.
Like, obviously they would kinda want to if/when this sort of thing becomes commonplace, but how do you do it while allowing your software to be agent-friendly for non-cloning uses, and who is going to want to tell their customers they can't use agents with their SaaS for fear of their software being cloned? Kind of a bad situation with either option.
Someone has vibe coded the entire Adobe suite btw: https://getartcraft.com/apps
> There are no noticeable bugs and all features of the original game are present.
>
> The remaining functions have been reworked repeatedly. While they still don’t match byte for byte, we believe their semantics are correct.
As there are no way to confirm this outside of ”believinh”, what you are left is with a end product that might contain thousands of micro changes that essentially make the game something else than it was intended to be.
Since even Fable 5.1 cannot reliably remember what it wrote into "memory" files before the last compaction, why would we trust it to not forget about some side-effect when handling assembly that is so verbose it'll surely exceed the context and, hence, necessitate compaction.
https://github.com/ValveSoftware/halflife/blob/master/pm_sha...
The LLMs have presumably already consumed information like this as part of their training sets. Converting between Hammer and Unity scale is a fairly trivial linear operation. You can dump the BSPs to recover geometry and rapidly accelerate map development using timings that are already known to work.
Attacking the raw binary directly is certainly impressive, but it's totally unnecessary.
So, if you reverse-engineer game X and post reverse-engineered code, what exactly do you infringe, how and in which jurisdiction? What changes if it is done via LLM?
(I understand that LLM decompilation is absolutely out of hand right now and something surely will come to trample the fun. But what and when? I suppose american LLMs will have their system prompt updated to forbid any reversing help and report suspicious activity straight to legal hotline)
In 1986 there was a federal court case, Whelan Associates Inc. v. Jaslow Dental Laboratory, in which it was ruled that the "structure, sequence, and organization" of a computer program was protected by copyright, and thus independently produced software could be found infringing if it copied these elements, even if the code were not copied (or mechanically translated) verbatim. This led to a six-year period in which computer software enjoyed generous copyright protection, such that "clones" of copyrighted software were effectively infringing. It wouldn't be until the early nineties that other court rulings would tighten the rules again, notably Computer Associates International, Inc. v. Altai Inc.. The 3-step "abstraction-filtration-comparison" test has been used by most courts since 1992 to determine whether nonliteral parts of program code are eligible for copyright protection, and whether another, independently written program is infringing.
HOWEVER, the Whelan standard was never actually overturned or stricken from U.S. law due to legislation or litigation! And companies have sued and won under the Whelan standard! Most notably, Oracle in their copyright and patent case against Google regarding Java APIs in Android. Google ultimately prevailed, but only because the Supreme Court ruled that Google's use of the APIs was fair use; they did not overturn the Federal Circuit's finding that the "structure, sequence, and organization" of the Java declaration code was ineligible for copyright! And I doubt that the Supreme Court would similarly smile on a reverse-engineered complete video game!
I believe that these reverse-engineered projects infringe copyright, under the Whelan standard and perhaps under the stricter Altai standard as well. You do not get a free pass because the original code was in C and yours is in Rust; or because you created a slightly different version of each function in the original code.
And let's say there's a future where an AI model could zero-shot the game itself. What then?
How can one prove that X is a copy when none of the source code match original? Clean room re-implementations are not illegal after all, if we go that way (yeah, yeah, I know about clean room argument).
While I personally do not think it would be different, but a lot of LLM folks are, and even some companies (rushing to blatantly clone and de-compile stuff). Kinda strange feeling sitting and looking around in the midst of it, y'know.
PIF owns EA, when they invite you to a meeting you might start to worry about foil lined room and carpentry tools lying around
- please someone do it
https://www.reddit.com/r/ReverseEngineering/comments/1vxig19...
Even after building whole infra for that and not pursuing identical byte output it's been pretty token hungry process
Normal development isn't so lucky, you do care about code quality. So you have to review the code the agents write and handle the testing yourself. The agents will not do what you want, they'll do what they think you want, and you're the only person who knows what you think.
Obviously matching functions byte-byte is not it, but I think with enough creativity, you can have guardrails and oracles almost everywhere.
E.g. before diving into feature development, you can have agents generate a test suite based on your implementation first. Then, when developing features, this test suite shows what broke and agents can double check if that was intended.
How that test suite looks like obviously highly depends on the environment.
>At current 2026 API rates, 500 billion AI tokens would cost roughly $100,000–$750,000 depending on the model, with most flagship models in the $150–$400 per million input tokens range
Yet, it is now profitable.
The bet is that people realize how valuable these services are, and despite complaining, they still would pay the higher price. This realization would not happen without this initial subsidy from investors.
It isn't too different from drug dealer's first sample free...
Taxis were an established profitable business model and the uber subsidisation wasn’t anywhere near as much as AI subsidies.
????
Uber and Lyft were paying drivers $X but charging users way less than $X so they were literally burning investor money to get market share.
Why does this point come up over and over again, pretending that you can truly separate training and inference costs. Yes, they are separate things but the value OpenAI and Anthropic are presenting to the world is "we're the absolute best, no one else comes close." Well, to keep that up you can't just not train for extended periods of time. You have to keep that engine going non-stop when there's free Chinese models nipping at your heels. You can be profitable on "just inference" all you want but if training expenses dwarf that, you're not going to be profitable overall, and that's the bottom line.
So my conclusion is that it only works because the majority of users doen't fully utilize their allowance. Last month I used almost nothing of my Claude 20x plan (did use Codex though).
You think developers won't/don't stretch subscriptions to the absolute limit? The AI subscription model is like the gym model except a large percentage of the customers work out 24/7/365.
Also, Uber was peak ZIRP + COVID, money was cheap and growth was easy. I think the mountain to climb is a lot steeper now.
Reading between the lines:
"The avid reader of my blog might have noticed that I had previously written two posts that have since been removed. Everyone else might now be wondering which game I am talking about. To both of you I can only say that corporate America was here to ruin our fun."
Keeping it private likely wasn't the original plan...
https://developers.openai.com/api/docs/pricing https://platform.claude.com/docs/en/about-claude/pricing
OpenAI and Anthropic are both 10/mil in.
https://openrouter.ai/z-ai/glm-5.3#providers https://openrouter.ai/moonshotai/kimi-k3#providers Other frontier models are like 1-3/mil in.
Also, 500 billion is 500,000 millions. At the lower end of your 140/mil estimate that's 70 million dollars. Even at my 2/mil lookup for Chinese frontier models that's one million. Show your math for 100k-750k, please.