Statically Recompiling NES Games into Native Executables with LLVM and Go (2013)
andrewkelley.me
andrewkelley.me
It's amusing seeing this in a machine which gets its code from a ROM.
http://webassembly.org/docs/future-features/#platform-indepe...
What about java/.net/regex engines? Is it required that they exist outside the browser, or are performance limited to interpreting, or is there a better solution?
What's the problem with the control flow of assembly?
These tricks are IMO an indicator that they had a good idea that the game code was going to be a tight fit (or had possibly already discovered that the hard way), so they were optimizing for size. A jump would be faster than executing the BIT instruction.
The manually loading an address to the stack and running RTS is an interesting trick. Wozniak used it in his SWEET-16 interpreter, pushing the interpreter loop entry point to the stack before jumping to the subroutines that execute the byte code, just so that he could use RTS to return to the interpreter loop. In that sense it fulfills an entirely different purpose than a JMP indirect and also saves some bytes, again at the expense of execution speed and non-gray hair. For a table based interpreter, it fills gap left by the lack of a JSR indirect instruction.
Some games, because of bugs, allow exactly that. User input might be directly executable. That means that static compilation cannot be achievable in general. Either you have different behaviour in the event of such bugs or you have a full interpreter to fall back on.
[1] https://www.polygon.com/2014/1/14/5309662/bizarre-super-mari...
I wonder whether something like this might be better attempted in rpython which will build a JIT compiler for you.
I hope all of these NES ROMs were coded well. Seems like you could do something like...
if False:
jump_to_data_address
... then he'd be interpreting data as code.Not true at all. There's nothing in reflection and serialization[1] that is hard to implement with an AOT compiler.
The real problem with AOT-compilation in Java is that there are many optimisations, like devirtualization, that can only be done (or done much more easily/effectively) at runtime.
You can get some of this back with whole-program/closed-world optimisation, but in reality that's highly impractical for Java. Many Java programs are very large and people want the ability to update them quickly and easily, often without even restarting let alone recompiling.
An LLVM-based Java runtime could use a hybrid AOT/JIT model, however. Recompiling parts of the program as needed at runtime based on profiler data, you'd get the fast startup of AOT combined with the high performance of JIT.
Don't get any crazy ideas about beating Hotspot's performance, though.
[1] one exception is classes/methods that are defined at runtime using dynamically generated bytecode. But that's pretty rare.
Oracle Labs is also pushing forward their research of meta-circular JVM, which makes use of AOT compilation. Although it restricts it to a "native Java" subset.
http://cr.openjdk.java.net/~jrose/metropolis/Metropolis-Prop...
Have you ever profiled Excelsior JET, Aonix, Jamaica, IBM Websphere Real Time for example?
Their products are almost as old as Java, so I imagine they are worth their price.
Regarding OpenJDK, according to the Java Languages Summit talk, the Java 10 dev branch already has better optimizations than version 9.
https://github.com/graalvm/sulong
Which is much better, because we need more language research with type safe languages, not with C and C++.
Emulation is just plain fun.
http://lists.llvm.org/pipermail/llvm-dev/2008-July/015629.ht...
However, thanks for sharing the link to the previous discussion!
Maybe "Guidelines" and "FAQ" should simply be merged?
Edit: Opened an "Ask HN" submission about this: https://news.ycombinator.com/item?id=15255270