Like, what is the vision here, a LuaJIT that targets RISC-V, running inside my x86-64 game, where the emulator translates the RISC-V back into x86-64... for reasons? Just isolation?
Like, what is the vision here, a LuaJIT that targets RISC-V, running inside my x86-64 game, where the emulator translates the RISC-V back into x86-64... for reasons? Just isolation?
Yes. When something's intellectually stimulating you work on it more, that's it.
Most game development is a grind. These huge distractions, like a scripting backend, well if you spend 100h working on the scripting engine only to spend 1h authoring actual scripts, you still spent 1h authoring scripts those 2 weeks instead of 0h. The game gets delivered sooner even if you spend 10x as long working on it.
It's a quintessential misunderstanding about indie game development. HN readers think Jonathan Blow is wasting his time writing a whole new programming language and engine, and that's why his games take 6 years to make. No: his games would take 20 years to make if they weren't intellectually engaging to make. They wouldn't be made at all!
It's the same energy as rewriting everything in another programming language.
I think this happens at giant companies too, all the time. So called Not Invented Here syndrome: it's as much about laundering open source code as it is about keeping things interesting enough to make the extreme boredom and grind worth it for otherwise smart and healthy people.
Any of those would be better supported and have a faster, battle tested JIT engine.
2. Regardless, this sub thread isn’t about TFA.
It is indeed true that most VMs, including JavaScript, WebAssembly, the JVM, etc., have significant latency on the boundary. For example, many wasm VMs will install a signal handler, and that needs to be enabled/disabled on each entry/exit from the VM.
But the latency can be fixed. You can sandbox WebAssembly without a signal handler, in particular (at the cost of a few % throughput overhead for bounds checks). I'm not sure if the author benchmarked that, but it should be very fast.
(As for "why RISC-V" in the project, it looks like that's because it's easy to write an interpreter for, but it could have been any compiler target, it seems.)
I agree that it could as well have been MIPS or ARM.
It is the youngest of the lot, meaning it is carrying around the least legacy compatibility stuff while still incorporating a lot of modern design lessons.
Additionally, modern ARM and MIPS are not free and (potentially?) patent encumbered as well.
You could go with some old, free MIPS or ARM, but then how good is performance, and how good is compiler support for these nowadays?
Wasm was designed to have very fast calls, both internally and externally. That's why the basic types correspond directly to machine types. The goal is to have 0 overhead, except for sandboxing.
Regarding sandboxing, there is the stack overflow check that happens in Wasm on each call. But you can disable that check if you don't want it.
Btw, you can see that Wasm has no extra overhead on calls in a simple way: compile Wasm using wasm2c and inspect the C output. That C is what a Wasm VM would run. Some details:
https://kripken.github.io/blog/wasm/2020/07/27/wasmboxc.html
Note the section there about LTO: The Wasm can even be inlined into the code calling into it, which means there is zero call overhead. (But even without LTO, it's just another C function to call.)
It's cool that you can inline sandboxed functions like that. I'm curious about how you arbitrarily select and call functions though? A string-lookup kind of thing?
RISC-V seems to be used here more as a bytecode format of sorts, and at least compared to x86_64 it should be much easier to implement (first of all because all of the opcodes have the same size.)
But if you are doing a lot of communication between the script and the host system, I'm not sure virtualisation does so well?
(Although... I do see value in having a second option. I guess it might be worth seeing if it can actually be better, even though it's not the thing I would have written)
Doesn't he explicitly address WASM as an unfit backend in the article ?
> Lua, Luau and even LuaJIT have fairly substantial overheads when making function calls into the script, especially when many arguments are involved. The same is true for WebAssembly emulators that I have measured, eg. wasmtime.
If the scope is small, the answer is easy: write some naive code, incrementally profile. The project isn't big so full rebuild times can stay light, and your team isn't large so you really can "do whatever" and get somewhere, especially if you take the route of writing in a native language that compiles fast - a Pascal or something more "new and hip" like Beef.
But when you are asked to do it on a AAA project you end up with every imaginable kind of feature: it needs to support a team of hundreds, it needs to be fast on console hardware, it needs to be straightforward to debug, it needs to support modding, it needs to be fast to iterate on, it needs to be flexible about the memory layouts of save data or assets.
So you do end up in this kind of space where you're like, "we'll target a VM that is really low level, and that'll reduce the friction and granularity of switching between debug and release profiles, and that lets us stay in control of when we want to sandbox and when we want to go fast and we'll still be able to control every byte". That it happens to be RISC-V is not super relevant - it could be WASM or a custom bytecode like the Hashlink target in Haxe. It just needs to be a thing to compile to, that you can reasonably expect to implement and maintain.
It's not the only approach that could be taken. Downscoping the ambition by 50% and hardcoding a little more of your spec is a good way to get through the technical stuff 10x faster.
It could definitely still emulate itself, but it'd be far higher performing to run the code directly in some sort of sandbox.