>> I was curious why you swapped over to a register based VM
Not the OP but there are real performance benefits, I've been poking at a wasm VM and it has two jit backends where one is pure copy-and-patch while the other caches the locals in registers using the function args + copy-and-patch and there is a significant performance gain just from that alone. A push/pop from a stack is fairly expensive while the register caching keeps things in the CPU's happy place. The smallest gain was ~2x over the interpreter on memory bound tasks while the largest was ~20x on math heave kernels. Admittedly, the interpreter isn't the fastest thing ever as its one and only goal is conformance with the spec to use for differential testing but the difference between the the two jit levels are somewhere in the neighborhood of 1.5-5x depending what the code is up to.
The three biggest performance gains, from the random benchmarks, are quality of the bytecode out of the compiler, the jit itself and register caching from what I can tell from the fancy chart I had Claude make and a good squint. Tail-calling would be somewhere on that list too but I can't measure that as all the opcode do the tail-calls between each other as that's just how it was all put together, the code the interpreter runs is the same code the copy-and-patch jit stitches together as they are both generated from the same DSL. Which is also the biggest cost with the register caching as the code template file grew from tens of kilobytes for the 407(?) wasm opcodes to ~3MB for all the specialized ones to pass the locals in eight args but that's really just a binary size thing, the stitched together functions just pick and chose the ones they need.
Long winded way to say CPUs like when you keep things in registers, I suppose...