This would still be pretty slow (see: Microsoft’s version of this under Windows ARM) due to the need to issue a ton of memory fence instructions to make ARM’s looser memory model behave like Intel’s, except that Apple baked the ability to switch the CPU into an Intel-like memory model directly into the silicon.
So in practice it is shockingly fast.
No it's not. Apple is not immune from fundamental computer science principles whatever their marketing team says, and even the original keynote acknowledged that Rosetta 2 emulates some instructions at runtime.
Imagine you're running Python under Rosetta. The original Python interpeter takes Python code, translates it into x86 assembly, and runs that x86 assembly. Those x86 instructions did not exist prior to execution! Even if Rosetta could translate the entire interpreter into ARM code, the interpreter would still be producing x86 assembly.
Other types of programs produce code at runtime as well. Rosetta 2 is able to cache a very impressive amount of instructions ahead of time, but it's still doing emulation.
You're still right that that's not sufficient, for example anything that generates Intel code will definitely need JIT (just in time) translation. But presumably a lot of code will still hit the happy AOP path.
That being said, a JIT does not have to be super slow. The early vwmware products, back before there was virtualization support in Intel products, actually had to do some translation as well: https://www.vmware.com/pdf/asplos235_adams.pdf
I mean, we can call it an ARM binary or we can call it an instruction cache. I generally prefer the latter term, because what Rosetta produces are not standalone executables, they're incomplete. I don't know how often the happy path is used, but Rosetta can always be observed doing work at runtime.
JITs are great and Rosetta 2 is incredible! I just can't imagine it working over any sort of shared filesystem, that would add an incredible amount of latency.
So yeah, no real back and forth to the host platform.
Not dynamically. They just call predefined C (or whatever the interpreter was written in) functions based on some internal mechanism.
> Or else what does the CPU execute?
Usually either the interpreter is just walking the AST and calling C functions based on the parse tree’s node type (this is very slow), or it will convert the AST into an opcode stream (not x86-64 opcodes, just internal names for integers, like OP_ADD = 0, OP_SUB = 1, etc) when parsing the file, and then the interpreter’s “core” will look something like a gigantic switch state statement with case OP_ADD: add(lhs, rhs) type cases. “add” in this case being a C function that implements the add semantics in this language. (The latter approach, where the input file is converted to some intermediate form for more efficient execution after the parse tree is derived, is more properly termed a virtual machine and “interpreter” generally only refers to the AST approach. People tend to use “interpreter” pretty broadly in informal conversations, but Python is strictly speaking a VM, not an interpreter)
In either case, the only thing emitting x86-64 is the compiler that built the interpreter’s binary.
> Am I totally misunderstanding how interpreters work?
You’re confusing them with JITs.
If every interpreter had to roll their own dynamic binary generation, they’d be a hell of a lot less portable (like JITs).