If a CPU ran on its own, then there's no end to our optimization potential. If it were a strict Harvard architecture where instructions could only run in immutable (ROM) areas, we could translate the instructions, remove all unnecessary flag calculations, and then statically recompile the result to a native program.
But when the CPU needs to talk to the audio chip, the video chip, the other audio CPU running alongside it, etc ... the reason everything ends up so slow is that there's no viable path for speeding up that synchronization.
You try and do it with real hardware resources: a real CPU core/thread for each emulated chip, and performance falls off a cliff. Even with simple atomic reads/writes for one thread/chip to set a "waiting on you" flag, and another thread/chip to clear it, modern CPUs can only do this around 100,000 times a second ... and that's before the overhead of your emulation. So, if you want to emulate two CPUs that run at more than 0.1 MIPS (which even the NES surpasses), you're out of luck.
So you try and do it with a single thread, but you get destroyed through context switches. You're in the middle of executing a 68K instruction, but that instruction writes to RAM that the Z80 can read from. You don't know if the Z80 is going to read there, so you have to run the Z80 until it's "ahead" in time to the 68K. Switching into the Z80 interpreter absolutely murders your performance.
Pretty much the primary key to optimizing CPU emulators is to synchronize less often. Making assumptions like, "it's very unlikely the Z80 is going to write to the CPU instruction stream in the middle of it executing instructions ... the 68K is most likely executing code in ROM anyway" and not synchronizing the Z80 when the 68K fetches an opcode or operand byte.
But these optimizations always come at a cost. When you get a library of thousands of programs designed to eke out every last drop of performance from old, legacy 2MHz hardware, there's always that one program ... either by design or by fluke, does something crazy and relies on something you optimized away as extremely improbable, and breaks as a result.
In my view, the way to optimize real-world emulation of machines is that we need to be able to spawn lots of 1 core = 1 thread objects, and have their synchronizations be as cheap as humanly possible. The cores do not have to be lightning fast, they just have to be able to synchronize as quickly and as cheaply as real code running on a 90s era 68K+Z80 machine (eg a Sega Genesis) is able to. Many, many bonus points if the "waiting on another chip" operation can result in sapping less power for that thread, without increasing the latency necessary for it to resume operating once the condition is met.
The industry has, for decades, been moving in the complete opposite direction, so I don't have a lot of hope. At this point, I think it's more viable (but still very unlikely) to put FPGAs into home computers for this purpose.
Architecturally, I think some of the things you want exist in some unusual microcontroller families (XMOS, GreenArrays, Cortex-R), but then if you get to pick the microcontroller I guess you might as well also pick the clock so that you can run in real time instead of simulated time.
In fact, I took one of the techniques (traces) directly from a paper describing how the Bochs x86 virtual machine works :-)
But it goes deeper than that. I believe the whole "trace" thing in both jits and vms comes from a few papers describing trace-based instruction predecoding for hardware CPUs.
https://en.wikipedia.org/wiki/Josh_Fisher#Trace_Scheduling
He combined this with a VLIW processor architecture to build Multiflow, a hardware startup. (Interesting history tidbit: Robert Colwell, who architected the P6, the first out-of-order Intel core, started his career at Multiflow before joining Intel. The P6 didn't have any trace-cache influences, but the P4, a few years later, infamously did...)
Similarly might you have links for those papers that describe "trace-based instruction predecoding for hardware CPUs."?
Cheers
The big problem was that Apple was lousy at specifying how much cache they needed. IBM told Apple that it needed a LOT more cache on the 603 unless everything was running native. Apple ignored the advice.
The result was that the 603 and 604 were performance dogs because so much was still running in emulation.
So, IBM went back and bumped the cache for the 603e, and performance went up dramatically. This then led to the unfortunate situation wherein the "low-power" chip from IBM (the 603e) tended to be far more performant than the "high-performance" chip (the 604) from Motorola.
The microinstruction is 23 bits wide (plus one for parity). Rather than doing a lot of masking, shifting, and testing to extract fields (some fields are non-contiguous even), I instead use 64b per microword.
bits [23:0] -- original microinstruction bits [31:24] -- 8b op-dependent predecode value bits [39:32] -- 8b interpreter dispatch index bits [47:40] -- 8b op-dependent predecode value bits [63:48] -- 16b op-dependent predecode value
There are a few oddball instructions which could use more than the 8b and 16b predecode fields, and instead do the old shift and mask on the raw microword to figure out what is needed.
Overall, it nearly doubled my performance.
I'm curious why does the microinstruction need a parity bit? What happens if the parity is wrong, a machine check exception?
>"bits [23:0] -- original microinstruction bits [31:24] -- 8b op-dependent predecode value bits [39:32]"
What do the pre-decode bits do exactly?
Which microcoded machine was this?
There are a number of microinstruction formats, and fields aren't always contiguous. Making up an example, say the instruction is "ADD R1,#imm8" to add an 8b immediate to the R1 register. But the 8b immediate value is stored in bits [14:10] and [4:2] of the microword. The straight-forward way would be to write "uint8_t imm = ((instr >> 7) & 0xF8) | ((instr >> 2) & 0x07;" Instead, when the writable control store is written, that quantity is decoded and stored in an 8b aligned predecode field, so getting the value is just "uint8_t imm = instr_struct.imm8;"
The machine is the Wang 2200. There were two architectures: the first used a 20b word in ROM, the second used a writable control store so the BASIC could have bug and feature updates by mailing out floppies.
Think of the popular 1980s computers: IBM PC (Intel 8086), Amiga (Motorola 68000), Commodore 64 (MOS Technology 6510), TRS-80 (Zilog Z80), Apple II (WDC 65C02), Acorn Electron (Synertek SY6502A).
The solution at that time was to use p-code ("portable" code), such as https://en.wikipedia.org/wiki/UCSD_Pascal
In fact, Pascal's popularity in the 1980s was probably due to the large number of p-code interpreters and Pascal -> p-code compilers.
--------------
Ironically, we still use "p-code", now called Bytecode, today. But we really don't move between systems aside from ARM and x86. GPU assembly is special, since its an entirely different model so you can't really port Java or Python to the GPU.
I guess LLVM shows that the high-level translation to LLVM-intermediate language just simplifies compiler optimization, to the point that its useful even if you're making code for a specific machine.
EDIT: I think the modern CPU have more or less settled on the same features. They're all Multithreaded, cache-coherent Modified 64-bit Harvard machines with out-of-order execution, super-scalar front-end with speculative branch prediction, with ~6 uop dispatch per clock tick and roughly 2 or 3 load/store units with 64kB L1 cache and 64-byte cache lines with a dedicated 128-bit SIMD unit
The above lines describes ARM Cortex-A72, Intel Skylake, AMD Zen, and POWER9... except Skylake has 256-bit SIMD units I guess, and Apple's A12 has 96kB L1 cache. Not very big differences anymore between CPUs.
Was there a JIT compiler for p-code back in the day? I thought all implementations were interpreters or AOT compilers.
See the paper: "Efficient Implementation of the Smalltalk-80 System", which suggests that code was generated on the fly.
http://pascal.hansotten.com/niklaus-wirth/recollections-abou...
There is then some marketing hype around JIT from late 90's that is mostly only relevant for Java, which implies that "true JIT" does things like hot spot detection, deoptimalization traps and trace-based program flow reconstruction. This is mostly only done in production only by JVM implementations and fallback interpreters in hypervisors and is based on 80's research projects.
Somewhat notably in early 90's there was HP Dynamo, which was userspace PA-RISC emulator that ran on PA-RISC host by means of trace based JIT which was able to agressively optimize the code to the extent that it was actually faster than running natively in not insignificant number of (real world!) cases.