[1] Someone else please speak up, but the Pentium II seems to be when Intel began leaning on microcode to optimize execution. https://en.wikipedia.org/wiki/Intel_Microcode
[2] Not always, of course. Sometimes local, runtime information is more important, as JIT enthusiasts love to point out. The problem is that the instrumentation to analyze and utilize that information is costly and incurred perpetually, whereas with static compilation it's a one-time deal.[3] As Rust and C++ has proven, many developers are willing to put up with several orders of magnitude more costs up-front than could ever be tolerated in a JIT, especially the JITs inside the CPU.
[3] This is a similar issue to GCs (mark & sweep and reference counting), where memory must be scanned and sometimes even copied many times during runtime whereas in static memory management there's exactly one operation for allocation and one for deallocation. So GC is usually (but, again, not always) slower. Optimizations like bump allocations (which in static compilation is often simply "the stack") that aggregate and amortize some costs don't change the fundamental costs.
Most important instructions map 1:1 to CPU micro-ops, with memory operands (on x86) being mapped to separate load/store micro-ops. There are cases where instructions map to long sequences (e.g. VMENTER) or get combined into one micro-op (e.g. compare-and-branch), but this is far from a JIT compiler.
From a wholistic viewpoint I think it's fair to characterize modern CPUs as JIT-like. It's easy to trivialize individual aspects of the whole. You can do the same for software JITs. They're not magic (notwithstanding the breathless claims and river of thesis papers), and often simpler is better, anyhow (e.g. LuaJIT). For obvious reasons the tactics employed by CPUs will tend toward the simpler end of the spectrum.
I'm trying to figure out if there is anything generally useful for now in these papers.
Wasn’t that the Pentium Pro?