For one example, about 25 years ago Intel made a processor called Itanium which used 'VLIW' (very long instruction word) instructions as opposed to x86. In this case, a lot of the complexity of operating a processor is shifted into the compiler, since individual CPU instructions may perform complex operations. In a modern processor, instructions are 'simple', and the processor is responsible for finding the data dependencies that allow multiple instructions to be executed in parallel (this is called 'superscalar' processing). With Itanium, the compiler author can encode such individual operations into more complex instructions which can be executed in parallel, all at once. So, where in x86 I might have in 3 different instructions (add r1, r2, sub r3, r4, mul r5, r6), with Itanium I might be able to combine all of those into a single instruction at compile time. By doing this, in theory we can remove a lot of the complex superscalar logic from the processor, devote more of the chip to execution units and achieve more overall performance.
One problem here is that the compiler may not actually know what parallelism is possible. For example, the latency of a given instruction may vary from chip to chip. The number of execution units for a single operation may vary. Our code is now very highly optimised for a particular CPU, we may not see natural speedups as newer processors are made, newer processors might not even support the code we've written without a translation layer. Taken to extremes, we lose the universality of software (the ability to change the hardware without rebuilding the software). VLIW chips are viable, but they're typically used in more isolated use cases (like DSPs) where it's reasonable to recompile the software for every chip (and where perhaps there is lots of parallelism in the problem space).
State machines essentially have similar issues. High performance processors tease out the data dependencies between instructions in order to perform superscalar processing, which is key to good performance in a modern processor. Stack machines make it very difficult to understand data dependencies because in principle all data depends on all other data, a post-processing step is required (which is considerably more expensive than the processor equivalent of converting to SSA form). Modern register machines rely heavily on register renaming to ensure that both superscalar processing and pipelining work well with each other - maintaining similar performance with a stack machine would lead to similar solutions, and so very likely switching from a register machine to a stack machine would likely look like having a translation layer on the chip, from the stack machine back into a register-like language.
Stacks do exist in a few places these days. The JVM clearly does it - Java is 'register oriented' for the programmer, the JVM instruction set is stack oriented, and the runtime converts it back into a register oriented machine code. Another example is the ARM Jazelle extensions which allow some ARM processors to directly execute Java bytecode. Here, the bytecode is converted into the machine code as an extra stage in the processor pipeline, adding overhead.
BLAB: - There isn't much advantage to stack machines. It's relatively straightforward to translate between the two. - Modern processors are architecturally closer to register machines, and a lot of the techniques used to make them fast would have to be replicated on a stack machine. - Because any decisions at this level of the stack are effectively invisible to the user in modern tech, it's beneficial to go for the simplest solutions to any given problem, which here is to use a register machine at the hardware level.