But it's more specifically focused on virtual machines than physical hardware.
http://users.ece.cmu.edu/~koopman/stack_computers/index.html
My memory of the conclusion is that register machines are a bit faster for general programs. Stack machines are better in some special niches. (High volume/rapid Interrupt handling comes to mind)
I should read the book again myself...
If you examine any stack language code, you should consider any stack manipulation to be a NOP type of inefficiency. And when a virtual stack machine is implemented on non-stack hardware, these inefficiencies are compounded.
We've switched to RISC (despite the above I'm a big fan have built several) these days largely because there came a point where we could push everything onto a chip, RISC started to make sense at about the time where cache went on-chip (or was SRAM closely coupled to a single die) - and for the record I think x86 has survived because its ISA was the most RISCy of it's original stable-mates (68k, 32k, z8k etc) - x86 instructions make at most 1 memory access (with one exception) with simple operands
Stack machines made more sense when memory was limited (small opcodes) and there were no hardware multiplies, but these days they make no sense.
I don't understand all the vague nostalgia for something that never could have panned out outside of the creaky old Apollo flight computer or something.
I'd kinda like to see a machine with a intermediate, one-operand style of instructions. Eg:
add tos stN # *sp += sp[N]
add stN pop # sp[N] += *sp++
ld [stN] # *--sp = mem[sp[N]]