1. Modern CPU microarchitectures issue multiple instructions/micro-instructions per clock.
2. Modern CPU microarchitectures rely heavily on register renaming in order to avoid bottlenecking on register allocation.
3. High code density is a huge win because: a) highly dense code reduces the cache footprint for a given amount of expressed functionality, which in turn reduces the amount of die area that need be dedicated to instruction cache in a well-balanced processor. b) highly dense code reduces the amount of bus traffic filling the instruction cache from main memory, thus reducing the overall memory bandwidth required to execute a given amount of functionality. Taking a & b together, they explain, in part, why X86 is still king of the hill, and it certainly explains why X86 won over Itanium. X86 code, for all its accumulated cruft, is very dense.
4. Stack machine code is also very dense, because exactly zero bits are dedicated to register names in every instruction.
So... given that the register numbers in typical 2-address or 3-address machines are getting renamed away anyway, why have them at all? Let the compiler generate 0-address code for a stack machine, and let the instruction decoder parse a long string of tokens during a single clock, using standard register renaming techniques to assign virtual stack locations from a large buffer pool. It could issue multiple stack ops per clock, but unwound into simple 3-address micro-ops for the CPU back end. It's really just the final step in moving the register allocator from the compiler to the CPU.
These are all solved problems for the X86 instruction set. Given that existence proof, it is difficult to argue that it would be hard to do for a clean stack-machine ISA.