At the same time, presumably, common processor architectures have all undergone decades of targeted optimisation to improve their performance on the typical imperative code that is encountered in system and user programs. What are the implications of this? Should we expect a theoretical cap on the real-world performance of concatenative programs that is well below what we can achieve with mainstream programming languages?