Firstly, every microarchitecture will have a different optimal schedule. This means a significant shift in how binaries are deployed. Either the compiler is built into a distribution system aware of microarchitectures (for example, bitcode and the Apple App Store), or these tasks are brought online with a JIT compiler. For the sake of software freedom I hope for the latter.
Secondly, this paper shows 88% of the speedup is static. That means there will always be a little extra speed available to a chip maker that ships an OOO processor. For the same dollars, will you take the CPU that is 10% slower. Today in the middle of Intel's security disaster, sure. But ten years from now when CPU security is stable again, that free 10% performance will be compelling.