I've spent the last decade working on and thinking about JITs, runtime hidden type modeling, and other optimization problems relating to reactive dynamic systems dealing with heavy polymorphism (specifically Javascript as a programming language).
Now, different people in the industry will have different perspectives, so I'll preface this by saying that this is _my_ view on where the future of JIT compilation leads.
The order of magnitude gains were mostly plumbed with the techniques JIT engineers have brought into play over the last decade or so.
One aspect that remains relevant now is responsiveness to varying workloads, and nible reaction to polymorphic code, and finding the right balance between time spent compiling and the productivity of the optimized code that comes out the other end. There is significant work yet to be done in finding the right runtime type models that quickly and effectively distinguish between monomorphism, polymorphism, and megamorphism, and are able to respond appropriately.
The other major avenue for development, and the one that I have yet to see a lot of talk about, is the potential to move many of these techniques and insights out of the domain of language runtime optimization, and into _libraries_, and allowing developers direct API-level access to the optimization strategies that have been developed.
If you work on this stuff, you find very quickly that a huge amount of a JIT VM's backend infrastructure has nothing to with _compilation_ per se, and much more to do with the support structures that allow for the discovery of fastpaths for operations over "stable" data structures.
In Javascript, the data structure that is of most relevance is a linked list of property-value-maps (javascript objects linked together by proto chains). We use well-known heuristics about the nature of those structures (e.g. "many objects will share the _keyset_ component of the hashtable", and "the linked list of hashtables will organize itself into a tree structure"). Using that information, we factor out common bits of the data-structure "behind the scenes" to optimize object representation and capture shared structural information in hidden types, and then apply techniques (such as inline caches) to optimize on the back of that shared structure.
There's no reason that this sort of approach cannot be applied to _user specified structures_ of different sorts. Different "skeletal shapes" that are highly conserved in programs.
For me, the big promise that these technologies have brought to the fore is the possibility of effectively doing partial specialization of data structures at runtime. Reorganizing data structures under the hood of an implementation to deliver, transparently to the client developer, optimization opportunities that we simply don't even consider as possible today.