The key problems with this approach is not mentioned in the blog post, but is shown in figures 6 and 7 of the paper - https://stefan-marr.de/downloads/acmsac23-huang-et-al-optimi...
Basically, the code handler ordering does not generalise well across benchmark nor processor, so to get the speedup they see, you'd need a specialised interpreter for your specific benchmark and processor. That puts this into the "interesting, but not very practical" category.