Part of what we are learning, in AI/ML, is the number of problems where the ML is relevant and where it is not. (Or, to put it another way, where we can find statistical relevance between features and targets).
So I think you need some kind of "meta-metric" that measures the growth of the taxonomy itself. And perhaps some kind of weighting for the impact of the solution.
There is also an interaction effect (for instance, Natural Language Processing is powerful, and "common sense reasoning" is powerful, but put them together and you have a knockout), but I don't know how to go about measuring that.