https://github.com/stassa/metasplain
So that's "metasplain" a little program that explains "invented predicates", which is what I say above, predicates that are automatically constructed by a symbolic machine learning system in the process of learning. Think of them as invented features that are relevant to the learning task. These are given automatic names so they're difficult to read, especially if you have lots of them. Metasplain starts by automatically assigning meaningful names to invented predicates by combining the (not invented) symbols of their literals, then asks the user for improved names. It can go all the way automatically, without interaction, but the results are a bit meh. With a human in the loop you get the best of both worlds.
And that's what I think is the best way to solve interpretability problems in machine learning: instead of automating them, which is like trying to create a chicken so you can get an egg so you can hatch a chicken, put the human back in the loop and make it easy for her to provide meaningful explanations, even if she's not an expert.