But the neuro-symbolic concept learning paper I gave already shows the potential translucency of these kinds of hybrid systems: the linguistic interface (its a VQA task) allows one to simply look up the feature vectors and programs associated with particular phrases or English nouns. Similarly, the scene is parsed in an explicitly interpretable way, with bounding boxes for the various objects on which reasoning will commence. This 'bridge' theme between natural language and the underlying task space is really powerful, and it probably makes sense to figure out how we include them for systems that have nothing to do with natural language.
https://arxiv.org/abs/1901.11390 contains another great example of how interpretable such models can be, especially if they are generative. Take a look at those segmentations!
Lastly, https://arxiv.org/abs/1604.00289 lays out this vision in a lot more detail.
https://github.com/stassa/metasplain
So that's "metasplain" a little program that explains "invented predicates", which is what I say above, predicates that are automatically constructed by a symbolic machine learning system in the process of learning. Think of them as invented features that are relevant to the learning task. These are given automatic names so they're difficult to read, especially if you have lots of them. Metasplain starts by automatically assigning meaningful names to invented predicates by combining the (not invented) symbols of their literals, then asks the user for improved names. It can go all the way automatically, without interaction, but the results are a bit meh. With a human in the loop you get the best of both worlds.
And that's what I think is the best way to solve interpretability problems in machine learning: instead of automating them, which is like trying to create a chicken so you can get an egg so you can hatch a chicken, put the human back in the loop and make it easy for her to provide meaningful explanations, even if she's not an expert.
ML models work in the same way. Difference is they aren't human so it's harder to just use your empathy muscles (though even humans cut off from the rest of the world can come to pretty wild perspectives and ways of thinking). But they are logical, in some respect, and they are modeling our human perspective on some process in the world. But as a sibling comment posted in a link to the Distill paper, we need a lot of tools to make that process easier.
For example, very often researchers will probe single neurons and find they do something we find conceptually understandable, like a neuron detecting when you're inside a parenthetical when generating text so that the parenthesis is eventually closed. I'd expect the "symbols" to be very similar, because after all the neurons are symbols too. Both require you to either relate it to a small concept or put some together to create bigger concepts (or both).
(By the way, probably you have the coolest job on earth)