> Models can be of arbitrarily large sizes to the point where people really can’t understand them. How do you go about dissecting a 2 layer NN with 10^40th nodes?
I'd probably take the output of the first layer, and cluster it.
It would very much depend on what this NN was attempting to do.
Like, all of the work in this field does suggest that people value this feature, and its incredibly useful for debugging which is generally pretty hard in ML/statistics.
> It’s a vastly larger solution space. So it’s really the reverse that would be surprising.
Imagine a world in which linear models returned the entire matrix by observation rather than the coefficients. People would argue that it was uninterpretable, but it's a problem of tools.
I actually think that if you can instrument a model appropriately, then you can definitely build an interpretation layer on top of it. Clearly that doesn't make the model perform worse.
Even if you can't instrument it, you can run thousands of experiments changing one feature at a time and then estimate the impact of this feature on the model. Granted, that's not practical on many problems, but neither were deep NN's a decade ago.