Zoom In: Speculative claims about neural circuits
distill.pub
distill.pub
The literature on this is too hard to summarize in a post, but basically in turns into an empirical-scientific question, of making predictions about model features and testing these predictions scientifically.
A notable researcher privately told me that they think all interpretability research is nonsense. As someone who's dedicated the last six years of my life to this field, that was pretty uncomfortable to hear. But I think it's important to pay attention to, because I think it's actually a pretty common, unspoken view.
As a result, this has been on my mind a great deal. I think two important questions are:
(1) How can we surface the disagreements that are leading to such divergent views between different members of the research community? (Especially when people are generally too polite to say that they think something is total nonsense.)
(2) What would a more epistemically stable foundation for interpretability look like?
I'm not sure what the right answers to these are, but I think they're important to discuss.