Better interpretability, I suppose. Could give insights into how cognition works.
Better interpretability, I suppose. Could give insights into how cognition works.
The gains in parameter efficiency was a surprise even to us when we first tried it out.
More work is being done on this as we speak.
I.e., is this something that could (and therefore, will) be turned towards identifying toxic concepts as understood by the chinese or us government, or to identify (say) pro-union concepts so they can be down-weighted in a released model, etc?
This is true. The features closer together now have much stronger semantic overlap. You can watch how the weights self-organize in a GPT here: https://toponets.github.io/webpage_assets/banner_video.mp4
We're already studying the effects of topographic structure on polysemanticity.