Toy Models of Superposition (2022)
transformer-circuits.pub
transformer-circuits.pub
> the neural networks we observe in practice are in some sense noisily simulating larger, highly sparse networks
This seems somewhat related to a point made by Ilya Sutskever here [1]: NNs can be though of an approximation to the Kolmogorov compressor. Speculating, one could say any network is a projection of the ideal compressor (which argualy perfectly represents all n-features in an n-dimensional ONB) into a lower dimentional space, hence the interference. But why is there not always such an interference?
To the authors if they happen to find themselves here, I say: bravo!