Here a "feature" might be seen as an abstract, very, very high dimensional vector space. The team is pretty deep in investigating the idea of superposition, where individual neurons encode for multiple concepts. They experiment with a toy model and toy data set where the latent features are represented explicitly and then compressed into a small set of data dimensions. This forces superposition. Then they show how that superposition looks under varying sizes of training data.
It's obviously a toy model, but it's a compelling idea. At least for any model which might suffer from superposition.
https://transformer-circuits.pub/2023/toy-double-descent/ind...
Wonder if it's a matter of perspective - that is, of transform. Consider an image. Most real-world images have pixels with high locality - distant pixels are less correlated than immediate neighbours.
Now take an FFT of that. You get an equivalent 2D image containing the same information, but suddenly each pixel contains information about every pixel of the original image! You can do some interesting things there, like erasing the centre of the picture (higher frequencies), which will give you blurred original image when you run FFT on the frequency-image to get proper pixels again.
Some more detail here: https://calculatedcontent.com/2019/12/03/towards-a-new-theor...
Honestly the fact that there doesn't seem to be a good explanation for this makes me think that we just fundamentally don't understand learning.