Friends don’t let friends train small diffusion models
nonint.com
nonint.com
Here's my intuition. The output of the first layer is mostly a linear matrix transform. These linearities form the basis for the more complex features. If the network doesn't have enough of these base linear features, it can't compose the higher-level (and lower-dimensional) features later down the line.
Ideally the training just derives non-intuitive relationships between these decorrelated input features to give you something useful.
Very cool work being done in ML these days, both inside large corps, and also by independent researchers/"hobbyists" (quotation marks because there is some really expert work being produced by them).
And from yesterday, an open-source version of Google's Imagen by lucidrains: https://github.com/lucidrains/imagen-pytorch
This is just begging to be text-guided with the lyrics. Then you could put in the lyrics and have the model generate a clip it thinks would go with the lyrics. :)
>To test out this theory, I needed a way to compress the output of the diffusion model. One idea that came to mind is using the concepts behind PixelShuffle. So I compressed the waveform by 16x using a 1D form of PixelShuffle, increased the base dimensionality to 256, and started another training run.
I'm not familiar with the author's previous projects or the terminology here. Do they just mean he increased the size of the first hidden layer and (base dimensionality to 256) and compressed the output layer by 16?
Utterly amazing.
So, are there some real words in there? Do the lyrics make sense? Or is it just "random" vocalizations that go along with the song?