Finally someone who did this, I've always thought this was a low hanging fruit. I wonder if you could make interesting and quickly trained diffusion models with this trick.
Note that it is from 2018. As someone here already mentioned there is a paper that applies the same idea to Vision Transformers published this year [1].