CoLT5: Faster Long-Range Transformers With Conditional Computation
arxiv.org
arxiv.org
Another recent paper on efficient architecture for long context lengths: https://arxiv.org/abs/2302.10866
Stable diffusion was possible because diffusion model architecture and then doing it in latent space instead of image space and then upscaling.
Of course someone still needed to train it, but they did it when it was theoretically possible.
> No model weights
For shame