Google, in collaboration with OpenAI, has published an impressive tour de force where they have throughly developed and validated at scale a sparse transformer architecture, applied to general language modeling task: https://arxiv.org/abs/2111.12763
This happened in November of 2021, and there is a public implementation of this architecture on the Google's public github.
Impressively, due to some reasons, other up-and-coming players are still not releasing models trained with this approach, even though it promises multiplicative payoff in inference economy. One boring explanation is conservatism for NN training at scale, where training runs cost O(yearly salary).
Let's hope the open source side of things catches up.
Of course from a memory bandwidth and model size perspective maybe there are benefits long before that.
(Edit: just to be clear here, I’m not saying I expect the whole field is full of dummies who missed something obvious or something like that, I don’t know much at all about machine learning so I’m sure I’m missing something).
I'm not saying it's impossible but if resources allow it makes a lot of sense to start with the biggest model you can still train. Especially since for whatever reason things seem to get a lot easier if you simply throw more computing power at it (kind of like how no matter how advanced your caching algorithm it's not going to be more than 2 times faster than the simplest LRU algorithm with double the amount of cache).
It sounds like there is a lot of work happening on sparse networks now, so it’ll be interesting to see how this changes in the near future.
There has been a lot of work on sparsity and discovering sparse subnetworks in trained dense networks. And intel even proposed some alternative cpu friendly architectures and torch/tf and gpus are starting to do okay with sparse matrixes so thing are changing.
A 4-bit quantize of a 1000B model should be <600GB, so would fit on a regular 8x80 DGX system.
Benchmarks: https://github.com/ggerganov/llama.cpp/issues/34